Method for generating harmonious objects in a 3D scene based on physically based rendering and uncertainty estimation
By adopting a method based on physical rendering and uncertainty estimation in a three-dimensional scene, combining scene semantic information and ambient lighting, the generated 3D objects are optimized, and the problem of disharmonious fusion between objects and scenes in the prior art is solved, and the quality and fidelity of generated objects are improved.
Patent Information
- Application Number
- CN202510296870.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The prior art does not fully utilize scene semantic information and ambient lighting when generating objects in three-dimensional scenes, resulting in disharmonious integration between objects and scenes and low generation quality.
Using a method based on physical rendering and uncertainty estimation, a diffusion model is generated by building feature networks and multi-view images, combining scene semantic information and ambient lighting, the generated 3D objects are optimized to keep them harmonious with the scene in multiple perspectives.
It effectively solves the problem of dissonance between objects and scenes, improves the quality and fidelity of generated objects, and allows objects to be naturally and seamlessly integrated into the scene.
Smart Images

Figure CN119850847B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and generative models, and relates to a method for generating harmonious objects in a three-dimensional scene, and particularly relates to a method for generating harmonious objects in a three-dimensional scene based on physically based rendering and uncertainty estimation. Background Art
[0002] In recent years, with the development of the fields of computer vision and generative models, technologies such as text-driven 3D object generation and 3D scene editing have emerged and are widely used in multiple fields such as film and television, games, architectural design, VR / AR, etc. In these applications, by inputting specific text prompts and other information by the user, 3D objects desired by the user can be generated at specified positions in the scene, greatly improving the modeling efficiency and reducing costs. However, although certain achievements have been made in the research on generating 3D objects in a three-dimensional scene, it is still a difficult problem to make the generated 3D objects visually natural and harmonious with the scene. Some existing methods, such as:
[0003] An article in 2024, InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes, proposed a method for inserting 3D objects generated by text into a three-dimensional scene reconstructed by a neural radiance field (NeRF). This method first takes the input rectangular box and prompt text given by the user in the three-dimensional scene as the basis, obtains a single image of the object harmonious with the scene through an image completion prior model, then uses an existing method for reconstructing 3D objects from a single image to obtain the 3D object, and then rescales and places the object at the corresponding position in the scene according to the estimated object ratio and depth. However, although this method generates a 2D object harmonious with the scene from a single perspective using the image completion method, the subsequent process of reconstructing the 3D object from a single image cannot guarantee the harmony of the object with the scene from multiple other perspectives, resulting in the final generated 3D object not being harmonious and natural with the scene.
[0004] A 2024 article "GO-NeRF: Generating Virtual Objects in Neural Radiance Fields" proposed a method to generate 3D objects that are harmonious with the 3D scene under multi-view conditions. By replacing the commonly used Score Distillation Sampling Loss (SDS Loss) in the text-to-3D process with the Inpainting Score Distillation Sampling Loss (Inpainting SDS Loss), this method utilizes a multi-view image inpainting prior model to gradually optimize the generated objects in the scene. Additionally, some traditional image harmonization loss functions, such as saturation loss and style consistency loss, are used to optimize the harmony between the objects and the scene. Although this method considers the harmony between the generated objects and the scene from multiple perspectives, it ignores the environmental lighting of the scene. Moreover, traditional image harmonization losses have drawbacks such as reduced color accuracy, poor generalization ability, and possible detail loss, which can easily lead to suboptimal or low-quality generation results. Summary of the Invention
[0005] Aiming at the problems in the prior art that when generating objects in a 3D scene, the scene semantic information is not fully utilized and the environmental lighting is ignored, resulting in disharmonious fusion between the object and the scene and low generation quality, the present invention aims to provide a new method that can fully utilize the scene semantic information and environmental lighting, and combine and integrate the scene semantic information and environmental lighting into the generated 3D object through estimated uncertainty to ensure that the object naturally integrates into the scene, remains harmonious with the scene under multiple perspectives, and at the same time improves the quality and fidelity of the generated object.
[0006] The present invention proposes a method for generating harmonious objects in a 3D scene based on physically based rendering and uncertainty estimation. This method mainly models the geometry and material of the generated 3D object through physically based rendering, then re-illuminates the 3D object according to the given scene, and combines the estimated uncertainty to perform harmonious optimization on each perspective of the 3D object in combination with the scene, ensuring that the generated object is naturally harmonious with the scene and improving the quality and fidelity of the generated object.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A method for generating harmonious objects in a 3D scene based on physically based rendering and uncertainty estimation, comprising the following steps:
[0009] Given a 3D scene, a text description of the target object to be generated, and a rectangular frame indicating the position of the target object in the 2D rendering view of the 3D scene;
[0010] According to the text description of the generated target object, a deformable tetrahedral mesh of the geometric appearance of the target object is constructed based on a two-stage method using a pre-trained multi-view image generation diffusion model;
[0011] Construct a feature network, the feature network includes a material term prediction module and an uncertainty term prediction module; the input of the feature network is the sampling points on the surface of the deformable tetrahedral mesh, and the output is the material term and the uncertainty term of the object;
[0012] Use physical rendering technology to obtain the rendered picture of the target object, and train and optimize the feature network;
[0013] According to the three-dimensional scene, the deformable tetrahedral mesh and the material term output by the trained feature network, obtain the multi-view scene rendering dense view and the 3D model file containing the target object.
[0014] Furthermore, the three-dimensional scene is a pre-trained three-dimensional scene represented by a neural radiance field NeRF.
[0015] Furthermore, the specific steps of constructing a deformable tetrahedral mesh of the geometric appearance of the target object based on a two-stage method using a pre-trained multi-view image generation diffusion model according to the text description of the generated target object are as follows:
[0016] Construct a 3D target object represented by NeRF;
[0017] Taking the camera parameters as the input, use the multi-layer perceptron of NeRF to process the input, output the color field and the density field, and obtain the 2D picture img1 of the 3D object represented by NeRF under the camera parameter conditions according to the color field and the density field through volume rendering technology;
[0018] Taking the text description of the generated target object and the camera parameters as the input, use the pre-trained multi-view image generation diffusion model to output the 2D picture img0 of the target object that conforms to the text description under the camera parameter conditions;
[0019] Calculate the SDS loss (Score Distillation Sampling Loss) using the picture img1 and the picture img0, train the 3D target object represented by NeRF, and obtain the 3D target object represented by NeRF with a rough geometric appearance;
[0020] Extract the initial deformable tetrahedral mesh from the 3D target object represented by NeRF with a rough geometric appearance;
[0021] Obtain the normal map normal1 of the initial deformable tetrahedral mesh and the normal map normal2 of the 3D target object with a rough geometric appearance under the camera parameter conditions;
[0022] Calculate the SDS loss using the picture img0 and the normal map normal1, calculate the MSE loss (Mean Squared Error Loss) using the normal map normal1 and the normal map normal2, and train the initial deformable tetrahedral mesh to obtain a deformable tetrahedral mesh including the refined geometric appearance of the target object.
[0023] Further, the specific steps for extracting the initial deformable tetrahedral mesh from the 3D target object with a rough geometric appearance represented by NeRF are as follows:
[0024] Convert the density field of the 3D target object with a rough geometric appearance represented by NeRF into a signed distance function SDF (Signed Distance Function), subtract the signed distance function SDF with a non-zero constant to obtain the initial SDF, and use the differentiable tetrahedral marching algorithm (Deep Marching Tetrahedra, DMTet) based on the initial SDF to obtain the initial deformable tetrahedral mesh.
[0025] Further, the feature network includes an input layer, a hash network encoding layer, a linear layer, an activation function layer, and an output layer.
[0026] Further, the material terms include albedo, metallicity, and roughness.
[0027] Further, the specific process of obtaining the rendered picture of the target object using physical rendering technology specifically includes:
[0028] Based on the image-based lighting technology (IBL), calculate the ambient light on the surface of the target object according to the environment map, combine the material terms and the ambient light, and obtain the rendered picture of the target object through the physical rendering technology (PBR).
[0029] Further, the specific process of training and optimizing the feature network specifically includes:
[0030] Randomly initialize the environment map, obtain the initial rendered picture of the target object according to the initial environment map and the initial material terms output by the feature network, and use the initial rendered picture to train the material term prediction module of the feature network through the material term loss function;
[0031] Obtain the spatial coordinates of the target object in the 3D scene based on the 3D scene, the initial rendering image of the target object, and the rectangular frame of the position of the target object in the 2D rendering view of the 3D scene;
[0032] Obtain the scene environment texture map according to the spatial coordinates, and obtain the new rendering image of the target object according to the scene environment texture map and the material item output by the trained feature network;
[0033] Use the pre-trained 2D image harmonization network to obtain the result image after harmonizing the target object and the 3D scene;
[0034] Calculate the uncertainty term loss and the harmonization loss according to the new rendering image of the target object and the result image after harmonization, and train the albedo part of the uncertainty term prediction module and the material item prediction module of the feature network.
[0035] Further, obtaining the spatial coordinates of the target object in the 3D scene according to the 3D scene, the rendering image of the target object, and the rectangular frame of the position of the target object in the 2D rendering view of the 3D scene specifically includes:
[0036] Obtain the 2D rendering image and the scene depth map of the 3D scene;
[0037] According to the rectangular frame of the position of the target object in the 2D rendering view of the 3D scene, use the 2D rendering image of the 3D scene as the background and the rendering image of the target object as the foreground to construct a fused image;
[0038] Use the monocular depth estimation method to estimate the depth map of the fused image, and use the least squares method to align the depth map of the fused image and the scene depth map to determine the depth of the foreground target object;
[0039] Back-project the center point of the rectangular frame into the 3D scene, and calculate the coordinates of the center point according to the depth of the foreground target object as the spatial coordinates of the target object in the 3D scene.
[0040] Further, the specific steps of using the pre-trained 2D image harmonization network to obtain the result image after harmonizing the target object and the 3D scene include: inputting the new rendering image of the target object and the 2D rendering image of the 3D scene under the same camera parameter conditions into the pre-trained 2D image harmonization network, and outputting the result image after harmonizing the target object.
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] 1. The method proposed by the present invention makes full use of scene semantic information and environmental illumination, and integrates the two into the generated 3D object through the estimated uncertainty, which can effectively solve the problem of disharmonious integration of the target object and the scene in the prior art, ensure that the generated target object naturally and seamlessly integrates into the scene, and enhance the realism and visual effect of the target object in the scene.
[0043] 2. The technical means adopted by the present invention can decouple the geometry and material modeling of the target object. During the generation process, through a carefully designed optimization method, it can ensure the generation of high-quality, high-fidelity and multi-view consistent geometry and material information, overcoming the problems of low quality and poor fidelity of the generated target object in the prior art.
[0044] 3. The present invention realizes a simple user interaction design. The user only needs to provide a text description and a rectangular frame of the position of the target object in the 2D picture of the scene rendering view. Through spatial position and scale calculation, the target object can be accurately placed in the 3D space. According to the depth calculation, the occlusion relationship is reduced, and the disharmonious occlusion caused by perspective switching is reduced, enhancing the adaptability and stability of the object in the complex 3D scene, and improving the flexibility and creativity of scene editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is the overall flow chart of the generation method in the embodiment of the present invention.
[0046] Figure 2 is the sub-flow chart of the geometric learning stage of the generation method in the embodiment of the present invention.
[0047] Figure 3 is the sub-flow chart of the material modeling stage of the generation method in the embodiment of the present invention.
[0048] Figure 4 is the structural diagram of the feature network for predicting materials and uncertainties of the generation method in the embodiment of the present invention.
[0049] Figure 5 is the sub-flow chart of the positioning of the object in the scene of the generation method in the embodiment of the present invention.
[0050] Figure 6 is the sub-flow chart of the harmonization stage of the generation method in the embodiment of the present invention.
[0051] Figure 7 is the test result diagram (1) of generating a "pumpkin head" in the embodiment of the present invention.
[0052] Figure 8 is the test result diagram (2) of generating a "pumpkin head" in the embodiment of the present invention.
[0053] Figure 9It is the result diagram (3) of generating "Pumpkin Head" in the embodiments of the present invention.
[0054] Figure 10 It is the result diagram (1) of generating "Charmander" in the embodiments of the present invention.
[0055] Figure 11 It is the result diagram (2) of generating "Charmander" in the embodiments of the present invention.
[0056] Figure 12 It is the result diagram (3) of generating "Charmander" in the embodiments of the present invention. Detailed implementation manners
[0057] The technical solution of the present invention will be further clearly and detailedly described below in conjunction with the accompanying drawings and specific examples.
[0058] A method for generating harmonious objects in a 3D scene based on physically based rendering and uncertainty estimation provided in this embodiment includes the following steps:
[0059] S1. Given a 3D scene, a text description of the target object to be generated, and a rectangular frame of the position of the target object in the 2D rendering view of the 3D scene; the 3D scene is a pre-trained 3D scene represented by a neural radiance field NeRF.
[0060] S2. According to the text description of the target object to be generated, use a pre-trained multi-view image generation diffusion model to construct a deformable tetrahedral mesh of the geometric appearance of the target object based on a two-stage method; the specific steps are as follows:
[0061] S2.1. Construct a 3D target object represented by NeRF.
[0062] S2.2. Using the camera parameters as input, process the input by the multi-layer perceptron of NeRF, output the color field and density field, and obtain the 2D picture img1 of the 3D object represented by NeRF under the camera parameter conditions according to the color field and density field through volume rendering technology;
[0063] S2.3. Using the text description of the target object to be generated and the camera parameters as input, use a pre-trained multi-view image generation diffusion model to output the 2D picture img0 of the target object that conforms to the text description under the camera parameter conditions;
[0064] S2.4. Calculate the SDS loss (Score Distillation Sampling Loss) using the picture img1 and the picture img0, and train the 3D target object represented by NeRF to obtain a 3D target object represented by NeRF with a rough geometric appearance;
[0065] Extract an initial deformable tetrahedral mesh from the 3D target object represented by the NeRF with a rough geometric appearance. Specifically, convert the density field of the 3D target object represented by the NeRF with a rough geometric appearance into a signed distance function (SDF), subtract the signed distance function SDF from a non-zero constant to obtain an initial SDF, and use the differentiable tetrahedral marching algorithm (Deep Marching Tetrahedra, DMTet) based on the initial SDF to obtain an initial deformable tetrahedral mesh.
[0066] Obtain the normal map normal1 of the initial deformable tetrahedral mesh and the normal map normal2 of the 3D target object represented by the NeRF with a rough geometric appearance under the camera parameter conditions.
[0067] Calculate the SDS loss using the picture img0 and the normal map normal1, calculate the MSE loss (Mean Squared Error Loss) using the normal map normal1 and the normal map normal2, and train the initial deformable tetrahedral mesh to obtain a deformable tetrahedral mesh including the refined geometric appearance of the target object.
[0068] S3. Construct a feature network including an input layer, a hash network encoding layer, a linear layer, an activation function layer, and an output layer. The feature network includes a material term prediction module and an uncertainty term prediction module. The input of the feature network is the sampling points on the surface of the deformable tetrahedral mesh, and the output is the material term and uncertainty term of the object. The material term includes albedo, metallicity, and roughness.
[0069] S4. Use physical rendering technology to obtain the rendered picture of the target object and train and optimize the feature network.
[0070] The specific process of using physical rendering technology to obtain the rendered picture of the target object includes: calculating the ambient light on the surface of the target object based on the environment map using image-based lighting (IBL), and combining the material term and the ambient light to obtain the rendered picture of the target object through physical-based rendering (PBR).
[0071] The specific process of training and optimizing the feature network includes:
[0072] Randomly initialize the environment map, obtain the initial rendered picture of the target object according to the initial environment map and the initial material term output by the feature network, and use the initial rendered picture to train the material term prediction module of the feature network through the material term loss function.
[0073] Obtain the spatial coordinates of the target object in the 3D scene based on the 3D scene, the initial rendered image of the target object, and the rectangular box of the position of the target object in the 2D rendered view of the 3D scene; specifically including: obtaining the 2D rendered image and the scene depth map of the 3D scene; according to the rectangular box of the position of the target object in the 2D rendered view of the 3D scene, using the 2D rendered image of the 3D scene as the background and the rendered image of the target object as the foreground, construct a fused image; use the monocular depth estimation method to estimate the depth map of the fused image, use the least squares method to align the depth map of the fused image and the scene depth map, and determine the depth of the foreground target object; back-project the center point of the rectangular box into the 3D scene, and calculate the coordinates of the center point according to the depth of the foreground target object, which is used as the spatial coordinates of the target object in the 3D scene.
[0074] Obtain the scene environment texture map according to the spatial coordinates, obtain the new rendered image of the target object according to the scene environment texture map and the material item output by the trained feature network, use the pre-trained 2D image harmonization network to obtain the result image after harmonizing the target object and the 3D scene, calculate the uncertainty term loss and the harmonization loss according to the new rendered image of the target object and the result image after harmonization, and train the albedo part of the uncertainty term prediction module and the material item prediction module of the feature network.
[0075] S5. Obtain the multi-view scene rendering dense view and the 3D model file containing the target object according to the 3D scene, the deformable tetrahedral mesh, and the material item output by the trained feature network.
[0076] In another embodiment of the present invention, as Figure 1 shown, a method for generating a harmonized object in a 3D scene based on physically based rendering and uncertainty estimation includes the following steps:
[0077] Step 1, obtain the input: The input is a given 3D scene, the text description of the generated target object, and the rectangular box of the position of the target object in the 2D rendered view of any perspective of the 3D scene.
[0078] (1.1) The 3D scene is a pre-trained 3D scene represented by NeRF. According to the NeRF characteristics, the sampled spatial coordinates, azimuth angle, and tilt angle can be used to calculate the 2D rendered view and the depth map of any perspective of the scene through the volume rendering integral equation, which is convenient for subsequent use in calculating the spatial position and scale of the object in the scene during the harmonization stage;
[0079] (1.2) Generate a text description of the target object, which will be used as the input of the image generation diffusion model to calculate the SDS loss and guide the generation of the 3D object. It should be noted that to avoid the common "multiple heads" problem in text-to-3D, that is, due to factors such as unstable models or improper parameter settings, the generated results are complex and chaotic and lack consistency. All subsequent processes use an existing multi-view image generation diffusion model (such as MVDream) pre-trained on a large-scale three-dimensional shape dataset;
[0080] (1.3) The user needs to mark the position where the target object is desired to be placed in the 2D rendering view of a certain perspective in the three-dimensional scene by drawing a rectangular box. Subsequently, when calculating the spatial position and scale of the target object in the scene, the position of the object will be converted to three-dimensional space coordinates through back-projection. This design of interacting in the 2D view enables users to complete complex 3D object generation and positioning tasks in a more intuitive and simple way, bringing a smoother and more efficient operation experience to users.
[0081] Step 2, geometric learning stage: According to the generated text description of the target object, use the pre-trained multi-view image generation diffusion model to construct a deformable tetrahedral mesh of the object's geometric appearance based on a two-stage method. Among them, in the first stage, the SDS loss is used to guide the generation of a 3D object represented by NeRF. In the second stage, a deformable mesh is extracted from NeRF and fine-tuned to improve the geometric accuracy of the generated object. The two-stage geometric learning method adopted by the present invention is similar to the existing method. The first stage generates a rough geometric appearance, and the second stage fine-tunes the rough geometry. However, compared with the existing method, the present invention adopts a carefully designed optimization method, which can improve the geometric accuracy of the generated object on the one hand and accelerate the training process to achieve fast convergence on the other hand.
[0082] Refer to Figure 2 , the implementation of this step includes the following:
[0083] (2.1) The first stage of geometric learning, generating a rough geometric appearance:
[0084] (2.1.1) When training a 3D object represented by NeRF, first sample the camera parameter c as the input. The camera parameter represents the perspective (denoted by the symbol v) of observing the 3D target object and is a 5D vector composed of spatial coordinates, azimuth angle, and tilt angle. Then use the multi-layer perceptron (MLP) of NeRF to process the input and output a color field and a density field, which respectively represent the color information and density distribution of the object in space. Through the volume rendering technique, the 2D picture img1 of the 3D target object represented by NeRF in the perspective v can be calculated;
[0085] (2.1.2) Taking the text description of the target object and the camera parameter c as inputs, through a pre-trained multi-view image generation diffusion model, an image img0 that conforms to the text description can be generated under the view v. The multi-view image generation diffusion model adds noise to the original image step by step and then learns to restore the data from the noise. In this process, the time step represents the degree of noise addition. In the initial stage of training, the range of the sampled time step is set to U(0.02, 0.98), and later it is reduced to U(0.02, 0.5). This is because in the initial stage of training, img1 is obtained through randomly initialized NeRF and does not conform to the text description. A larger time step is beneficial for NeRF to learn the general features of the generated object. Setting a smaller time step in the later stage is beneficial for NeRF to make fine adjustments;
[0086] (2.1.3) Use img0 and img1 to calculate SDS_Loss(img0, img1) and train NeRF to obtain the rough geometric appearance in the first stage;
[0087] (2.2) The second stage of geometric learning: Fine-tuning the rough geometry:
[0088] (2.2.1) Extract a deformable mesh from the NeRF representation generated in the first stage. First, subtract a non-zero constant from the density field of NeRF to obtain the initial SDF, and then use the differentiable tetrahedron marching algorithm to obtain the initial deformable tetrahedron mesh;
[0089] (2.2.2) Obtain the normal map normal1 of the initial deformable tetrahedron mesh and the normal map normal2 of NeRF under the condition of view v. The normal is directly approximated by calculating the gradient of the NeRF density field;
[0090] (2.2.3) Using the multi-view diffusion model and normal2 as supervision, calculate SDS_Loss(img0, normal1) and MSE_Loss(normal1, normal2) respectively to train the deformable tetrahedron mesh. In the second stage, using the normal map obtained by taking the gradient of the NeRF density field as supervision helps to accelerate the convergence speed of the mesh and can retain rich geometric details of NeRF.
[0091] Through the two-stage geometric modeling method, we can obtain high-quality 3D objects in the geometric learning stage. These objects not only accurately reflect the input text description but also have a reasonable and accurate geometric structure, laying a solid foundation for subsequent material modeling and harmonization stages.
[0092] Step 3, Material Modeling Phase: Construct a feature network F including a material item prediction module and an uncertainty item prediction module to predict the material items on the surface of the object, including albedo, metallicity, and roughness. Render the appearance of the 3D object based on Physically Based Rendering (PBR) and Image-Based Lighting (IBL), and use the SDS loss and other loss functions to optimize the material item prediction module.
[0093] Refer to Figure 3 , the implementation of this step includes the following:
[0094] (3.1) Predict the material items on the surface of the object. First, sample points on the mesh surface according to the camera parameter c, and then input them into the feature network F. The structure of the feature network F refers to Figure 4 as shown. First, perform hash grid encoding processing on the input data, then pass through a linear layer to map the input to a feature space with a dimension of 64 and apply the Rectified Linear Unit (ReLU). Then, pass through another linear layer to map the input to a feature space with a dimension of 5 to further transform the data. The left branch applies the Sigmoid Function to the output of the linear layer to obtain the output of the material item. The right branch first passes through a linear layer to map the input to a feature space with a dimension of 1, and then applies the Softplus Function to obtain the output of the uncertainty item.
[0095] By using hash grid encoding and an MLP with fewer channels and layers, the generation speed can be greatly improved while ensuring the generation quality. After being processed by the feature network F, the material items and uncertainty items on the surface of the target object are output. The feature network for generating the uncertainty item part will be trained and used in the harmonization phase;
[0096] (3.2) Obtain the initial rendered image img2 based on Physically Based Rendering (PBR). Specifically: Adopt the classic Cook - Torrance Bidirectional Reflectance Distribution Function (Cook - Torrance BRDF) model, divide the rendering of the target object surface into two independent parts: diffuse reflection and specular reflection, and calculate the diffuse reflection term and specular reflection term according to the realistic shading technology in Unreal Engine 4 combined with the predicted material items. Finally, obtain the initial rendered image img2 of the target object under the viewing angle v.
[0097] It should be noted that in order to decouple the material of the generated object from the environmental lighting so that the predicted material is more accurate, the environmental map in the process of IBL uses a trainable environmental map tensor with a random initialization range of [0.25, 0.75];
[0098] (3.3) Loss calculation and network parameter update. Using the results image img0 generated by the multi-view diffusion model as supervision, calculate SDS_Loss(img0, img2), which is used to train the feature network F for generating the material item part and the environment map tensor. In addition, since there is no prior material information, it is challenging to directly decouple the material of an object from the environmental illumination. The present invention adopts a material smoothness loss and an illumination loss to guide the optimization in order to obtain a well-separated result of the material and the environmental illumination:
[0099] (3.3.1) Calculate the material smoothness loss. Taking the albedo a as an example, the material smoothness loss is defined as:
[0100] ,
[0101] where a(x) represents the albedo parameter at the surface x of the object, is a random three-dimensional displacement vector of x, which follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.01, represents the number of dimensions of the albedo. Adopting the material smoothness loss can prevent local mutations or unreasonable fluctuations from occurring, which helps to ensure the natural and continuous transition of the material on the object surface and makes the generated 3D object more realistic in appearance;
[0102] (3.3.2) Calculate the illumination loss. In order to reduce the separation of high-intensity illumination and sharp shadows onto the albedo and the separation of the base color of the object onto the environmental illumination during the process of decoupling the material and the environmental illumination of the object, the illumination loss designed by the present invention is defined as:
[0103] ,
[0104] where and are the diffuse and specular radiance sampled from the environment map respectively, is a simple luminance operator, which is calculated by taking the average of the three color channels of the input illumination, while is the luminance calculation formula in the hue-saturation-value model (HSV). The first term of the illumination loss approximates the luminance of the illumination with the result image img2 of PBR, and then propagates the gradient to and , which is beneficial to baking the high-intensity illumination and sharp shadows onto the environment map; while the second term penalizes the change in the illumination color, making the environment map tend to be black-and-white light.
[0105] Step 4, Object localization in the scene: Regarding the target object as the foreground and the 3D scene as the background according to the input 3D scene and the rectangular frame where the target object is located, construct a fused image. Determine the depth information of the foreground object through monocular depth estimation and the least squares method, and then calculate the spatial coordinates of the target object in 3D space using back-projection.
[0106] Refer to Figure 5 , the implementation of this step is as follows:
[0107] (4.1) Construct the fused image. First, render the 2D image of the given 3D scene and the scene depth map according to the camera parameters c, and then render the rendered image of the target object obtained in step 3 . Then, regarding the rendered image of the target object as the foreground and the 2D image of the 3D scene as the background, construct the fused image ;
[0108] (4.2) Calculate the depth of the foreground object. Use a common monocular depth estimation method (such as ominidata) to estimate the depth map of the fused image . Since the estimated depth is non-metric, the least squares method can be used to align this depth map with the scene depth map. Through this alignment, the depth of the foreground object can be determined;
[0109] (4.3) Calculate the spatial coordinates of the target object in the scene. Back-project the center point p of the rectangular frame into 3D space as the spatial coordinates of the generated object. The projection formula is:
[0110] ,
[0111] where represents the coordinates of a point in 3D space; is the inverse matrix of the extrinsic camera parameter matrix, where is the rotation matrix, describing the rotation attitude of the camera, is the translation vector, describing the translation position of the camera; represents the depth value of point p in the depth map; is the inverse matrix of the intrinsic camera parameter matrix, used to convert the homogeneous coordinate point on the image plane to the normalized coordinate in the camera coordinate system, represents the homogeneous coordinate of point on the image plane, and are the pixel coordinates of this point on the image plane. The last dimension of 1 is to satisfy the representation form of homogeneous coordinates.
[0112] Step 5, Harmonization Phase: Obtain the scene environment map according to the spatial coordinates of the object to re-illuminate the generated object, and use the estimated uncertainty and the pre-trained 2D image harmonization network to optimize the harmony between the generated object and the scene, ensuring that the generated object naturally integrates into the scene.
[0113] Refer to Figure 6 , the implementation of this step includes the following:
[0114] (5.1) Obtain the scene environment map according to the spatial coordinates of the object. Take the spatial coordinates of the generated object as the position of the camera, take the directions of the six faces of the cube as viewpoints, render six scene pictures to form an environment cube map, and then construct these six pictures into a panoramic picture, which is the scene environment map;
[0115] (5.2) Randomly sample camera parameters, and obtain a new rendered result picture img3 of the target object under the condition of these camera parameters based on physically based rendering (PBR). The process is similar to (3.2). The difference is that in the process of IBL, the scene environment map is used to replace the previous trainable environment map tensor to re-illuminate the target object and integrate the environmental light of the scene into the target object;
[0116] (5.3) Harmonization optimization. According to the estimated uncertainty u and the pre-trained 2D image harmonization network, optimize the harmony between the target object and the scene, calculate the uncertainty loss and the harmonization loss respectively, and train the uncertainty term prediction module and the albedo part of the material term prediction module of the feature network F:
[0117] (5.3.1) Calculate the uncertainty loss, which is used to train the part of the feature network F that estimates uncertainty. Since in the process of text-to-3D, its supervision mechanism is not based on the images rendered from real 3D objects, but relies on the multi-view diffusion model to calculate the SDS_Loss for control. This unique supervision method may cause many unexpected situations, resulting in poor quality and low realism of the generated objects, and also problems such as disharmony in the overall presentation. The present invention introduces uncertainty to transform the original loss function. When the calculated local loss is large due to model instability, the estimated uncertainty will also increase, and the transformed loss will make the impact of this part on the loss smaller. And when calculating the harmonization loss, local harmonization optimization can be performed according to the estimated uncertainty. The uncertainty loss function is defined as:
[0118] ,
[0119] where represents the estimated uncertainty term, represents An exponential function with a base, the first term is to transform the original with uncertainty , the second term is an uncertainty regularization term to prevent predicting infinite uncertainty, which leads to zero loss;
[0120] (5.3.2) Calculate the harmonization loss for training the part of the feature network F to predict albedo. Since the traditional harmonization loss has problems such as low color accuracy and poor generalization ability, the present invention uses a pre-trained 2D image harmonization network (such as PCTNet) to guide the harmonization optimization process. First, input the new rendered image of the target object and the 2D rendered image of the 3D scene under the same camera parameter conditions into the pre-trained 2D image harmonization network, and output the result image img4 of the target object harmonized with the scene. Then, calculate the difference between the new rendered result image img3 and the image img4 as a penalty, and combined with the estimated uncertainty u, the harmony degree between the target object and the scene can be further optimized for local details. The harmonization loss function is defined as:
[0121] ,
[0122] where represents the estimated uncertainty term, represents the mean square error loss function.
[0123] Use the rendered result images img3 and the harmonized result images img4 under different camera parameter conditions for multiple iterative trainings to obtain the finally trained feature network.
[0124] Step 6, obtain the output: The output is a multi-view scene rendering dense view containing the target object and a 3D model file of the target object. Calculate the occlusion relationship through the depth of the target object and the 3D scene in these views, and still be able to maintain harmony when switching different views. Some test experimental results can be referred to Figures 7 - 12 as shown. According to these dense views, various forms of 3D representations can be easily reconstructed. When generating the 3D model file of the target object, first perform parametric UV unwrapping on the deformable tetrahedral mesh, then use the trained feature network to predict the material properties based on the sampling points in the texture space, and then generate seamless textures through the method of adaptive image inpainting, and finally output the 3D model file of the object.
[0125] The above are only the preferred embodiments of the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the scope of the technical solution of the present invention, or modify it into equivalent embodiments with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for generating three-dimensional scenes and harmonized objects based on physical rendering and uncertainty estimation, characterized in that: The following steps are involved: Given a three-dimensional scene, generate a text description of a target object and a rectangular box of the location of the target object in a 2D rendering view of the three-dimensional scene; Based on the generated text description of the target object, a deformable tetrahedral mesh of the geometric appearance of the target object is constructed based on a two-stage method using a pre-trained multi-view image generation diffusion model; Constructing a feature network, wherein the feature network includes a material item prediction module and an uncertainty item prediction module; The input of the feature network is the sampling points of the deformable tetrahedral mesh surface, and the output is the material items and uncertainty items of the object, wherein the material items include albedo, metallicity and roughness; Using physical rendering technology to obtain a rendered image of the target object, and training and optimizing the feature network; Obtaining a multi-view scene rendering dense view and a 3D model file containing a target object according to the three-dimensional scene, the deformable tetrahedral mesh and the material items output by the trained feature network; The method of obtaining a rendered image of a target object by using physical rendering technology specifically includes: Image-based lighting technology calculates the ambient light on the surface of the target object according to the environment map, combines the material item and the ambient light, and obtains the rendered image of the target object through physical rendering technology; Training and optimizing the feature network specifically includes: Randomly initialize the environment map, obtain an initial rendered image of the target object according to the initial environment map and the initial material item output by the feature network, and use the initial rendered image to train the material item prediction module of the feature network through a material item loss function; Acquire the spatial coordinates of the target object in the three-dimensional scene according to the three-dimensional scene, the initial rendering image of the target object, and the rectangular frame of the position of the target object in the 2D rendering view of the three-dimensional scene; Acquire a scene environment map according to the spatial coordinates, and acquire a new rendered image of the target object according to the scene environment map and the material item output by the trained feature network; Use the pre-trained 2D image harmonization network to obtain the resulting image after the target object and the 3D scene are harmonized; The uncertainty term loss and the harmonization loss are calculated according to the new rendered image of the target object and the harmonized result image, and the albedo parts of the uncertainty term prediction module and the material term prediction module of the feature network are trained.
2. The method for generating three-dimensional scenes and harmonized objects based on physical rendering and uncertainty estimation according to claim 1, characterized in that: The three-dimensional scene is a pre-trained three-dimensional scene represented by a neural radiation field NeRF.
3. The method for generating three-dimensional scenes and harmonized objects based on physical rendering and uncertainty estimation according to claim 1, characterized in that: The step of generating a text description of the target object and using a pre-trained multi-view image generation diffusion model to construct a deformable tetrahedral mesh of the geometric appearance of the target object based on a two-stage method is as follows: Construct a 3D target object represented by NeRF; Taking the camera parameters as input, using NeRF's multi-layer perceptron to process the input, outputting a color field and a density field, and obtaining a 2D image img1 of the 3D object represented by NeRF under the camera parameter conditions through volume rendering technology according to the color field and the density field; Taking the text description and camera parameters of the generated target object as input, a pre-trained multi-view image generation diffusion model is used to output a 2D image img0 of the target object that meets the text description under the camera parameter conditions; Using the image img1 and the image img0 to calculate the SDS loss, training the 3D target object represented by the NeRF, and obtaining the 3D target object represented by the NeRF with a rough geometric appearance; Extracting an initial deformable tetrahedral mesh from the 3D target object represented by NeRF having a rough geometric appearance; Obtaining a normal map normal1 of the initial deformable tetrahedral mesh under the camera parameter condition and a normal map normal2 of the 3D target object represented by NeRF with a rough geometric appearance; The SDS loss is calculated using the image img0 and the normal map normal1, the MSE loss is calculated using the normal map normal1 and the normal map normal2, and the initial deformable tetrahedral mesh is trained to obtain a deformable tetrahedral mesh including a refined geometric appearance of the target object.
4. The method for generating three-dimensional scenes and harmonized objects based on physical rendering and uncertainty estimation according to claim 3, characterized in that: The specific steps of extracting the initial deformable tetrahedral mesh from the 3D target object represented by NeRF with a rough geometric appearance are: The density field of the 3D target object represented by NeRF with a rough geometric appearance is converted into a signed distance function SDF, the signed distance function SDF is subtracted from a non-zero constant to obtain an initial SDF, and an initial deformable tetrahedral mesh is obtained according to the initial SDF using a differentiable tetrahedron marching algorithm.
5. The method for generating three-dimensional scenes and harmonized objects based on physical rendering and uncertainty estimation according to claim 1, characterized in that: The feature network includes an input layer, a hash network coding layer, a linear layer, an activation function layer and an output layer.
6. The method for generating three-dimensional scenes and harmonized objects based on physical rendering and uncertainty estimation according to claim 1, characterized in that: Acquiring the spatial coordinates of the target object in the three-dimensional scene according to the three-dimensional scene, the rendered image of the target object, and the rectangular frame of the position of the target object in the 2D rendering view of the three-dimensional scene specifically includes: Obtaining a 2D rendered image and a scene depth map of the three-dimensional scene; According to the rectangular frame of the position of the target object in the 2D rendering view of the 3D scene, a fused image is constructed with the 2D rendering image of the 3D scene as the background and the rendering image of the target object as the foreground; Estimate the depth map of the fused image using a monocular depth estimation method, align the depth map of the fused image with the scene depth map using a least squares method, and determine the depth of the foreground target object; The center point of the rectangular frame is reversely projected into the three-dimensional scene, and the coordinates of the center point are calculated according to the depth of the foreground target object as the spatial coordinates of the target object in the three-dimensional scene.
7. The method for generating three-dimensional scenes and harmonized objects based on physical rendering and uncertainty estimation according to claim 1, characterized in that: The method of using the pre-trained 2D image harmonization network to obtain a result image after the target object and the three-dimensional scene are harmonized includes the following specific steps: The newly rendered image of the target object and the 2D rendered image of the 3D scene under the same camera parameter conditions are input into the pre-trained 2D image harmonization network, and the harmonized result image of the target object is output.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN115100339A
Image harmonization method and device
CN115797536A