New view angle image generation method based on depth estimation and diffusion model
By combining the new perspective image generation method of depth estimation and diffusion model, the problems of low quality, insufficient diversity and poor real-time performance in complex scenes and real-world applications are solved, and new perspective image generation with high quality, high diversity and high detail are achieved.
Patent Information
- Application Number
- CN202510369935.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing new perspective image generation technology has problems such as low quality, insufficient diversity and poor real-time performance when dealing with complex scenarios and real-world applications, especially when special equipment is not required and high efficiency is provided.
A new perspective image generation method based on depth estimation and diffusion model is adopted. By generating a training data set and filling the model and depth completion model with pre-trained image content, the monocular depth of the input image is estimated, the grid representation is constructed, and the new perspective image and depth with masks is rendered to fill the mask content.
New perspective image generation with higher quality, higher diversity and higher detail is achieved, able to handle wider perspective transformation, breaks through the limited small perspective transformation limitations of the existing technology, and performs well in complex scenarios and real-world applications.
Smart Images

Figure CN120219633A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and relates to a method for generating novel view images based on depth estimation and diffusion models, specifically to a method for synthesizing novel view images by integrating a depth estimation model and a vision diffusion model. Background Art
[0002] The novel view task, that is, using computer vision technology to generate novel view images from a set of images, is an important topic that has received extensive attention in the fields of computer vision and graphics in recent years. This technology is crucial for understanding and reconstructing 3D scenes, not only being able to display objects from different perspectives but also providing support for various applications such as virtual reality, augmented reality, 3D modeling, and film production. Currently, novel view synthesis mainly relies on methods such as deep learning and 3D reconstruction. Although these methods perform well in dealing with simple scenes or limited view transformations, they usually require a large amount of labeled training data and perform poorly in complex real-world scenes. In recent years, with the rapid development of artificial intelligence-generated content (AIGC), especially the great progress made in text-to-image generation technology, it has provided new opportunities for novel view image generation.
[0003] However, the existing novel view synthesis technologies based on AIGC models still have many defects. For example, previous methods mainly utilized generative adversarial networks (GANs) and achieved remarkable success in synthesizing novel view images on datasets with specifically class-aligned images. At the same time, some methods explored the use of other 3D representations such as neural radiance fields (NeRF) and their derivatives. These methods usually involve training a 3D scene generator, with the support of a co-trained discriminator, for an effective 3D scene representation. However, a common limitation of these methods is calibrating image poses or establishing the prerequisite for pose estimation, which is often impractical in open-world scenarios. To address this issue, some methods introduced a method called score distillation sampling (SDS), which utilized the diffusion model as 2D and then extracted NeRF from text prompts. Subsequently, some methods improved the quality, diversity, and resolution of the output. Despite these advancements, a significant drawback remains the long optimization time required for a single scene, making these methods less feasible for real-time applications. Recent research has extended 3D-aware image synthesis to more diverse and extensive datasets, pursuing an end-to-end training method from single-view images to 3D representations. However, challenges still exist in achieving high resolution and optimal image fidelity, mainly due to the inherent limitations of 3D or multi-view image data with text annotations. In the method, the power of large text-to-image diffusion models was utilized in combination with monocular depth estimation. This combination facilitated the generation of 3D-aware training pairs from single-viewpoint images. Summary of the Invention
[0004] The object of the present invention is to overcome the defects existing in the prior art. In order to obtain high-quality novel-view images, a method for generating novel-view images based on depth estimation and diffusion models is creatively proposed. The present invention can fill the gap in the current novel-view synthesis technology without the need for special equipment and with high efficiency.
[0005] First, the present invention provides a method for generating novel-view images based on depth estimation and diffusion models, including the following steps:
[0006] Generate a training dataset and use it to pre-train an image content filling model and a depth completion model;
[0007] Use a monocular depth estimation model to estimate the monocular depth of the input image, construct its grid representation, and render a novel-view image with a mask and depth; use the pre-trained image content filling model and depth completion model to fill the masked content in the novel-view image with a mask and depth.
[0008] Preferably, the process of generating the training dataset is as follows:
[0009] Construct a key feature text description of a single-view image;
[0010] Input the key feature text description into the diffusion model to synthesize a single-view image;
[0011] Use a monocular depth estimation model to estimate the depth information of the single-view image;
[0012] Utilize the depth information estimated by the monocular depth estimation model, combined with forward-backward warping operations, to construct a novel-view image with a real 3D motion mask as the model label of the single-view image;
[0013] The training dataset is composed of the single-view image and the novel-view image with a real 3D motion mask.
[0014] Preferably, the key features include scene, object, color, and atmosphere.
[0015] Preferably, the diffusion model uses Stable-DiffusionXL.
[0016] Preferably, the monocular depth estimation model uses MiDaS.
[0017] Preferably, the image content filling model uses Stable-Diffusion 2.0, and the depth completion model uses CompleteFormer.
[0018] Preferably, the monocular depth of the input image is estimated using a monocular depth estimation model, and its grid representation is constructed as follows:
[0019] Process the input image I using a monocular depth estimation model i , obtaining the depth map D of the image i ;
[0020] Perform grid processing on the depth map D i to obtain a grid model;
[0021] Map the texture of the original monocular image onto the grid model to effectively convert a single-view image into a grid model with a three-dimensional geometric structure.
[0022] Preferably, render the new view image with a mask and depth; use a pre-trained image content filling model and depth completion model to fill the masked content in the new view image with a mask and depth as follows:
[0023] 1) Use three-dimensional rendering technology to render the grid model from the new view, generating a two-dimensional depth image D' of the three-dimensional scene seen from the new view i+1 ;
[0024] 2) Perform a forward warping operation on the image D' i+1 to obtain an image I' with a mask identifier i+1 ;
[0025] 3) Use the pre-trained image content filling model to fill the masked content of the image I' with a mask identifier i+1 to obtain the image I i+1 ;
[0026] 4) If the maximum number of iterations has been reached, use the current image I i+1 as the final new view image, otherwise perform depth completion on the current image I i+1 to obtain the depth map D i+1 ;
[0027] 5) Re-render the depth map D i+1 to generate a two-dimensional depth image D' of the three-dimensional scene seen from the new view i+2 , then repeat the forward warping operation and masked content filling in steps 2)-3) to obtain the image I i+2 ;
[0028] 6) If the maximum number of iterations has been reached, use the current image I i+2 as the final new view image, otherwise update i = i + 1, perform depth completion on the current image I i+1 to obtain the depth map D i+1, repeat step 5).
[0029] Preferably, the forward warping operation is performed on the image D’ i+1 to obtain the image I’ with a mask identifier i+1 The process is as follows:
[0030] According to the internal and external parameters of the camera, convert the pixel coordinates in the depth information into three-dimensional coordinates;
[0031] Project the three-dimensional coordinates onto the two-dimensional image plane through the camera with a new view to obtain new pixel coordinates;
[0032] Assign the pixel value in the source image to the corresponding position in the target image, and mark that the position has been filled in the mask.
[0033] The second object is to provide a new view image generation system, including:
[0034] A model pre-training module, responsible for generating a training data set and using its pre-trained image content to fill the model and the depth completion model;
[0035] A new view image generation module, responsible for using a monocular depth estimation model to estimate the monocular depth of the input image, constructing its grid representation, and rendering the new view image with a mask and depth; using the pre-trained image content to fill the model and the depth completion model, and filling the mask content in the new view image with a mask and depth.
[0036] Compared with the prior art, the present invention has the following advantages:
[0037] 1. By combining depth estimation and diffusion models, the method of the present invention can estimate depth more accurately, thereby providing higher 3D structure quality and details when generating new view images.
[0038] 2. Compared with the prior art that only focuses on the target object, the method of the present invention considers the background more comprehensively, which makes the generated images more visually realistic and coherent.
[0039] 3. The method of the present invention can handle a wider range of view transformations, breaking through the limitation of the prior art that can only handle limited small view transformations.
[0040] 4. By integrating the latest AIGC technology, the method of the present invention has significantly improved both the quality and diversity of new view image generation.
[0041] Therefore, the method of the present invention is more competitive in the field of new view image generation, especially in dealing with complex scenes and real-world applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of the method of the present invention.
[0043] Figure 2 It is a detailed schematic diagram of the present invention. Specific embodiments
[0044] To better illustrate the purpose and advantages of the present invention, the present invention will be further described below with reference to the accompanying drawings.
[0045] The present invention proposes a text-guided diffusion model training framework dedicated to novel view synthesis for controllable and continuous 3D perception. Two stages are introduced to effectively bridge the gap between 2D and 3D representations. The training data generation stage synergizes the functions of three large pre-trained models - GPT-4 for language modeling, ZoeDepth for monocular depth estimation, and Stable-Diffusion for text-to-image translation - to create a comprehensive dataset containing image warping and content filling training pairs. In terms of data, an expensive multi-view or 3D training data is avoided by developing a data generation pipeline. This innovative pipeline generates diverse and high-quality single-view images. This is achieved by leveraging the synergistic capabilities of large language models and text-to-image diffusion. Depth information is essential and is attached to these generated single-view images through monocular depth estimation. Subsequently, using the estimated depth, multi-views are made by applying the warp-back strategy. The image content filling model, trained on this generated dataset, is used to draw the invisible content in the novel view. This not only ensures the high-fidelity maintenance of pixel values in the novel view corresponding to the visible regions in the source image, but also a novel text-guided content is drawn in the diffusion module to address the drawing of invisible content. In the novel view synthesis stage, to achieve continuous multi-view synthesis, a mesh completion strategy with a depth completion network is developed. This network refines the warped depth in the novel view and merges these depths into a unified mesh, thus playing a crucial role in ensuring the geometric consistency of multi-views. Specifically as follows:
[0046] Step 1: Use the training data generation stage, utilize GPT-4 to generate a large number of text prompts for describing application scenarios; then use the Stable-Diffusion text-to-image model and depth estimation model to generate images corresponding to the text prompts and estimate their depth values; finally, use forward-backward warping to preprocess the synthesized images according to the predicted depth, and train the novel view content filling model and depth estimation model;
[0047] Step 2: After completing the training of the content filling model and depth estimation model, use the depth estimation model to estimate the monocular depth of the input image, construct its mesh representation, and render the novel view with masked images and depths. Then use the trained content filling model and depth completion model to fill the masked content in the novel view images and depths.
[0048] Step 3: Add the filled content in the synthesized new view to the grid representation, and then synthesize other new view images according to Step 2.
[0049] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0050] At least one embodiment provides a method for generating new view images based on depth estimation and diffusion models, as Figure 1-2 shown, including the following steps:
[0051] Step 1: Training data generation.
[0052] Step 1.1: Generate key feature text descriptions of single-view images. The key features include scenes, objects, colors, and atmospheres.
[0053] First, determine the theme or content of the image you want to generate. This can be a specific scene, object, activity, or any element you want the image to represent. Based on the given theme and leveraging the powerful language model capabilities of GPT-4, generate text descriptions. You can provide initial prompts or keywords and let GPT-4 expand them into complete descriptions based on this information. Write a series of detailed text prompts. These texts should clearly describe the key features of the image you want to generate, such as scene settings, objects, colors, atmospheres, etc. The generated text descriptions may need to be fine-tuned to ensure they are detailed enough and meet your requirements. You can edit the generated text to make it more in line with the expected goals.
[0054] Step 1.2: Input the generated text prompts (i.e., key feature text descriptions) into the diffusion model to synthesize single-view images.
[0055] As an example, the diffusion model can adopt Stable-Diffusion XL.
[0056] Input the text prompts generated by GPT-4 into Stable Diffusion to synthesize single-view images. It is necessary to ensure that the text prompts are detailed, specific, and in line with the type of image you want to generate. These prompts should be generated by GPT-4 and optimized. Then start the Stable Diffusion model. This usually needs to be carried out in a supported software environment, such as a machine learning framework with sufficient computing power. In the user interface of Stable Diffusion, input the prepared text prompts. Ensure that the prompts conform to the input format requirements of the model. After submitting the text prompts, Stable Diffusion will start generating images according to the description you provided. This process may take some time, depending on the complexity of the image and the computing resources. The generated images may need further adjustment or optimization to achieve the desired effect. You can adjust the text prompts as needed or use other functions of the model for fine-tuning.
[0057] Step 1.3: Use forward-backward warping operations to synthesize paired data from single-view images for training the model.
[0058] First, use a monocular depth estimation model to estimate the depth of the single-view image. This step is to estimate the depth information of each pixel by analyzing the visual cues in the image.
[0059] As an example, the monocular depth estimation model can adopt MiDaS.
[0060] Next, using the estimated depth information and combining forward-backward warping operations, it is possible to construct image filling training data pairs with real 3D motion masks as the model labels for single-view images. This involves simulating the movement of the camera to generate the scene seen from a new perspective. At the same time, depth completion training data pairs can also be created to repair the loss or inaccuracy of depth information caused by the warping operation.
[0061] The training dataset includes single-view images, new perspective images with real 3D motion masks, and masks.
[0062] Finally, use these generated training data pairs to train the image content filling model and the depth completion model. These models will learn how to effectively fill in the blank areas caused by perspective changes or inaccurate depth estimation, thus generating more complete and realistic new perspective images.
[0063] As an example, the image content filling model can adopt Stable-Diffusion 2.0, and the depth completion model can adopt CompleteFormer.
[0064] Step 2: New perspective synthesis.
[0065] Estimate the monocular depth of the input image using a monocular depth estimation model, construct its grid representation, and render the image and depth with a mask for the new view; use the image content filling model and depth completion model pre-trained in step 1 to fill the masked content in the image and depth with the mask for the new view, and perform new view synthesis. Specifically:
[0066] Step 2.1: Construct the grid representation of the monocular image.
[0067] 2.1.1 Initialize the iteration count i = 0, and initialize the original monocular image P as the input image I i 。
[0068] 2.1.1 Process the input image I using the monocular depth estimation model i to obtain the depth map D of the image i 。This model estimates the depth value of each pixel point based on the visual content of the image. These depth values can be represented as a depth map.
[0069] 2.1.2 Perform meshing on the depth map:
[0070] According to the depth map D i convert each pixel point of the image into a point in three-dimensional space, thus forming a point cloud. The position of each point is determined by its pixel position and depth value.
[0071] For the point cloud data, connect adjacent points to form the faces of the grid. Sometimes additional processing may be required, such as denoising, smoothing, or filling missing regions, to generate a coherent and complete three-dimensional grid and obtain the grid model.
[0072] 2.1.3 Map the texture of the original monocular image onto the grid model to effectively convert a single-view image into a grid model with a three-dimensional geometric structure. This requires adjusting the texture coordinates of the original image to match the structure of the three-dimensional grid. Through this process, a single-view image can be effectively converted into a mesh with a three-dimensional geometric structure, providing a basis for subsequent view transformation and three-dimensional rendering. Specifically, it is mainly divided into the following steps:
[0073] First, perform UV mapping to assign UV coordinates to each vertex of the grid model. UV coordinates are coordinates on a two-dimensional texture image. Then perform texture coordinate calculation. According to the internal and external parameters of the camera, project the vertices of the three-dimensional grid model onto the two-dimensional image and calculate the coordinates of each vertex on the image. Finally, perform texture sampling. According to the calculated UV coordinates, sample the texture information from the two-dimensional image and apply it to the grid model.
[0074] Step 2.2: Render and fill the new view image.
[0075] First, determine the new perspective position and direction. This involves setting the position and orientation of a virtual camera. Then, use 3D rendering techniques such as ray tracing or rasterization to render the mesh from the new perspective. This will generate a 2D image of the 3D scene as seen from that perspective. Since the original image provides only a limited perspective, some regions may be invisible from the new perspective. A mask can be generated to identify these regions. The mask can be calculated based on depth information and the perspective transformation, marking the parts that are occluded or do not exist in the new perspective. Apply the mask to the rendered image to occlude the invisible regions. In this way, a new perspective image that is visually more realistic and consistent can be obtained.
[0076] Through this process, it is possible to generate a new perspective image with a certain degree of realism and integrity using the mesh constructed from a single image, while effectively processing and hiding the invisible regions. Then, use the image content filling and depth completion model trained in the first step to fill the masked regions of the image and depth in the new perspective, and the new perspective image and its depth can be obtained.
[0077] Specifically:
[0078] Step 2.2.1: Use 3D rendering technology to render the mesh model from the new perspective, generating a 2D depth image D’ of the 3D scene as seen from the new perspective i+1 .
[0079] Perform a forward warping operation on the image D’ i+1 to obtain an image I’ with mask identification; specifically, according to the internal and external parameters of the camera, convert the pixel coordinates in the depth image to 3D coordinates. Then project the 3D coordinates onto the 2D image plane through the camera of the new perspective to obtain new pixel coordinates. Assign the pixel values in the source image to the corresponding positions in the target image, and mark that position as filled in the mask. i+1 ;
[0080] Then, use the pre-trained image content filling model to fill the masked content of the image I’ with mask identification i+1 to obtain the image I i+1 .
[0081] Step 2.2.2: Determine whether the current iteration count has reached the maximum iteration count. If so, use the current image I i+1 as the final new perspective image; if not, execute Step 2.2.3.
[0082] Step 2.2.3: Input the current image I i+1 and the depth image D’ i+1 into the depth completion model to obtain the depth map D i+1, and then for the depth map D i+1 Perform meshing processing, then map the texture of the original monocular image to the mesh model, and finally use 3D rendering technology to render the mesh model from a new perspective to generate a 2D depth image D’ of the 3D scene seen from the new perspective i+2 Perform a forward warping operation on the image D’ i+2 to obtain an image I’ with a mask identifier i+2 ; for the image I’ with a mask identifier i+2 Use the pre-trained image content filling model to fill the masked content to obtain the image I i+2 ; determine whether the current iteration number has reached the maximum iteration number. If so, use the current image I i+2 as the final new perspective image; if not, update i = i + 1 and repeat step 2.2.3
[0083] This embodiment also provides a new perspective image generation system, including:
[0084] A model pre-training module, responsible for generating a training data set and using it to pre-train the image content filling model and the depth completion model
[0085] A new perspective image generation module, responsible for using a monocular depth estimation model to estimate the monocular depth of the input image, constructing its mesh representation, and rendering the image and depth with a mask from the new perspective; using the pre-trained image content filling model and the depth completion model to fill the masked content in the image and depth with a mask from the new perspective
[0086] The above specific description further details the purpose, technical solution and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention
Claims
1. A method for generating new perspective images based on depth estimation and diffusion model, characterized in that The following steps are involved: Generate a training dataset and use its pre-trained image content to fill the model and the depth completion model; The monocular depth estimation model is used to estimate the monocular depth of the input image, construct its grid representation, and render the masked image and depth of the new perspective. The pre-trained image content filling model and depth completion model are used to fill the mask content in the masked image and depth of the new perspective.
2. The method according to claim 1, characterized in that: The process of generating the training data set is as follows: Construct a text description of the key features of a single-view image; Input the key feature text description into the diffusion model to synthesize the single-view image; Use a monocular depth estimation model to estimate the depth information of a single-view image; Using the depth information estimated by the monocular depth estimation model and combining it with the forward-backward warping operation, a new view image with a true 3D motion mask is constructed as the model label of the monocular image. The training dataset consists of single-view images and new-view images with true 3D motion masks.
3. The method according to claim 2, characterized in that: The key features include scene, object, color, atmosphere.
4. The method according to claim 2, characterized in that: The diffusion model adopts Stable-Diffusion XL.
5. The method according to claim 1 or 2, characterized in that: The monocular depth estimation model adopts MiDaS.
6. The method according to claim 1, characterized in that: The image content filling model adopts Stable-Diffusion 2.0, and the depth completion model adopts CompleteFormer.
7. The method according to claim 1, characterized in that: The monocular depth estimation model is used to estimate the monocular depth of the input image, and its grid representation is constructed as follows: Use the monocular depth estimation model to process the input image I i , get the depth map D of the image i ; For the depth map D i Performing mesh processing to obtain a mesh model; The texture of the original monocular image is mapped onto the mesh model, which effectively converts a monocular image into a mesh model with a three-dimensional geometric structure.
8. The method according to claim 1, characterized in that: The process of rendering the masked image and depth of the new perspective; and filling the masked content in the masked image and depth of the new perspective using the pre-trained image content filling model and the depth completion model is as follows: 1) Use 3D rendering technology to render the mesh model from a new perspective and generate a 2D depth image D' of the 3D scene seen from the new perspective i+1 ; 2) For image D' i+1 Perform a forward warping operation to obtain an image I' with a mask mark i+1 ; 3) For the image I' with mask mark i+1 Use the pre-trained image content filling model to fill the mask content and get image I i+1 ; 4) If the maximum number of iterations has been reached, the current image I i+1 As the final new perspective image, on the contrary, the current image I i+1 Perform depth completion to obtain the depth map D i+1 ; 5) Depth map D i+1 Re-render to generate a 2D depth image D' of the 3D scene seen from a new perspective i+2 Then repeat the forward warping operation and mask content filling of steps 2)-3) to obtain image I i+2 ; 6) If the maximum number of iterations has been reached, the current image I i+2 As the final new perspective image, otherwise update i = i + 1, and for the current image I i+1 Perform depth completion to obtain the depth map D i+1 , repeat step 5).
9. The method according to claim 8, characterized in that: The pair of images D' i+1 Perform a forward warping operation to obtain an image I' with a mask mark i+1 The process is as follows: According to the intrinsic and extrinsic parameters of the camera, the pixel coordinates in the depth information are converted into three-dimensional coordinates; Project the three-dimensional coordinates onto the two-dimensional image plane through the camera of the new perspective to obtain new pixel coordinates; Assign the pixel value in the source image to the corresponding position in the target image and mark the position as filled in the mask.
10. A new perspective image generation system implementing the method according to any one of claims 1 to 9, characterized in that include: The model pre-training module is responsible for generating training data sets and using their pre-trained image content to fill the model and the depth completion model; The new-perspective image generation module is responsible for estimating the monocular depth of the input image using the monocular depth estimation model, constructing its grid representation, and rendering the masked image and depth of the new perspective; and using the pre-trained image content filling model and depth completion model to fill in the mask content in the masked image and depth of the new perspective.