Diffusion model-based image generation dynamic three-dimensional Gaussian scene method

By combining diffusion model and three-dimensional Gaussian splashing, the problems of slow rendering speed, insufficient dynamic controllability and low detail quality of image generation in the prior art are solved, and more efficient and controllable dynamic three-dimensional scene generation is achieved.

CN119941955APending Publication Date: 2025-05-06GUANGDONG BOHUA UHD INNOVATION CENT CO LTD

Patent Information

Application Number
CN202510074767.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing method of generating dynamic three-dimensional scenes in image has problems such as slow rendering speed, insufficient dynamic controllability and low detail quality.

Method used

Combining the diffusion model and 3D Gaussian splatter, the time correlation and detail quality are enhanced through the diffusion model, and leveraging the explicit characteristics of 3D Gaussian splatter and the advantages of efficient and microrenderable, dynamic 3D Gaussian scenes are generated.

Benefits of technology

The speed, controllability and quality of dynamic three-dimensional scenes generated through images are improved, and the problems of slow rendering speed, insufficient dynamic controllability and low detail quality are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941955A_ABST
    Figure CN119941955A_ABST
Patent Text Reader

Abstract

The invention provides a method for generating a dynamic three-dimensional Gaussian scene based on an image of a diffusion model. The method comprises the following steps: S1, generating a static three-dimensional Gaussian scene; s2, generating a dynamic video: generating the dynamic video according to the input image by using a stable video diffusion model; s3, optimizing video quality: processing the dynamic video frame by frame by using a diffusion model, and obtaining an optimized dynamic video image through back propagation; and S4, generating a dynamic three-dimensional Gaussian scene: combining a six-way feature plane (Hexplane) and three-dimensional Gaussian splashing as dynamic scene representation, extracting features from the dynamic scene representation, and introducing a multi-layer perceptron decoder to regress displacement, rotation and zooming information of three-dimensional Gaussian, thereby obtaining the dynamic three-dimensional Gaussian scene. According to the method, the diffusion model and the three-dimensional Gaussian splash are combined for use, the diffusion model is used for strengthening the time correlation and the detail quality, and the problem of low rendering speed is solved based on the three-dimensional Gaussian splash by using the explicit characteristic and the advantages of high efficiency and micro rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of computer vision, and in particular to a method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model. Background Art

[0002] Three-dimensional models are very important in many fields today and have a wide range of applications, such as virtual reality, game development, and video production. Three-dimensional Gaussian splatter (3DGS) [1] was proposed in 2023 and is the most revolutionary innovation in the field of new perspective synthesis in recent years. Compared with the implicit expression of the previous NeRF [2]-based method, 3D Gaussian splatter is an explicit 3D expression method because 3D Gaussian splatter is based on point clouds. Compared with NeRF, which requires querying neural networks to calculate the scene in real time, 3DGS has faster rendering speed and shorter training time. Furthermore, the research on four-dimensional scenes, that is, dynamic three-dimensional scenes, has made significant progress in recent years, including dynamic three-dimensional scene representation methods based on three-dimensional Gaussian splatter, such as four-dimensional Gaussian splatter [3] and deformable three-dimensional Gaussian [4].

[0003] Diffusion models[5] make it possible to generate images from text and videos from images. In the past two years, some methods[6],[7] have combined dynamic 3D scenes with diffusion models to achieve the generation of dynamic 3D scenes using images. However, the main problems of existing methods are slow rendering speed, insufficient dynamic controllability and low detail quality, which need to be solved urgently.

[0004] The difficulty of solving the above problems and defects is: in order to realize the generation of dynamic three-dimensional Gaussian scenes using images, it is necessary to ensure the controllability and stability of the generation, and multiple modules are required to operate on the input images and generated videos, as well as to enhance the details of the generated content.

[0005] The significance of solving the above problems and defects is: it provides a practical idea of ​​using images to generate dynamic three-dimensional Gaussian scenes, solves many problems of the old method, especially reduces rendering time and improves detail quality, which is of great help for the rapid output of digital content. Summary of the invention

[0006] The present invention provides a method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model. The present invention combines the diffusion model with three-dimensional Gaussian splashing, uses the diffusion model to enhance time correlation and detail quality, and based on three-dimensional Gaussian splashing, utilizes its explicit characteristics and the advantages of efficient and differentiable rendering to solve the problem of slow rendering speed.

[0007] The technical solution of the present invention is as follows: The method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model of the present invention comprises the following steps: S1. generating a static three-dimensional Gaussian scene: using a diffusion model and three-dimensional Gaussian splashing to obtain a static three-dimensional Gaussian scene according to an input image; S2. generating a dynamic video: using a stable video diffusion model to generate a dynamic video according to an input image; S3. optimizing video quality: using a diffusion model to process the dynamic video frame by frame, and obtaining an optimized dynamic video image through back propagation; and S4. generating a dynamic three-dimensional Gaussian scene: using a six-way feature plane (HexPlane) in combination with three-dimensional Gaussian splashing as a dynamic scene representation, extracting features therefrom, and introducing a multi-layer perceptron decoder to regress the displacement, rotation and scaling information of the three-dimensional Gaussian to obtain a dynamic three-dimensional Gaussian scene.

[0008] Optionally, in the above-mentioned method for generating dynamic three-dimensional Gaussian scenes from images based on diffusion models, in step S1, a dimensionality increase module consisting of three parts, namely, multi-view generation, background processing and three-dimensional Gaussian splashing, is used to obtain a static three-dimensional Gaussian scene according to the input image.

[0009] Optionally, in the above-mentioned method for generating dynamic three-dimensional Gaussian scenes from images based on the diffusion model, in step S3, a video optimization module is designed to optimize the dynamic video. Optimize dynamic video Each frame is decomposed into a set of images, and then the diffusion model is used Each frame is optimized and used in the image Supervision is performed to ensure that the image content is consistent, and the expression is:

[0010] represents random noise, For optimized dynamic video images; Design a loss function to enforce consistency and smoothness using backpropagation:

[0011] Output as optimized dynamic video images .

[0012] Optionally, in the above-mentioned method for generating a dynamic three-dimensional Gaussian scene from an image based on a diffusion model, in step S4, a six-directional characteristic plane (HexPlane) is used as a dynamic representation of a three-dimensional Gaussian, and the deformation process of the dynamic three-dimensional Gaussian is regarded as four variables: ,in Indicates location, Represents the timestamp. These four variables are combined in pairs to obtain six feature planes. Features are extracted from the six feature planes of HexPlane, and displacement, rotation and scale changes are regressed from the multi-layer perceptron. This deformation process is defined as the deformation module :

[0013] A gradient back propagation is used to control the deformation process of the 3D Gaussian with the optimized dynamic video image as a guide. The loss function is as follows:

[0014] in, represents the rendering function of a dynamic 3D Gaussian, Indicates a reference view.

[0015] According to the technical solution of the present invention, the beneficial effects produced are: The method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model of the present invention aims at solving the problems of slow rendering speed, insufficient dynamic controllability and low detail quality of existing methods for generating dynamic scenes from images. A dynamic three-dimensional scene generation model based on a diffusion model and three-dimensional Gaussian splashing is constructed, which effectively improves the speed, controllability and quality of generating dynamic three-dimensional scenes from images. In the prior art, the rendering speed is slow, even up to ten hours, and the time correlation is not high and the detail quality is poor. The method of the present invention uses a diffusion model to generate static images of multiple perspectives with images as input through a series of operations to render static three-dimensional Gaussian scenes, and uses an image-to-video diffusion model to generate a controllable video, which is used as a guide to deform the static three-dimensional Gaussian scene, so as to realize the generation of dynamic three-dimensional Gaussian scenes using images. The present invention uses a three-dimensional Gaussian splashing model with explicit characteristics to improve the quality of the results and the rendering speed.

[0016] In order to better understand and illustrate the concept, working principle and effect of the present invention, the present invention is described in detail below through specific embodiments in conjunction with the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific implementation of the present invention or the technical solution in the prior art, the drawings required for use in the specific implementation or the description of the prior art are briefly introduced below.

[0018] Figure 1 It is a flow chart of a method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model of the present invention; Figure 2 Step S1 of the method of the present invention is a dimension-raising module Flowchart of Figure 3is a flow chart of step S2 of the method of the present invention for generating a dynamic video; Figure 4 The step S4 of the method of the present invention is a deformation process of a static three-dimensional Gaussian; Figure 5 It is a comparison between the present invention and the existing method. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical method and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific examples. These examples are only illustrative and not limiting of the present invention.

[0020] The method of the present invention provides a method for generating a dynamic three-dimensional Gaussian scene from an image with high speed and quality. A diffusion model is used to generate static images of multiple viewing angles according to an input image to output a static three-dimensional Gaussian scene. A diffusion model from image to video is used to generate a controllable video and the diffusion model is used to optimize the texture quality of each frame, so that the static three-dimensional Gaussian scene is deformed with the generated video as a guide to obtain a dynamic three-dimensional Gaussian scene.

[0021] The principle of the method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model of the present invention is: 1) using a diffusion model and three-dimensional Gaussian splashing to realize the generation of dynamic three-dimensional Gaussian scenes from images; 2) using the diffusion model to generate static images of multiple perspectives to output static three-dimensional Gaussian scenes; 3) using the diffusion model from image to video to generate a controllable video; 4) using the diffusion model to optimize the texture quality of each frame; 5) using the generated video as a guide to deform the static three-dimensional Gaussian scene to realize the generation of dynamic three-dimensional Gaussian scenes from images. The present invention uses the three-dimensional representation of the diffusion model and three-dimensional Gaussian splashing to realize the generation of dynamic three-dimensional Gaussian scenes from images. For the input image, a diffusion model is used to generate static images of multiple perspectives to output a static three-dimensional Gaussian scene, a diffusion model from image to video is used to generate a controllable video, and a diffusion model is used to optimize the texture quality of each frame, and then the static three-dimensional Gaussian scene is deformed with the generated video as a guide to realize the generation of dynamic three-dimensional Gaussian scenes from images. Using dynamic three-dimensional Gaussian as a way to express dynamic three-dimensional scenes improves rendering speed and quality. Many problems of the old methods are solved.

[0022] like Figure 1 As shown, the method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model of the present invention comprises the following steps: S1. Generate a static 3D Gaussian scene: Use a diffusion model and 3D Gaussian splashing to obtain a static 3D Gaussian scene according to an input image.

[0023] In this step, a dimensionality-enhancing module consisting of multi-view generation, background processing, and 3D Gaussian splashing is used. , according to the input image Get a static 3D Gaussian scene .

[0024] S2. Generate dynamic video: Use the stable video diffusion model to generate dynamic video based on the input image.

[0025] In step S2, a stable video diffusion model is used , according to the input image Generate dynamic video .

[0026] S3. Optimize video quality: Use the diffusion model to process dynamic video frame by frame, and obtain the optimized dynamic video image through back propagation.

[0027] In step S3, the diffusion model is used Process dynamic video frame by frame , and obtain the optimized dynamic video image through back propagation .

[0028] S4. Generate a dynamic three-dimensional Gaussian scene: Use a six-way feature plane (HexPlane) combined with a three-dimensional Gaussian splash as a dynamic scene representation to extract features. The present invention introduces a multi-layer perceptron decoder to regress the displacement, rotation and scaling information of the three-dimensional Gaussian to obtain a dynamic three-dimensional Gaussian scene.

[0029] In step S4, a six-way feature plane (HexPlane) is trained to be combined with a three-dimensional Gaussian splash as a dynamic scene representation, from which features are extracted, and the information of displacement, rotation and scaling is regressed from the multi-layer perceptron decoder to obtain a dynamic three-dimensional Gaussian scene. .

[0030] The method of the present invention uses a diffusion model to generate static images of multiple perspectives with images as input to render static three-dimensional Gaussian scenes, and uses an image-to-video diffusion model to generate a controllable video, which is used as a guide to deform the static three-dimensional Gaussian scene, thereby realizing the generation of dynamic three-dimensional Gaussian scenes using images. Using dynamic three-dimensional Gaussian as a way to express dynamic three-dimensional scenes improves rendering speed and quality, thereby solving the problems of slow rendering speed, insufficient dynamic controllability, and low detail quality of dynamic three-dimensional scenes generated from images in existing methods.

[0031] The specific implementation steps of the present invention are as follows: S1. Generate a static 3D Gaussian scene: First, obtain the user's image input, and the input image is recorded as The present invention adopts a dimension-raising module , according to the input image Get a static 3D Gaussian scene , whose expression is:

[0032] Specifically, the dimension-raising module used in the present invention It consists of three parts: multi-view generation, background processing and 3D Gaussian splashing. Figure 2 Using the diffusion model , according to the input image Generate a set of sixteen views of a 2D image sample , the expression of this step is:

[0033] in, is a text prompt, such as "generate multiple views of the object in this image".

[0034] In order to avoid other colors of the background introducing additional noise when rendering, the two-dimensional image samples Do background processing , set the background uniformly to white.

[0035] Then, this set of two-dimensional image samples Input to the 3D Gaussian splatter model In the static three-dimensional Gaussian scene The expression is:

[0036] So far, the output of step S1 is obtained .

[0037] S2. Generate dynamic video: The present invention uses a stable video diffusion model , according to the input image Generate dynamic video , whose expression is:

[0038] here represents random noise, Indicates time, users can choose different random seeds to generate high-quality videos with better time correlation, thus achieving controllability. Its design is as follows Figure 3 shown.

[0039] S3. Optimize video quality: The dynamic video obtained by step S2 There may be problems with low texture quality and inconsistency. To solve this problem, a video optimization module is designed to optimize dynamic videos. Optimize, the optimized dynamic video image expression is:

[0040] Specifically, the dynamic video Each frame is decomposed into a set of images, and then the diffusion model is used Each frame is optimized and used in the image Supervision is performed to ensure that the image content is consistent, and the expression is:

[0041] here represents random noise, Optimized for dynamic video images.

[0042] In order to control the optimization process of this step, a loss function is designed to use backpropagation to enhance consistency and smoothness:

[0043] The output of this step is an optimized dynamic video image .

[0044] S4. Generate dynamic 3D Gaussian scene: We trained a HexPlane combined with a 3D Gaussian splash as a dynamic scene representation, extracted features from it, and regressed the displacement, rotation, and scaling information from the multi-layer perceptron decoder to obtain a dynamic 3D Gaussian scene. .

[0045] Specifically, in order to further transform the static 3D Gaussian scene Deformed into a dynamic 3D Gaussian field, HexPlane is used as the dynamic representation of the 3D Gaussian. More specifically, the deformation process of the dynamic 3D Gaussian is considered as four variables ,in Indicates location, Represents the timestamp. Combining these four variables in pairs can get six feature planes, as follows: Figure 4 Controlling the axis resolution in space and time can regularize the local spatial resistance and temporal coherence of motion, leading to better results.

[0046] Specifically, the present invention extracts features from the six feature planes of HexPlane and regresses displacement, rotation and scaling changes from a multi-layer perceptron. The present invention defines this deformation process as a deformation module :

[0047] The present invention uses a gradient back propagation to control the deformation process of the three-dimensional Gaussian with the optimized dynamic video image as a guide, and the loss function is as follows:

[0048] in, represents the rendering function of a dynamic 3D Gaussian, Indicates a reference view.

[0049] Step S4 is the core invention of the present invention.

[0050] Figure 5 This is an experimental comparison chart of the present invention and the existing image generation method of dynamic three-dimensional scene. It can be seen intuitively that the present invention retains the details and material texture of the input image as much as possible, while the old method loses a lot of detail information and material information during the generation process. And from the new perspective images generated subsequently, the dynamic three-dimensional scene generated by the present invention has a strong time correlation, and the generated new perspective also has a very good quality, while the dynamic three-dimensional scene generated by the old method has a decline in image quality as the perspective deviates from the initial perspective, and the action is less and less close to reality. This reflects that the present invention has better generation quality and time correlation than the old method, and is very usable.

[0051] The following Table 1 compares the result parameters of the present invention and the existing image generation dynamic three-dimensional scene method. The learning-perceived image block similarity is a measurement method for evaluating image quality from the perspective of human perception through learning. It is often used in generative model evaluation, especially in image generation and image reconstruction tasks. The lower the value of the learning-perceived image block similarity in the evaluation of the present invention, the better the quality of the generated image is, and it is visually closer to the real image.

[0052] The contrast language-image pre-training model is mainly used to evaluate the matching degree between the image and the text description, which usually measures the distance between the image and the text embedding space. The present invention uses some text maps to generate dynamic three-dimensional Gaussian scenes for experiments, and randomly samples the perspective from the results to calculate the matching degree with the text description. The present invention obtains higher results on this parameter, which shows that the present invention has better generation quality.

[0053] The Fréchet video distance is an evaluation index for measuring the quality of a video generation model. It measures the distance between the generated video and the real video in the latent space, focusing on the quality and diversity of the generated video. In particular, in the task of evaluating the video generation model, it evaluates the performance of the model by calculating the statistical difference between the generated video and the real video. The low value of the Fréchet video distance of the present invention indicates that the video quality captured by the dynamic three-dimensional Gaussian scene generated by the present invention is higher, that is, the quality of the dynamic three-dimensional Gaussian scene is high.

[0054] Table 1: Comparison results of parameters involved in image generation of dynamic 3D scenes between the present invention and the existing model

[0055] In summary, the method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model of the present invention benefits from the explicit characteristics of three-dimensional Gaussian and the generation ability of the diffusion model, and designs a series of modules to realize the generation of dynamic three-dimensional scenes from images. Specifically, the present invention solves the problem of slow rendering speed of the old method based on the combination of a three-dimensional Gaussian splash model and a six-way feature plane (HexPlane) as an expression method for dynamic three-dimensional scenes; a stable video diffusion model that can convert images into videos is used to obtain a highly controllable video through the image. In order to further improve the temporal relevance of the generated content, the present invention uses a diffusion model to enhance the quality of each frame, solving the problem of low detail quality of the old method. In order to improve the temporal relevance of the generated sequence, the present invention uses a diffusion model to optimize the texture quality of each frame. The present invention can generate a dynamic three-dimensional Gaussian scene using the input image. Compared with the existing method, the present invention reduces the rendering time, increases the dynamic controllability and improves the detail quality. After the image is input, the diffusion model is used to generate images of other perspectives and process the background, so as to generate a static three-dimensional Gaussian scene using the three-dimensional Gaussian splash model, and a stable video is generated according to the input image using the stable video diffusion model, and the static three-dimensional Gaussian scene is deformed with the generated video as a guide.

[0056] The above description is the best embodiment based on the concept and working principle of the invention. The above embodiment should not be understood as limiting the protection scope of the present claims, and other implementations and combinations of implementations according to the concept of the present invention belong to the protection scope of the present invention.

[0057] References [1] B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3DGaussian Splatting for Real-Time Radiance Field Rendering,” Aug. 08, 2023, arXiv : arXiv:2308.04079. Accessed: Oct. 18, 2023. [Online]. Available: http: / / arxiv.org / abs / 2308.04079 [2] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R.Ramamoorthi, and R. Ng, “NeRF: Representing Scenes as Neural Radiance Fieldsfor View Synthesis,” Aug. 03, 2020, arXiv : arXiv:2003.08934. Accessed: Oct.18, 2023. [Online]. Available: http: / / arxiv.org / abs / 2003.08934 [3] G. Wu et al. , “4D Gaussian Splatting for Real-Time Dynamic SceneRendering,” Jul. 15, 2024, arXiv : arXiv:2310.08528. doi: 10.48550 / arXiv.2310.08528. [4] Z. Yang, X. Gao, W. Zhou, S. Jiao, Y. Zhang, and X. Jin,“Deformable 3D Gaussians for High-Fidelity Monocular Dynamic SceneReconstruction,” Nov. 19, 2023, arXiv : arXiv:2309.13101. doi: 10.48550 / arXiv.2309.13101. [5] J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion ProbabilisticModels,” Dec. 16, 2020, arXiv : arXiv:2006.11239. doi: 10.48550 / arXiv.2006.11239. [6] Y. Zheng et al. , “A Unified Approach for Text- and Image-guided4D Scene Generation,” May 07, 2024, arXiv: arXiv:2311.16854. doi: 10.48550 / arXiv.2311.16854. [7] S. Bahmani et al. , “4D-fy: Text-to-4D Generation Using HybridScore Distillation Sampling,” May 26, 2024, arXiv : arXiv:2311.17984. doi:10.48550 / arXiv.2311.17984.

Claims

1. A method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model, characterized in that: The following steps are involved: S1. Generate a static 3D Gaussian scene: use a diffusion model and 3D Gaussian splashing to obtain a static 3D Gaussian scene according to an input image; S2. Generate dynamic video: Use the stable video diffusion model to generate dynamic video based on the input image; S3. Optimizing video quality: processing the dynamic video frame by frame using a diffusion model, and obtaining an optimized dynamic video image through back propagation; and S4. Generate a dynamic three-dimensional Gaussian scene: Use a six-way feature plane (HexPlane) combined with a three-dimensional Gaussian splash as a dynamic scene representation, extract features from it, and introduce a multi-layer perceptron decoder to regress the displacement, rotation and scaling information of the three-dimensional Gaussian to obtain the dynamic three-dimensional Gaussian scene.

2. The method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model according to claim 1, characterized in that: In step S1, a dimensionality increase module consisting of three parts, namely, multi-view generation, background processing and three-dimensional Gaussian splashing, is used to obtain the static three-dimensional Gaussian scene according to the input image.

3. The method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model according to claim 1, characterized in that: In step S3, a video optimization module is designed to optimize the dynamic video Optimize dynamic video Each frame is decomposed into a set of images, and then the diffusion model is used Each frame is optimized and used in the image Supervision is performed to ensure that the image content is consistent, and the expression is: represents random noise, For optimized dynamic video images; Design a loss function to enforce consistency and smoothness using backpropagation: Output as optimized dynamic video images .

4. The method for generating dynamic three-dimensional Gaussian scenes from images based on a diffusion model according to claim 1, characterized in that: In step S4, the six-directional characteristic plane (HexPlane) is used as a dynamic representation of the three-dimensional Gaussian, and the deformation process of the dynamic three-dimensional Gaussian is regarded as four variables ,in Indicates location, Represents the timestamp. These four variables are combined in pairs to obtain six feature planes. Features are extracted from the six feature planes of the six-way feature plane (HexPlane), and the displacement, rotation and scaling changes are regressed from the multi-layer perceptron. This deformation process is defined as the deformation module : A gradient back propagation is used to control the deformation process of the three-dimensional Gaussian with the optimized dynamic video image as a guide, and the loss function is as follows: in, represents the rendering function of a dynamic 3D Gaussian, Indicates a reference view.

Citation Information

Patent Citations

  • Dynamic real-time rendering method for large assembly scene based on three-dimensional Gaussian splashing

    CN119229031A

Cited By

  • Scene reconstruction method, electronic equipment, storage medium and program product

    CN121190638A

  • Scene reconstruction method, electronic device, storage medium, and program product

    CN121190638B

  • Four-dimensional scene generation method and electronic equipment

    CN121190684A