New view angle picture generation method based on diffusion model

By acquiring and identifying camera parameters of multi-view pictures, and using the mapping network to generate conditional embedding vectors, the diffusion model is trained, which solves the problem of high consumption of new view image generation in the existing technology that intensive image acquisition and computing resources is consumed, and fast and efficient new view image generation is achieved.

CN120182408APending Publication Date: 2025-06-20HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510245286.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art relies heavily on intensively acquired images when generating new perspective pictures, and the computing resources for generating new perspective pictures using diffusion models are huge, which cannot meet the needs of quickly generating new perspective pictures.

Method used

The camera collects pictures corresponding to multiple perspectives in a single scene, and identifies the camera parameters of each picture, and generates a test dataset and a training dataset. The mapping network is used to convert camera parameters into conditional embedding vectors, and the diffusion model is performed forward noise addition and reverse diffusion denoising with picture embedding vectors, thereby training the diffusion model. For new perspectives, new conditional embedding vectors are generated through the mapping network, and together with the noise implicit vectors are used to iterative reverse denoising, outputting a full rendered picture.

Benefits of technology

Reduce the dependence of generating new perspective pictures on intensively acquired images, improve the efficiency of computing resources, and realize the rapid generation of high-quality new perspective pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182408A_ABST
    Figure CN120182408A_ABST
Patent Text Reader

Abstract

The invention relates to the field of machine learning, in particular to a new view angle picture generation method based on a diffusion model, which comprises the following steps: acquiring a plurality of pictures of a single scene corresponding to a plurality of view angles through a camera, identifying camera parameters corresponding to each picture, and generating a test data set and a training data set; converting a camera parameter corresponding to the selected training view angle into a conditional embedding vector through a mapping network; generating a picture embedding vector based on the picture of the selected training view angle; based on the condition embedding vector and the picture embedding vector, performing forward noise addition and reverse diffusion denoising on the diffusion model, and training the diffusion model; converting a camera parameter corresponding to the new view angle into a new condition embedding vector through a mapping network; and performing iterative reverse denoising on the diffusion model based on the generated noise implicit vector and the new condition embedding vector, and outputting a complete rendering picture corresponding to the new view angle. According to the method, the demand of quickly generating the new view angle picture can be met by utilizing the least computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and particularly to a method for generating new perspective images based on a diffusion model.

Background Art

[0002] Traditional new perspective image generation methods based on image rendering or physical rendering rely heavily on densely acquired images when inferring scene geometry and appearance. In contrast, pre-trained large models such as diffusion models have significant advantages in generating diverse images using prior knowledge. However, there is still a gap in the field of using diffusion models to generate new perspective images depicting the same scene, and the main difficulty is how to accurately control according to the camera position of the new perspective during the diffusion model generation process.

[0003] In order to achieve precise control of the diffusion model and adapt to diverse task requirements, it is crucial to fine-tune pre-trained large models such as diffusion models and text-to-image models with task-specific data. However, fine-tuning and storing for each task data consume huge computing resources and time, and cannot meet the requirement of generating new perspective images with the least computing resources in the shortest time. It can be seen that how to effectively utilize the prior knowledge of the pre-trained diffusion model, learn the latent vectors related to the global scene content and observation direction from existing multi-perspective images, and guide the diffusion model to generate new perspective images is of great significance for meeting the demand of quickly generating new perspective images with the least computing resources.

Summary of the Invention

[0004] In order to reduce the heavy dependence of generating new perspective images on densely acquired images and meet the demand of quickly generating new perspective images with the least computing resources, the present invention provides a method for generating new perspective images based on a diffusion model, and the method includes the following steps:

[0005] Collect a plurality of images corresponding to multiple perspectives of a single scene through a camera, and label the camera parameters corresponding to each image; generate a test data set and a training data set based on the plurality of images;

[0006] Through a mapping network, convert the camera parameters corresponding to the selected training perspective into conditional embedding vectors; generate image embedding vectors based on the images of the selected training perspective; perform forward noise addition and reverse diffusion denoising on the diffusion model based on the conditional embedding vectors and the image embedding vectors, so as to train the diffusion model;

[0007] Through a mapping network, convert the camera parameters corresponding to the new perspective into new conditional embedding vectors; perform iterative reverse denoising on the trained diffusion model based on the generated noise latent vectors and the new conditional embedding vectors, so as to output a complete rendered image corresponding to the new perspective.

[0008] Preferably, several pictures corresponding to multiple perspectives of a single scene are collected by a camera, and the camera parameters corresponding to each picture are identified. Specifically:

[0009] Several pictures with a unified resolution corresponding to multiple perspectives of a single scene are collected by a camera;

[0010] The internal parameters and external parameters corresponding to the camera at the time of collecting each picture are obtained by using the structure from motion recovery algorithm; wherein, the internal parameters include the focal length of the camera, and the external parameters include the pose of the camera; then each picture is identified based on the internal parameters and the external parameters.

[0011] Preferably, a test data set and a training data set are generated based on several pictures. Specifically:

[0012] All the pictures that have been identified are divided into a test data set and a training data set according to a ratio of 1:8.

[0013] Preferably, the camera parameters corresponding to the selected training perspective are converted into a conditional embedding vector through a mapping network. Specifically:

[0014] A prompt template with several placeholders is input into the multi-modal text and image pre-training model CLIP to obtain an initial word embedding vector;

[0015] Through the mapping network, the camera parameters corresponding to the selected training perspective, the relevant parameters of the diffusion model, and the relevant parameters of the initial word embedding vector are processed to obtain a new word embedding vector; wherein, the relevant parameters of the diffusion model include the current layer index number of the reverse denoising network of the diffusion model and the current denoising step; the relevant parameters of the initial word embedding vector include the current index number of the initial word embedding vector.

[0016] Then, after replacing the initial word embedding vector with the new word embedding vector, iterative processing is performed through the mapping network to obtain a conditional embedding vector.

[0017] Preferably, a picture embedding vector is generated based on the pictures of the selected training perspective. Specifically:

[0018] The ground truth of the pictures of the selected training perspective is input into an image encoder to generate a picture embedding vector.

[0019] Preferably, based on the conditional embedding vector and the picture embedding vector, forward noise addition and reverse diffusion denoising are performed on the diffusion model, thereby training the diffusion model. Specifically:

[0020] Based on the conditional embedding vector, Gaussian noise is applied to the diffusion model for forward diffusion;

[0021] Perform reverse diffusion denoising on the diffusion model that has completed forward noise addition based on the image embedding vector;

[0022] Based on the noise from the reverse diffusion denoising and the Gaussian noise applied in the forward diffusion, determine the diffusion loss of the diffusion model, thereby determining whether the diffusion model has completed training.

[0023] Preferably, the step of applying Gaussian noise to the diffusion model in the forward diffusion based on the conditional embedding vector is specifically:

[0024] Convert the conditional embedding vector into a noise latent vector with a Gaussian distribution, and then based on the noise latent vector, apply Gaussian noise to the diffusion model in the forward diffusion;

[0025] The step of performing reverse diffusion denoising on the diffusion model that has completed forward noise addition based on the image embedding vector is specifically:

[0026] Input the noise latent vector into the reverse denoising network of the diffusion model, and use the image embedding vector to guide the reverse denoising network to perform reverse diffusion denoising;

[0027] The step of determining the diffusion loss of the diffusion model based on the noise from the reverse diffusion denoising and the Gaussian noise applied in the forward diffusion, thereby determining whether the diffusion model has completed training, is specifically:

[0028] Based on the difference between the noise from the reverse diffusion denoising and the Gaussian noise applied in the forward diffusion, determine the diffusion loss function of the diffusion model; then based on the result of the diffusion loss function, determine whether the diffusion model has completed training.

[0029] Preferably, the step of converting the camera parameters corresponding to the new view into a new conditional embedding vector through the mapping network is specifically:

[0030] Based on the new view, determine the camera parameters corresponding to the set virtual camera;

[0031] Input a prompt template with several placeholders into the multi-modal text and image pre-training model CLIP to obtain an initial word embedding vector;

[0032] Through the mapping network, process the camera parameters corresponding to the virtual camera, the relevant parameters of the diffusion model, and the relevant parameters of the initial word embedding vector to obtain a new word embedding vector; wherein, the relevant parameters of the diffusion model include the current layer index number and the current denoising step number of the reverse denoising network of the diffusion model; the relevant parameters of the initial word embedding vector include the current index number of the initial word embedding vector;

[0033] After replacing the initial word embedding vector with the new word embedding vector, iterative processing is performed through the mapping network to obtain a conditional embedding vector.

[0034] Preferably, generating the noise latent vector specifically includes:

[0035] Generating the noise latent vector based on random Gaussian noise.

[0036] Preferably, based on the generated noise latent vector and the new conditional embedding vector, iterative reverse denoising is performed on the trained diffusion model, so as to output a complete rendered image corresponding to the new perspective, specifically:

[0037] Input the generated noise latent vector into the reverse denoising network of the trained diffusion model, and use the new conditional embedding vector to guide the reverse denoising network to perform iterative denoising in the iterative direction of a preset number of denoising steps, so as to output a complete rendered image corresponding to the new perspective.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] Collect a number of pictures corresponding to multiple perspectives of a single scene through a camera, and label the camera parameters corresponding to each picture; based on the number of pictures, generate a test data set and a training data set. First, use the camera to actually shoot and collect a single scene, shoot the scene from multiple different perspective directions to obtain a number of pictures, so that at least one picture is collected in each perspective direction. At the same time, during the shooting process of the camera in different perspective directions, in order to ensure the accuracy of picture shooting, it is necessary to adjust the shooting parameters of the camera, and the shooting parameters of the camera corresponding to different perspective directions are also different. Labeling the camera parameters corresponding to each picture during its acquisition process can differentiate all pictures, which is convenient for subsequent dividing all pictures into a test data set and a training data set, so as to ensure the effective and accurate training and verification of the subsequent diffusion model.

[0040] Through the mapping network, convert the camera parameters corresponding to the selected training perspective into conditional embedding vectors; based on the pictures of the selected training perspective, generate picture embedding vectors; based on the conditional embedding vectors and the picture embedding vectors, perform forward noise addition and reverse diffusion denoising on the diffusion model, so as to train the diffusion model. After selecting a training perspective, encode the pictures corresponding to the training perspective into picture embedding vectors; use the conditional embedding vectors and the picture embedding vectors to perform forward noise addition and reverse diffusion denoising on the diffusion model. By predicting the diffusion loss between the forward noise addition and the reverse diffusion denoising, the training of the diffusion model can be effectively supervised and optimized, and the new perspective picture generation performance and quality of the diffusion model can be ensured.

[0041] Through the mapping network, the camera parameters corresponding to the new perspective are converted into a new conditional embedding vector; based on the generated noise latent vector and the new conditional embedding vector, iterative reverse denoising is performed on the trained diffusion model, so as to output a complete rendered picture corresponding to the new perspective. By setting a virtual camera at the new perspective and obtaining the camera parameters of the virtual camera, and inputting them into the mapping network to generate a new conditional embedding vector, and combining the generated noise latent vector, iterative reverse denoising is performed on the diffusion model. After a predetermined number of denoising steps, a complete rendered picture of the new perspective is finally output, effectively utilizing the multi-perspective picture data and the capabilities of the diffusion model to generate high-quality pictures of the new perspective.

Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0043] Figure 1 It is a flowchart of the method for generating pictures of new perspectives based on the diffusion model provided by the present invention.

[0044] Figure 2 It is a comparison chart of the rendering results of the ablation experiment versions of the method for generating pictures of new perspectives based on the diffusion model provided by the present invention for different numbers of placeholders.

[0045] Figure 3 It is a comparison chart of the rendering results of the ablation experiment versions of the method for generating pictures of new perspectives based on the diffusion model provided by the present invention for different numbers of training pictures.

Detailed Embodiments

[0046] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the drawings. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present invention are shown in the drawings rather than all the structures. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0047] The terms "comprise" and "have" and any variations thereof in the present invention are intended to cover non-exclusive inclusion. For example, a process, method, method, product or device that comprises a series of steps or units is not limited to the listed steps or units, but optionally further comprises steps or units not listed, or optionally further comprises other steps or units inherent to these processes, methods, products or devices.

[0048] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0049] Please refer to Figure 1 As shown, the present invention provides a method for generating new perspective pictures based on a diffusion model, and the method comprises the following steps:

[0050] S100, collecting a plurality of pictures corresponding to multiple perspectives of a single scene through a camera, and identifying the camera parameters corresponding to each picture; generating a test data set and a training data set based on the plurality of pictures.

[0051] Further, collecting a plurality of pictures corresponding to multiple perspectives of a single scene through a camera, and identifying the camera parameters corresponding to each picture, specifically:

[0052] Collecting a plurality of pictures with a unified resolution corresponding to multiple perspectives of a single scene through a camera;

[0053] Using a structure from motion recovery algorithm to obtain the internal parameters and external parameters corresponding to the camera when each picture is collected; wherein, the internal parameters include the focal length of the camera, and the external parameters include the pose of the camera; and then identifying each picture based on the internal parameters and external parameters.

[0054] Further, generating a test data set and a training data set based on the plurality of pictures, specifically:

[0055] Dividing all the pictures that have been identified into a test data set and a training data set according to a ratio of 1:8.

[0056] In actual operation, a digital camera is used to separately capture and collect images from different viewing directions of a single scene, obtaining at least one image corresponding to each viewing direction. At the same time, in order to ensure the unity of the picture features of all the collected images, the shooting resolution of the digital camera is adjusted so that several images corresponding to multiple viewing angles have the same resolution. Then, the Structure from Motion (SfM) algorithm is used to analyze and process each image to determine the internal parameters and external parameters of the digital camera corresponding to each image during shooting and collection. Thus, all the images are differentiated at the camera parameter level, and each image is identified using the internal parameters and external parameters of the camera, establishing the corresponding relationship between each image and its corresponding camera internal parameters and external parameters, realizing the differential and unique differentiation of all the images. Additionally, all the identified images are divided into a test data set and a training data set in a ratio of 1:8, providing sufficient and reliable data support for the subsequent training and verification of the diffusion model.

[0057] S200, through the mapping network, convert the camera parameters corresponding to the selected training viewing angle into conditional embedding vectors; based on the images of the selected training viewing angle, generate image embedding vectors; based on the conditional embedding vectors and image embedding vectors, perform forward noise addition and reverse diffusion denoising on the diffusion model, thereby training the diffusion model.

[0058] Furthermore, through the mapping network, convert the camera parameters corresponding to the selected training viewing angle into conditional embedding vectors, specifically:

[0059] Input the prompt template with several placeholders into the multi-modal text and image pre-training model CLIP to obtain the initial word embedding vectors;

[0060] Through the mapping network, process the camera parameters corresponding to the selected training viewing angle, the relevant parameters of the diffusion model, and the relevant parameters of the initial word embedding vectors to obtain new word embedding vectors; among them, the relevant parameters of the diffusion model include the current layer index number and the current denoising step of the reverse denoising network of the diffusion model; the relevant parameters of the initial word embedding vectors include the current index number of the initial word embedding vectors;

[0061] Then, replace the initial word embedding vectors with the new word embedding vectors and perform iterative processing through the mapping network to obtain conditional embedding vectors.

[0062] Furthermore, based on the images of the selected training viewing angle, generate image embedding vectors, specifically:

[0063] Input the ground truth of the images of the selected training viewing angle into the image encoder to generate image embedding vectors.

[0064] Furthermore, based on the conditional embedding vectors and image embedding vectors, perform forward noise addition and reverse diffusion denoising on the diffusion model, thereby training the diffusion model, specifically:

[0065] Based on the conditional embedding vector, Gaussian noise is applied to the diffusion model for forward diffusion;

[0066] Based on the image embedding vector, the diffusion model after forward noise addition is denoised by reverse diffusion;

[0067] Based on the noise of reverse diffusion denoising and the Gaussian noise applied in forward diffusion, the diffusion loss of the diffusion model is determined, thereby judging whether the diffusion model has completed training.

[0068] Furthermore, based on the conditional embedding vector, applying Gaussian noise to the diffusion model for forward diffusion is specifically as follows:

[0069] The conditional embedding vector is transformed into a noise latent vector with a Gaussian distribution, and then based on the noise latent vector, Gaussian noise is applied to the diffusion model for forward diffusion;

[0070] Based on the image embedding vector, denoising the diffusion model after forward noise addition by reverse diffusion is specifically as follows:

[0071] The noise latent vector is input into the reverse denoising network of the diffusion model, and the image embedding vector is used to guide the reverse denoising network for reverse diffusion denoising;

[0072] Based on the noise of reverse diffusion denoising and the Gaussian noise applied in forward diffusion, determining the diffusion loss of the diffusion model, thereby judging whether the diffusion model has completed training, is specifically as follows:

[0073] Based on the difference between the noise of reverse diffusion denoising and the Gaussian noise applied in forward diffusion, the diffusion loss function of the diffusion model is determined; then based on the result of the diffusion loss function, it is judged whether the diffusion model has completed training.

[0074] In actual operation, a method for generating conditional embedding vectors based on a mapping network is used to generate pictures that affect the diffusion model to generate new perspectives. Specifically, a prompt template with N p placeholder {}s is input into the multi-modal text and image pre-training model CLIP to obtain the initial word embedding vector; the camera parameters corresponding to the selected training perspective, the current layer index number of the reverse denoising network of the diffusion model, the current denoising step, and the current index number of the initial word embedding vector are input into the mapping network together. In this way, the mapping network processes all the inputs, outputs a new word embedding vector, and then replaces the initial word embedding vector with the new word embedding vector and performs iterative processing through the mapping network to obtain the conditional embedding vector, which can ensure that the conditional embedding vector matches the selected training perspective.

[0075] Input the ground truth of the picture of the selected training perspective into the image encoder to generate a picture embedding vector, so that the picture embedding vector is consistent with the picture of the selected training perspective. Also, based on the conditional embedding vector, apply Gaussian noise to the diffusion model for forward diffusion; based on the picture embedding vector, perform reverse diffusion denoising on the diffusion model that has completed forward noise addition; among them, transform the conditional embedding vector into a noise latent vector of a Gaussian distribution, and then based on the noise latent vector, apply Gaussian noise to the diffusion model for forward diffusion, and input the noise latent vector into the reverse denoising network of the diffusion model, and use the picture embedding vector to guide the reverse denoising network to perform reverse diffusion denoising. In this way, by repeatedly performing alternating forward noise addition processing and reverse diffusion denoising processing on the diffusion model, the diffusion model can be fully and accurately trained and verified. Also, based on the difference between the noise of the reverse diffusion denoising and the Gaussian noise applied in the forward diffusion, determine the diffusion loss function of the diffusion model, that is, the diffusion loss function of the diffusion model is equal to the difference between the noise of the reverse diffusion denoising and the Gaussian noise applied in the forward diffusion. When this difference is less than or equal to the preset difference threshold, it is determined that the diffusion model has completed training; otherwise, it is determined that the diffusion model has not completed training. At this time, continue to perform alternating forward noise addition processing and reverse diffusion denoising processing on the diffusion model until this difference is less than or equal to the preset difference threshold.

[0076] S300, through the mapping network, convert the camera parameters corresponding to the new perspective into a new conditional embedding vector; based on the generated noise latent vector and the new conditional embedding vector, perform iterative reverse denoising on the trained diffusion model, so as to output a complete rendered picture corresponding to the new perspective.

[0077] Furthermore, through the mapping network, convert the camera parameters corresponding to the new perspective into a new conditional embedding vector, specifically:

[0078] Based on the new perspective, determine the camera parameters corresponding to the set virtual camera;

[0079] Input the prompt template with several placeholders into the multi-modal text and image pre-training model CLIP to obtain an initial word embedding vector;

[0080] Through the mapping network, process the camera parameters corresponding to the virtual camera, the relevant parameters of the diffusion model, and the relevant parameters of the initial word embedding vector to obtain a new word embedding vector; among them, the relevant parameters of the diffusion model include the current layer index number and the current denoising step of the reverse denoising network of the diffusion model; the relevant parameters of the initial word embedding vector include the current index number of the initial word embedding vector;

[0081] Then, replace the initial word embedding vector with the new word embedding vector and perform iterative processing through the mapping network to obtain a conditional embedding vector.

[0082] Further, a noise latent vector is generated, specifically as follows:

[0083] Based on random Gaussian noise, a noise latent vector is generated.

[0084] Further, based on the generated noise latent vector and the new conditional embedding vector, the trained diffusion model is iteratively denoised in reverse, so as to output a complete rendered image corresponding to the new perspective, specifically as follows:

[0085] The generated noise latent vector is input into the reverse denoising network of the trained diffusion model, and the new conditional embedding vector is used to guide the reverse denoising network to perform iterative denoising in the iterative direction of a preset number of denoising steps, so as to output a complete rendered image corresponding to the new perspective.

[0086] In actual operation, during the process of rendering and generating a new perspective image, based on the new perspective, the camera parameters of the virtual camera for collecting and shooting the scene at the new perspective are determined; also based on random Gaussian noise, a noise latent vector is generated, so that the noise latent vector has Gaussian distribution characteristics.

[0087] The prompt template with several placeholders is input into the multi-modal text and image pre-training model CLIP to obtain an initial word embedding vector. The camera parameters corresponding to the virtual camera, the current layer index number of the reverse denoising network of the diffusion model, the current denoising step number, and the current index number of the initial word embedding vector are input into the mapping network for processing, and a new word embedding vector is output; then the new word embedding vector replaces the initial word embedding vector and is iteratively processed through the mapping network to obtain a conditional embedding vector. Finally, the generated noise latent vector is input into the reverse denoising network of the trained diffusion model, and the new conditional embedding vector is used to guide the reverse denoising network to perform iterative denoising in the iterative direction of a preset number of denoising steps. After a predetermined number of denoising steps, a complete rendered image regarding the new perspective is finally output, effectively utilizing the multi-perspective image data and the capabilities of the diffusion model to generate high-quality new perspective images.

[0088] Please refer to Figures 2 - 3 shown, which are respectively the comparison diagrams of the rendering results of the ablation experiment versions of the new perspective image generation method based on the diffusion model provided by the present invention for different numbers of placeholders, and the comparison diagrams of the rendering results of the ablation experiment versions for different numbers of training images.

[0089] Specifically, the experimental evaluation of the present invention comprehensively evaluates the Forward-facing dataset. This dataset contains 8 real-world forward-facing scenarios, and each scenario contains approximately dozens of images taken from similar perspectives. The rendering quality of the new perspective images is evaluated by metrics such as Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). Among these evaluation metrics, the higher the value of PSNR and / or SSIM, the better the quality, and the lower the value of LPIPS, the better the quality.

[0090] In one experimental instance, different numbers of placeholders are used to evaluate their impact on the quality of new perspective image generation. The quantitative results of the experiment are shown in Table 1 below. As can be seen from Table 1, generally, the quality of new perspective image rendering is improved by increasing the number of placeholders. Correspondingly, the visual rendering effect of the new perspective images is as Figure 2 shown.

[0091] Table 1

[0092] Number of placeholders 5 10 15 20 25 30 PSNR 11.954 12.240 12.390 12.344 12.511 12.511 SSIM 0.296 0.300 0.306 0.301 0.303 0.307 LPIPS 0.588 0.582 0.581 0.578 0.573 0.576

[0093] In another experimental instance, different numbers of training images are used to evaluate their impact on the quality of new perspective image generation. The quantitative results of the experiment are shown in Table 2 below. By increasing the number of training images, the quality of new perspective synthesis is significantly improved, showing improvement in all evaluation metrics. Correspondingly, the visual rendering effect of the new perspective images is as Figure 3 shown.

[0094] Table 2

[0095] Number of training images 2 3 4 5 6 7 8 PSNR 9.99 11.10 12.08 12.13 11.74 12.62 12.53 SSIM 0.18 0.26 0.27 0.28 0.27 0.30 0.29 LPIPS 0.7 0.63 0.62 0.62 0.62 0.6 0.6 Number of training images 9 10 11 12 13 14 15 PSNR 12.39 12.81 13.14 13.53 13.53 13.51 13.09 SSIM 0.30 0.31 0.32 0.35 0.34 0.34 0.33 LPIPS 0.59 0.60 0.56 0.57 0.56 0.56 0.56

[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it. Other embodiments can also be adopted. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A new perspective image generation method based on a diffusion model, characterized in that: The method comprises the following steps: Use a camera to collect several pictures of a single scene corresponding to multiple perspectives, and identify the camera parameters corresponding to each picture; based on several pictures, generate a test data set and a training data set; The camera parameters corresponding to the selected training perspective are converted into a conditional embedding vector through a mapping network; a picture embedding vector is generated based on the picture of the selected training perspective; and a diffusion model is forward denoised and backward diffused denoised based on the conditional embedding vector and the picture embedding vector, thereby training the diffusion model; Through the mapping network, the camera parameters corresponding to the new perspective are converted into a new conditional embedding vector; based on the generated noise latent vector and the new conditional embedding vector, the trained diffusion model is iteratively reverse denoised to output a complete rendered image corresponding to the new perspective.

2. The method according to claim 1, characterized in that: The camera is used to collect a number of pictures corresponding to multiple perspectives of a single scene, and the camera parameters corresponding to each picture are identified, specifically: A camera is used to collect a number of pictures with uniform resolution corresponding to multiple perspectives of a single scene; The motion structure recovery algorithm is used to obtain the intrinsic parameters and extrinsic parameters corresponding to the camera when each picture is captured; wherein the intrinsic parameters include the focal length of the camera, and the extrinsic parameters include the position and posture of the camera; and then each picture is identified based on the intrinsic parameters and the extrinsic parameters.

3. The method according to claim 2, characterized in that The test data set and the training data set are generated based on a number of pictures, specifically: All the marked images are divided into a test data set and a training data set in a ratio of 1:

8.

4. The method according to claim 1, characterized in that: The camera parameters corresponding to the selected training perspective are converted into a conditional embedding vector through the mapping network, specifically: Input the prompt template with several placeholders into the multimodal text and image pre-training model CLIP to obtain the initial word embedding vector; Through the mapping network, the camera parameters corresponding to the selected training perspective, the relevant parameters of the diffusion model and the relevant parameters of the initial word embedding vector are processed to obtain a new word embedding vector; wherein the relevant parameters of the diffusion model include the current layer index number and the current denoising step number of the reverse denoising network of the diffusion model; the relevant parameters of the initial word embedding vector include the current index number of the initial word embedding vector; The new word embedding vector replaces the initial word embedding vector and then iterates through the mapping network to obtain a conditional embedding vector.

5. The method according to claim 4, characterized in that The image embedding vector is generated based on the image of the selected training perspective, specifically: The true value of the image of the selected training view is input into the image encoder to generate the image embedding vector.

6. The method according to claim 5, characterized in that The step of performing forward denoising and reverse diffusion denoising on the diffusion model based on the conditional embedding vector and the image embedding vector, thereby training the diffusion model, is specifically as follows: Based on the conditional embedding vector, performing forward diffusion on the diffusion model and applying Gaussian noise; Based on the image embedding vector, performing reverse diffusion denoising on the diffusion model that has completed forward denoising; Based on the noise of the reverse diffusion denoising and the Gaussian noise applied by the forward diffusion, the diffusion loss of the diffusion model is determined, so as to judge whether the diffusion model has completed training.

7. The method according to claim 6, characterized in that The step of applying Gaussian noise to the diffusion model by forward diffusion based on the conditional embedding vector is specifically as follows: The conditional embedding vector is transformed into a noise latent vector of Gaussian distribution, and then based on the noise latent vector, Gaussian noise is applied to the diffusion model by forward diffusion; Based on the picture embedding vector, reverse diffusion denoising is performed on the diffusion model that has completed the forward denoising, specifically: Inputting the noise latent vector into the reverse denoising network of the diffusion model, and using the image embedding vector to guide the reverse denoising network to perform reverse diffusion denoising; The noise based on the reverse diffusion denoising and the Gaussian noise applied in the forward diffusion are used to determine the diffusion loss of the diffusion model, thereby judging whether the diffusion model has completed training, specifically: Based on the difference between the noise of the reverse diffusion denoising and the Gaussian noise applied by the forward diffusion, a diffusion loss function of the diffusion model is determined; and based on the result of the diffusion loss function, whether the diffusion model has completed training is determined.

8. The method according to claim 1, characterized in that The camera parameters corresponding to the new perspective are converted into a new conditional embedding vector through the mapping network, specifically: Based on the new perspective, determine the camera parameters corresponding to the set virtual camera; Input the prompt template with several placeholders into the multimodal text and image pre-training model CLIP to obtain the initial word embedding vector; Through a mapping network, the camera parameters corresponding to the virtual camera, the relevant parameters of the diffusion model and the relevant parameters of the initial word embedding vector are processed to obtain a new word embedding vector; wherein the relevant parameters of the diffusion model include the current layer index number and the current denoising step number of the reverse denoising network of the diffusion model; the relevant parameters of the initial word embedding vector include the current index number of the initial word embedding vector; The new word embedding vector replaces the initial word embedding vector and then iterates through the mapping network to obtain a conditional embedding vector.

9. The method according to claim 8, characterized in that Generate the noise latent vector, specifically: The noise latent vector is generated based on random Gaussian noise.

10. The method according to claim 9, characterized in that 。 The method performs iterative reverse denoising on the trained diffusion model based on the generated noise latent vector and the new conditional embedding vector, thereby outputting a complete rendered image corresponding to the new perspective, specifically: The generated noise latent vector is input into the reverse denoising network of the trained diffusion model, and the new conditional embedding vector is used to guide the reverse denoising network to perform iterative directional denoising for a preset number of denoising steps, thereby outputting a complete rendered image corresponding to the new perspective.