Three-dimensional image generation model and three-dimensional image generation method and device
By adopting dual-stream asynchronous joint diffusion technology and feature interaction module in the three-dimensional image generation model, the problem of inconsistency of three-dimensional images at different perspectives is solved, and higher perspective consistency and creativity are achieved.
Patent Information
- Application Number
- CN202311744258.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-20
AI Technical Summary
The three-dimensional images generated based on the diffusion model are inconsistent at different perspectives, making it difficult to maintain the overall consistency of the image.
Two diffusion models and feature interaction modules between the two are used to train the two diffusion models together through training data from different perspectives, and a three-dimensional image generation model is generated based on the trained diffusion model, and the dual-stream asynchronous joint diffusion technology is used to improve the consistency of the image.
Through feature interaction across images, the two diffusion models generate images consistent with each other, significantly improving the perspective consistency of the three-dimensional image generation model and enhancing the creative ability of the model by introducing neural radiation field models.
Smart Images

Figure CN120182464A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision, and particularly to a method and apparatus for generating a three-dimensional image generation model, and a method and apparatus for generating a three-dimensional image. Background Art
[0002] Generative models in machine learning models play an important role in data analysis. A generative model needs to reproduce the original data distribution by generating new instances, while a discriminative model directly generates predictions for the input. For this purpose, a generative model often needs to comprehensively model various features of the data, while a discriminative model only needs to learn task-related information. For example, a discriminative model may ignore task-irrelevant information (such as color) without sacrificing its performance, but we often expect a generative model to successfully generate every detail in an image (such as texture, color, etc.). From this perspective, generative models are more challenging than discriminative models.
[0003] A generative model can be obtained by using diffusion models. Diffusion models are mainly divided into two processes: forward diffusion and reverse diffusion. In the training stage, an image is gradually contaminated with introduced noise for forward diffusion until the image becomes completely random noise. Through the forward diffusion process, a mapping is formed between the training data distribution and the noise distribution. In the testing stage, through reverse diffusion, a series of Markov chains are used to gradually remove the predicted noise at each time step, thereby recovering the image data from Gaussian noise.
[0004] For the generative model obtained based on the diffusion model, the combined image data of each perspective generated by it sometimes does not look like the same object, that is, the problem of inconsistent three-dimensional images occurs. Summary of the Invention
[0005] In an embodiment of the present disclosure, based on two diffusion models and a feature interaction module therebetween, training data of different perspectives are used to jointly train the two diffusion models, and a three-dimensional image generation model is generated based on the trained diffusion models. Among them, cross-image feature interaction can prompt the two diffusion models to generate consistent images with each other. Thus, a three-dimensional image generation model is obtained based on the two-stream asynchronous joint diffusion technology, and the three-dimensional images generated by using the three-dimensional image generation model have good consistency. In addition, the three-dimensional image generation model may further include a neural radiance field model, and the reference image used to guide the diffusion can be generated in a manner using the neural radiance field model, so that the three-dimensional image generation model has a certain creative ability.
[0006] Some embodiments of the present disclosure propose a method for generating a three-dimensional image generation model, including:
[0007] Setting a feature interaction module between a first diffusion model and a second diffusion model;
[0008] Through the feature interaction processing of the feature interaction module, the first diffusion model and the second diffusion model are jointly trained using image training data from different perspectives;
[0009] A first 3D image generation model is obtained by adding a feedback module between the output end and the input end of the trained first diffusion model or second diffusion model, and the feedback module is configured to determine the noise-added image input at the previous time step based on the noise-added image input at the later time step and the predicted noise output.
[0010] In some embodiments, jointly training the first diffusion model and the second diffusion model using image training data from different perspectives includes:
[0011] Input the full text describing the object, the first time step representing the first true noise, the first reference image of the object from the first perspective, and the first noise-added image with the first true noise added into the first diffusion model to obtain the first predicted noise output by the first diffusion model;
[0012] Input the full text describing the object, the second time step representing the second true noise, the second reference image of the object from the second perspective, and the second noise-added image with the second true noise added into the second diffusion model to obtain the second predicted noise output by the second diffusion model;
[0013] Update the parameters of the first diffusion model according to the first predicted noise and the first true noise;
[0014] Update the parameters of the second diffusion model according to the second predicted noise and the second true noise;
[0015] After updating the values of the time step and the perspective, iteratively train the diffusion model, where the time step includes the first time step and the second time step, the perspective includes the first perspective and the second perspective, and the diffusion model includes the first diffusion model and the second diffusion model.
[0016] In some embodiments, the resolution of the object in the first reference image is lower than that in the first noise-added image; the resolution of the object in the second reference image is lower than that in the second noise-added image.
[0017] In some embodiments, it further includes:
[0018] Input the partial text describing the object and the first perspective into the neural radiance field model to obtain the first reference image of the object from the first perspective output by the neural radiance field model;
[0019] Input a partial text describing the object and a second perspective into the neural radiance field model to obtain a second reference image of the object from the second perspective output by the neural radiance field model;
[0020] After iteratively training the diffusion model, a second 3D image generation model is obtained, and the second 3D image generation model includes the neural radiance field model and the first 3D image generation model.
[0021] In some embodiments, the first diffusion model includes a plurality of first diffusion units, the second diffusion model includes a plurality of second diffusion units, and the feature interaction module includes a plurality of feature interaction units. Among them, the outputs of the first diffusion unit and the second diffusion unit at the upper level are fused through the feature interaction unit at the upper level and used as the inputs of the first diffusion unit and the second diffusion unit at the lower level.
[0022] In some embodiments, the outputs of the first diffusion unit and the second diffusion unit at the upper level are respectively weighted and summed using a first weighting coefficient and a second weighting coefficient to obtain a first calculation result, which is used as the input of the first diffusion unit at the lower level; the outputs of the second diffusion unit and the first diffusion unit at the upper level are respectively weighted and summed using a first weighting coefficient and a second weighting coefficient to obtain a second calculation result, which is used as the input of the second diffusion unit at the lower level.
[0023] In some embodiments, each first diffusion unit and each second diffusion unit include a convolution module and an attention module, and the output end of the convolution module is connected to the input end of the attention module.
[0024] Some embodiments of the present disclosure propose a method for generating a 3D image, including:
[0025] Load the first 3D image generation model;
[0026] Input the full text describing the object, any third perspective, the third reference image of the object from the third perspective, and the third noise image into the first 3D image generation model to obtain a third denoised image of the object from the third perspective output by the first 3D image generation model;
[0027] Synthesize a 3D image of the object based on the third denoised images of the object from each third perspective.
[0028] Some embodiments of the present disclosure propose a method for generating a 3D image, including:
[0029] Load the second 3D image generation model, and the second 3D image generation model includes the neural radiance field model and the first 3D image generation model;
[0030] Input partial text of the object and any fourth perspective into the neural radiance field model to obtain a fourth reference image of the object output by the neural radiance field model at the fourth perspective;
[0031] Input the full text of the object description, the fourth perspective, the fourth reference image of the object at the fourth perspective, and the fourth noise image into the first 3D image generation model to obtain a fourth denoised image of the object output by the first 3D image generation model at the fourth perspective;
[0032] Synthesize a 3D image of the object based on the fourth denoised images of the object at each fourth perspective.
[0033] Some embodiments of the present disclosure propose a method for generating a 3D image, including:
[0034] According to the full text of the description of the first object, the full text of the description of the second object, the mixed image of the first object and the second object at any fifth perspective, and the mask used to edit the mixed image, use the diffusion model in the 3D image generation model to predict the mixed noise of the mixed image;
[0035] Fix the diffusion model and optimize the mixed text variable to fit the mixed noise to obtain the optimal mixed text;
[0036] According to the optimal mixed text and any sixth perspective, use the 3D image generation model to obtain an image of the mixed image at any sixth perspective;
[0037] Synthesize an edited 3D image based on the mixed image at the fifth perspective and the images of the mixed image at each sixth perspective,
[0038] wherein, the 3D image generation model is the first 3D image generation model or the second 3D image generation model.
[0039] In some embodiments, predicting the mixed noise of the mixed image includes:
[0040] Based on the mixed image and the full text of the description of the first object, use the diffusion model in the 3D image generation model to predict a third predicted noise;
[0041] Based on the mixed image and the full text of the description of the second object, use the diffusion model in the 3D image generation model to predict a fourth predicted noise;
[0042] Mix the third predicted noise and the fourth predicted noise using the mask to obtain the mixed noise.
[0043] In some embodiments, fixing the diffusion model and optimizing the mixed text variable to fit the mixed noise to obtain the optimal mixed text includes:
[0044] Using the diffusion model, constructing a prediction function for predicting the mixed noise based on the mixed image and the mixed text variable;
[0045] Constructing a loss function for the gap between the mixed noise and the predicted value of the mixed noise output by the prediction function;
[0046] Calculating the value of the mixed text variable when the value of the loss function is minimized as the optimal mixed text.
[0047] In some embodiments, when the partial image of the first object is edited using the partial image of the second object to obtain the mixed image, if the edited part in the image of the first object is less than a preset ratio, the loss function is the sum operation of the first dot product result and the second dot product result, where the first dot product result is obtained by performing a first dot product operation on the difference between the third predicted noise and the prediction function and the difference between 1 and the mask, and the second dot product result is obtained by performing a second dot product operation on the difference between the fourth predicted noise and the prediction function and the product of the mask and the loss weight, and the loss weight is greater than 1.
[0048] Some embodiments of the present disclosure propose a generating device for a three-dimensional image generation model, including: a memory; and a processor coupled to the memory, the processor being configured to execute a generating method of the three-dimensional image generation model based on instructions stored in the memory.
[0049] Some embodiments of the present disclosure propose a three-dimensional image generating device, including: a memory; and a processor coupled to the memory, the processor being configured to execute a generating method of the three-dimensional image based on instructions stored in the memory.
[0050] Some embodiments of the present disclosure propose a generating device for a three-dimensional image generation model, including: a module for executing a generating method of the three-dimensional image generation model.
[0051] Some embodiments of the present disclosure propose a three-dimensional image generating device, including: a module for executing a generating method of the three-dimensional image.
[0052] Some embodiments of the present disclosure propose a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of a generating method of a three-dimensional image generation model or a generating method of a three-dimensional image. Description of the Drawings
[0053] The accompanying drawings required for use in the embodiments or the description of related technologies will be briefly introduced below. The present disclosure can be more clearly understood according to the following detailed description with reference to the drawings.
[0054] Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 and Figure 2 A schematic diagram showing a method for generating a three-dimensional image generation model according to some embodiments of the present disclosure.
[0056] Figure 3 A schematic diagram showing a three-dimensional image generation model according to some embodiments of the present disclosure.
[0057] Figure 4 A schematic diagram showing a method for generating a three-dimensional image according to some embodiments of the present disclosure.
[0058] Figure 5 A schematic diagram showing a method for generating a three-dimensional image according to some embodiments of the present disclosure.
[0059] Figure 6 and Figure 7 A schematic diagram showing a method for generating a three-dimensional image according to some embodiments of the present disclosure.
[0060] Figure 8 A schematic diagram showing the structure of a device for generating a three-dimensional image generation model according to some embodiments of the present disclosure.
[0061] Figure 9 A schematic diagram showing the structure of a three-dimensional image generation device according to some embodiments of the present disclosure.
[0062] Figure 10 A schematic diagram showing the structure of a device for generating a three-dimensional image generation model according to some embodiments of the present disclosure.
[0063] Figure 11 A schematic diagram showing the structure of a three-dimensional image generation device according to some embodiments of the present disclosure. Detailed implementation manners
[0064] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure.
[0065] Unless otherwise specified, the descriptions such as "first", "second", etc. in the present disclosure are used to distinguish different objects and do not represent meanings such as size or time sequence.
[0066] Figure 1 and Figure 2Schematic diagram of a method for generating a three-dimensional image generation model showing some embodiments of the present disclosure.
[0067] As Figure 1 and Figure 2 shown, the method for generating a three-dimensional image generation model of this embodiment includes the following steps. Among them, the three-dimensional image generation model generated according to steps 120 - 140 is called the first three-dimensional image generation model, and the three-dimensional image generation model generated according to steps 110 - 140 is called the second three-dimensional image generation model.
[0068] In step 110, a neural radiance field model is used to generate reference images, for example, including steps 110a and 110b.
[0069] In step 110a, partial text describing an object and a first viewing angle are input into the neural radiance field model to obtain a first reference image of the object at the first viewing angle output by the neural radiance field model.
[0070] In step 110b, partial text describing the object and a second viewing angle are input into the neural radiance field model to obtain a second reference image of the object at the second viewing angle output by the neural radiance field model.
[0071] The terms and specific implementations in steps 110a / b are described below.
[0072] The text describing an object includes full text and partial text. Among them, partial text is a way of speaking relative to full text. Partial text is also called "coarse-grained text", and full text is also called "fine-grained text". Partial text is, for example, a general descriptive text of an object extracted from the full text. For example, the full text describing a certain sedan is, for example, a red A1 brand B1 series C1 model sedan, and the partial text describing the sedan can be a red A1 brand sedan.
[0073] The neural radiance field model (Neural Radiance Fields, NeRF) is a model trained by two-dimensional images from different viewpoints for generating new viewpoint images. It uses volume rendering with an implicit neural scene representation through a multi-layer perceptron (Multi-Layer Perceptron, MLP). MLP is a network model composed of fully connected layers and activation layers.
[0074] The working process of the neural radiance field model is described below. Calculate the three-dimensional position and viewing direction of each camera ray according to the given camera view. The camera ray can be denoted as r(t) = o + td, where o and d are the origin and direction of the ray respectively, and t represents the distance of the ray. Then, r(t) is combined with y cConnect them and use the MLP to predict the density σ(t) and RGB (Red, Green, Blue) color c(t) of a given point in the NeRF implicit representation, i.e., σ(t), c(t) = MLP(r(t), d, y c ) where y c represents the partial text of the object. Then, the reference image (also known as the coarse-grained image) is obtained through volume rendering, and the formula can be expressed as: where t n and t f are the nearest boundary and the farthest boundary of t respectively, x c (r) represents the RGB value of the reference image x rendered through the camera ray r c . exp is the exponential function with the natural constant e as the base, s also represents the distance of the light ray, σ(t) and c(t) represent the density and RGB color of a given point in the NeRF implicit representation respectively, and t represents the distance of the light ray.
[0075] The reference image used to guide the diffusion is generated by the neural radiance field model, enabling the 3D image generation model to have a certain creative ability and generate 3D images with certain differences based on the same input data.
[0076] In step 120, a feature interaction module is set between the first diffusion model and the second diffusion model.
[0077] The first diffusion model includes multiple first diffusion units. The first diffusion model is, for example, a U-Net model. The second diffusion model includes multiple second diffusion units. The second diffusion model is, for example, a U-Net model. Each first diffusion unit and each second diffusion unit include a convolution module and an attention module, and the output end of the convolution module is connected to the input end of the attention module. The convolution module is configured to perform convolution processing on the input, and the attention module is configured to process the input according to the attention mechanism in deep learning.
[0078] A feature interaction module is provided between the first diffusion model and the second diffusion model. The first diffusion model, the second diffusion model, and the feature interaction module are collectively called the two-stream asynchronous joint diffusion module. According to the network structure, if there are multiple first / second diffusion units, the feature interaction module also includes multiple feature interaction units, where, in the arrangement order of the network structure, the diffusion units and the feature interaction module in the same position belong to the same level. Among them, the outputs of the first diffusion unit and the second diffusion unit at the upper level are fused through the feature interaction unit at the upper level and used as the inputs of the first diffusion unit and the second diffusion unit at the lower level.
[0079] An exemplary feature interaction and fusion method is as follows: The outputs of the first diffusion unit and the second diffusion unit at the previous level are respectively weighted and summed using the first weighting coefficient and the second weighting coefficient to obtain a first calculation result, which serves as the input to the first diffusion unit at the next level; The outputs of the second diffusion unit and the first diffusion unit at the previous level are respectively weighted and summed using the first weighting coefficient and the second weighting coefficient to obtain a second calculation result, which serves as the input to the second diffusion unit at the next level. Among them, according to needs, the first weighting coefficient can be set to be greater than the second weighting coefficient.
[0080] The forward diffusion processes of the images of the object in the first perspective and the second perspective can be respectively expressed as:
[0081]
[0082]
[0083] Among them, the superscripts 1 and 2 respectively represent and distinguish the parameters of the diffusion branches where the first diffusion model and the second diffusion model are located. The meanings of the superscripts 1 and 2 will not be separately explained in the subsequent parameter explanations. t is the time step. It can be obtained by cumulative variance calculation, that is represents the variance α at time step i i Cumulative variance is obtained by accumulating from i = 1 to t I is the identity matrix, whose dimension can be set according to the operation needs. x0 represents the original image, and x t represents the image of the original image after adding noise at time step t. represents Gaussian noise, := represents sampling. The first formula above means: In the first diffusion model, given the original image After being mixed with Gaussian noise several times, the sampled image Probability distribution The second formula above means: In the second diffusion model, given the original image After being mixed with Gaussian noise several times, the sampled image Probability distribution
[0084] In step 130, through the feature interaction processing of the feature interaction module, the first diffusion model and the second diffusion model are jointly trained using the image training data from different perspectives.
[0085] In some embodiments, the joint training includes, for example, the following steps.
[0086] In step 130-1a, the full text describing the object, the first time step representing the first true noise, the first reference image of the object from the first perspective, and the first noise-added image with the first true noise added are input into the first diffusion model to obtain the first predicted noise output by the first diffusion model, where the resolution of the object in the first reference image is lower than that in the first noise-added image.
[0087] Among them, the first noise-added image is obtained, for example, by adding the first true noise corresponding to the first time step to the object image described in the full text.
[0088] In step 130-1b, the full text describing the object, the second time step representing the second true noise, the second reference image of the object from the second perspective, and the second noise-added image with the second true noise added are input into the second diffusion model to obtain the second predicted noise output by the second diffusion model, where the resolution of the object in the second reference image is lower than that in the second noise-added image.
[0089] Among them, the second noise-added image is obtained, for example, by adding the second true noise corresponding to the second time step to the object image described in the full text.
[0090] Among them, the first / second reference image input into the first / second diffusion model in step 130-1a / b can be generated by the method of using the neural radiance field model in step 110a / b, or can be directly provided using prior knowledge.
[0091] In step 130-2a, according to the first predicted noise and the first true noise, update the parameters of the first diffusion model. Through iterative update, the first predicted noise gradually approaches the first true noise.
[0092] In step 130-2b, according to the second predicted noise and the second true noise, update the parameters of the second diffusion model. Through iterative update, the second predicted noise gradually approaches the second true noise.
[0093] Training the model ∈ θ To predict the applied noise, its optimization function L is: Among them, represents the first noise-added image obtained by adding noise to the original image of the object from the first perspective based on the t1 time step, represents the first reference image of the object from the first perspective, represents the second noise-added image obtained by adding noise to the original image of the object from the second perspective based on the t2 time step, represents the second reference image of the object from the second perspective, y represents the full text of the object, and the subtraction operation therein represents the gap between the predicted noise and the true noise ∈, denotes the mathematical expectation, and the defined optimization function aims to continuously narrow this gap.
[0094] In step 130-3, after updating the values of the time step and the perspective, the diffusion model is iteratively trained to obtain a three-dimensional image generation model, where the time step includes the first time step and the second time step, the perspective includes the first perspective and the second perspective, and the diffusion model includes the first diffusion model and the second diffusion model.
[0095] In step 140, a three-dimensional image generation model is generated based on the trained diffusion model.
[0096] Figure 3 A schematic diagram showing a three-dimensional image generation model according to some embodiments of the present disclosure, where the three-dimensional image generation model is the first three-dimensional image generation model or the second three-dimensional image generation model.
[0097] The first three-dimensional image generation model is obtained by adding a feedback module between the output end and the input end of the trained first diffusion model or the second diffusion model. The feedback module is configured to determine the noisy image input at the previous time step based on the noisy image input at the subsequent time step and the predicted noise output. Here, the noisy image input at the previous time step is also the image obtained by denoising the subsequent time step by one time step. The noise is iteratively removed step by step through the feedback module to finally obtain a denoised image.
[0098] Assume that adding a feedback module between the output end and the input end of the trained first diffusion model is selected as the first three-dimensional image generation model. Then, the denoising process of one of the time steps for gradually denoising based on the feedback module can be expressed by the formula: where is the noise at the t1 time step predicted by the first diffusion model. For the meaning of its parameters, refer to the description in step 130-2b and will not be elaborated here. Δt≥0 is a hyperparameter of the time step difference used to adjust the difference degree of noise addition between the two diffusion branches. represents the variance at the t1 time step in the first diffusion model. represents the variance at the t1-1 time step in the first diffusion model. represents the noisy image at the t1 time step. represents the noisy image at the t1-1 time step.
[0099] The second three-dimensional image generation model includes the neural radiance field model and the first three-dimensional image generation model cascaded after it.
[0100] In the embodiments of the present disclosure, based on two diffusion models and a feature interaction module therebetween, training data from different perspectives are used to jointly train the two diffusion models, and a three-dimensional image generation model is generated based on the trained diffusion models. Among them, cross-image feature interaction can prompt the two diffusion models to generate consistent images with each other. Thus, the three-dimensional images generated by the three-dimensional image generation model (the first / second three-dimensional image generation model) obtained based on the two-stream asynchronous joint diffusion technology have good consistency. In addition, the second three-dimensional image generation model may further include a neural radiance field model, and the reference image used to guide the diffusion can be generated in the way of the neural radiance field model, so that the second three-dimensional image generation model also has a certain creative ability.
[0101] Figure 4 A schematic diagram showing a method for generating a three-dimensional image according to some embodiments of the present disclosure.
[0102] As Figure 4 shown, the method for generating a three-dimensional image includes: steps 410-430.
[0103] In step 410, load the first three-dimensional image generation model.
[0104] In step 420, input the full text describing the object, any third perspective, the third reference image of the object in the third perspective, and the third noise image into the first three-dimensional image generation model, and obtain the third denoised image of the object in the third perspective output by the first three-dimensional image generation model.
[0105] Among them, the third perspective can be any perspective, including both the perspectives used in model training and the new perspectives not used in model training. The third reference image of the object in the third perspective is also the coarse-grained image of the object in the third perspective. The third noise image of the object in the third perspective is a noisy image of the object in the third perspective, which can be generated, for example, based on the full text describing the object and the third perspective using a certain image generation model, such as the NeRF model, but is not limited to this obtaining method.
[0106] Based on the information input, at each time step, the first three-dimensional image generation model predicts the noise through the diffusion model, and the feedback module feeds back the noisy image input in the previous time step based on the noisy image input in the next time step and the predicted noise output. After several iterations, an image with noise removed is finally obtained.
[0107] In step 430, synthesize the three-dimensional image of the object based on the third denoised images of the object in each third perspective.
[0108] The three-dimensional images generated by the first three-dimensional image generation model obtained based on the two-stream asynchronous joint diffusion technology have good consistency.
[0109] Figure 5 Schematic diagram showing a method for generating a three - dimensional image according to some embodiments of the present disclosure.
[0110] As Figure 5 shown, the method for generating a three - dimensional image includes steps 510 - 540.
[0111] In step 510, a second three - dimensional image generation model is loaded.
[0112] In step 520, a partial text describing the object and any fourth viewing angle are input into the neural radiance field model, and a fourth reference image of the object at the fourth viewing angle output by the neural radiance field model is obtained.
[0113] Among them, the fourth viewing angle can be any viewing angle, including both the viewing angles used in model training and new viewing angles not used in model training. The fourth reference image of the object at the fourth viewing angle is also the coarse - grained image of the object at the fourth viewing angle.
[0114] In step 530, the full text describing the object, the fourth viewing angle, the fourth reference image of the object at the fourth viewing angle, and a fourth noise image are input into the first three - dimensional image generation model, and a fourth denoised image of the object at the fourth viewing angle output by the first three - dimensional image generation model is obtained.
[0115] The fourth noise image of the object at the fourth viewing angle is a noisy image of the object at the fourth viewing angle, which can be generated, for example, based on the full text describing the object and the fourth viewing angle using a certain image generation model, such as the NeRF model, but is not limited to this obtaining method.
[0116] Based on the input information, at each time step, the first three - dimensional image generation model predicts noise through a diffusion model, and the feedback module feeds back the noisy image input at the previous time step based on the noisy image input at the later time step and the predicted noise output, and after several iterations, a noise - removed image is finally obtained.
[0117] In step 540, a three - dimensional image of the object is synthesized based on the fourth denoised images of the object at each fourth viewing angle.
[0118] The three - dimensional image generated by the second three - dimensional image generation model based on the two - stream asynchronous joint diffusion technology has good consistency, and the introduction of the neural radiance field model endows it with certain creative ability.
[0119] Figure 6 And Figure 7 Schematic diagram showing a method for generating a three - dimensional image according to some embodiments of the present disclosure.
[0120] As Figure 6 and Figure 7 shown, the method for generating a three-dimensional image of this embodiment includes steps 610-640.
[0121] In step 610, based on the full text describing the first object, the full text describing the second object, the mixed image of the first object and the second object from any fifth perspective, and the mask used to edit the mixed image, the mixed noise of the mixed image is predicted using the diffusion model in the three-dimensional image generation model. Among them, the three-dimensional image generation model is the first three-dimensional image generation model or the second three-dimensional image generation model.
[0122] Among them, the mixed image is obtained by editing a partial image of the first object in the fifth perspective with a partial image of the second object in the fifth perspective. For example, the full text describing the first object is, for example, "a red sedan of brand A1, series B1, and model C1", and the full text describing the second object is, for example, "a red sports car of brand A2, series B2, and model C2". Use the image of the window part of a red sports car of brand A2, series B2, and model C2 in the fifth perspective to edit and replace the image of the window part of a red sedan of brand A1, series B1, and model C1 in the fifth perspective, so as to obtain a mixed image with the window of a red sports car of brand A2, series B2, and model C2 and other parts of a red sedan of brand A1, series B1, and model C1.
[0123] Among them, Figure 7 in represents the original image without noise corresponding to the full text describing the first object.
[0124] In some embodiments, predicting the mixed noise of the mixed image includes steps 611-613. Steps 611-613 are not shown in the figure.
[0125] In step 611, based on the mixed image and the full text describing the first object, the third predicted noise is predicted using the diffusion model in the three-dimensional image generation model, which can be expressed by the formula: Among them, represents the third predicted noise, ∈ θ represents the diffusion model (the first diffusion model or the second diffusion model) in the three-dimensional image generation model, represents the mixed image, y o represents the full text describing the first object. Here, for the sake of simplicity, only the mixed image and the full text describing the first object required for the diffusion model prediction are shown, and the coarse-grained image and perspective required for the diffusion model prediction are not explicitly shown. These information can be obtained by referring to the foregoing method and will not be elaborated here.
[0126] In step 612, based on the mixed image and the full text describing the second object, using the diffusion model in the three-dimensional image generation model, the fourth predicted noise is obtained, which can be expressed by the formula: Wherein, represents the fourth predicted noise, ∈ θ represents the diffusion model in the three-dimensional image generation model, represents the mixed image, y n represents the full text describing the second object. Similarly, for the sake of simplicity, only the mixed image and the full text describing the second object required for the diffusion model prediction are shown here, and the coarse-grained image and perspective required for the diffusion model prediction are not explicitly shown. These information can be obtained by referring to the foregoing method and will not be elaborated here.
[0127] In step 613, the third predicted noise and the fourth predicted noise are mixed using the mask to obtain the mixed noise, which can be expressed by the formula: Wherein, represents the mixed noise, M represents the mask, and ⊙ represents the Hadamard product (i.e., dot product operation) represents the third predicted noise, represents the fourth predicted noise.
[0128] In step 620, the diffusion model is fixed, and the mixed text variable is optimized to fit the mixed noise to obtain the optimal mixed text.
[0129] In some embodiments, this step specifically includes steps 621-623. Figure 6 Steps 621-623 are not shown.
[0130] In step 621, using the diffusion model, a prediction function for predicting the mixed noise based on the mixed image and the mixed text variable is constructed;
[0131] In step 622, a loss function for the gap between the mixed noise and the predicted value of the mixed noise output by the prediction function is constructed;
[0132] In step 623, the value of the mixed text variable when the value of the loss function is minimized is calculated as the optimal mixed text.
[0133] In some embodiments, the loss function can be expressed by the formula: Wherein, * represents a fixed variable (a fixed variable is a variable whose quantity value can be precisely controlled), represents the mixed noise as a fixed variable, represents the mixed image as a fixed variable mixed text variable y b, a prediction function for predicting mixed noise. This loss function can be used when the edited part in the image of the first object is not less than a preset ratio (i.e., the edited area is relatively large).
[0134] In some other embodiments, the loss function is to perform a summation operation on a first dot product result and a second dot product result. Among them, the first dot product result is obtained by performing a first dot product operation on a first difference (the difference between the third predicted noise and the prediction function) and a second difference (the difference between 1 and the mask), and the second dot product result is obtained by performing a second dot product operation on the product of the difference between the fourth predicted noise and the prediction function and the mask and the loss weight (indicating the proportion of the loss part of the second dot product result in the total loss). The loss weight is greater than 1, for example, a value between (1, 2). The loss function can be expressed by the formula: where, * represents a fixed variable, λ represents the loss weight, and the meanings of other symbols are as described above and will not be elaborated here. This loss function can be used, for example, when the edited part in the image of the first object is less than the preset ratio (i.e., the edited area is relatively small), to avoid the loss being occupied by the image noise of the first object during the optimization process.
[0135] So far, the reversal of mixed noise to view-independent mixed text is achieved.
[0136] According to the foregoing text example, the optimal mixed text is, for example, a vehicle with the windows of a red sports car of brand A2, series B2, and model C2 and other parts of a red sedan of brand A1, series B1, and model C1.
[0137] In step 630, according to the optimal mixed text and any sixth view, using the three-dimensional image generation model, an image of the mixed image at any sixth view is obtained.
[0138] Based on the description of generating an image by the foregoing three-dimensional image generation model, it can be determined that: the three-dimensional image generation model can generate an image of the mixed image at any sixth view based on information such as the input full text (the optimal mixed text in this embodiment), the view (the sixth view in this embodiment), the reference image (the reference image obtained by NeRF based on the optimal mixed partial text (such as the red car) and the sixth view in this embodiment), and the noise image (which can be a noise image obtained by models such as NeRF based on the optimal mixed text and the sixth view) in this embodiment.
[0139] In step 640, an edited three-dimensional image is synthesized based on the mixed image at the fifth view and the images of the mixed image at each sixth view.
[0140] Embodiments of the present disclosure are based on the reversal of mixed noise to view-independent mixed text, and then generate images from various perspectives based on the mixed text, so that the image editing of a certain perspective is consistently extended to the images of other perspectives, so as to still form a three-dimensional image with good consistency after local image editing.
[0141] Figure 8 The structural schematic diagram of the generating device of the three-dimensional image generation model showing some embodiments of the present disclosure.
[0142] As Figure 8 As shown, the generating device 800 of the three-dimensional image generation model of this embodiment includes: a memory 810 and a processor 820 coupled to the memory 810. The processor 820 is configured to execute the generating method of the three-dimensional image generation model in any of the foregoing embodiments based on the instructions stored in the memory 810.
[0143] The device 800 may further include an input / output interface 830, a network interface 840, a storage interface 850, etc. These interfaces 830, 840, 850 and the memory 810 and the processor 820 may be connected through a bus 860, for example.
[0144] Among them, the memory 810 may include a system memory, a fixed non-volatile storage medium, etc. The system memory stores an operating system, application programs, a boot loader, and other programs, for example.
[0145] Among them, the processor 820 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or discrete hardware components such as transistors.
[0146] Among them, the input / output interface 830 provides connection interfaces for input / output devices such as monitors, mice, keyboards, and touchscreens. The network interface 840 provides connection interfaces for various networking devices. The storage interface 850 provides connection interfaces for external storage devices such as SD cards and USB flash drives. The bus 860 can use any bus structure among various bus structures. For example, the bus structure includes, but is not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, and the Peripheral Component Interconnect (PCI) bus.
[0147] Figure 9 Schematic diagram of the structure of a three-dimensional image generation device showing some embodiments of the present disclosure.
[0148] As Figure 9 shown, the three-dimensional image generation device 900 of this embodiment includes: a memory 910 and a processor 920 coupled to the memory 910. The processor 920 is configured to execute the method for generating a three-dimensional image in any of the foregoing embodiments based on instructions stored in the memory 910.
[0149] The device 900 may further include an input / output interface 930, a network interface 940, a storage interface 950, etc. These interfaces 930, 940, 950 and the memory 910 and the processor 920 may be connected through a bus 960, for example.
[0150] Among them, the memory 910 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a Boot Loader, and other programs.
[0151] Among them, the processor 920 may be implemented in the form of a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gates, or discrete hardware components such as transistors.
[0152] Among them, the input / output interface 930 provides connection interfaces for input / output devices such as monitors, mice, keyboards, and touchscreens. The network interface 940 provides connection interfaces for various networking devices. The storage interface 950 provides connection interfaces for external storage devices such as SD cards and USB flash drives. The bus 960 can use any bus structure among a variety of bus structures. For example, the bus structure includes, but is not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, and Peripheral Component Interconnect (PCI) bus.
[0153] Figure 10 Schematic diagram of the structure of a generating device for a three-dimensional image generation model showing some embodiments of the present disclosure.
[0154] As Figure 10 shown, the generating device 1000 for the three-dimensional image generation model of this embodiment includes the following modules.
[0155] The setting module 1010 is configured to set a feature interaction module between the first diffusion model and the second diffusion model.
[0156] The training module 1020 is configured to jointly train the first diffusion model and the second diffusion model by using image training data from different perspectives through the feature interaction processing of the feature interaction module.
[0157] The generating module 1030 is configured to obtain a first three-dimensional image generation model by adding a feedback module between the output end and the input end of the trained first diffusion model or the second diffusion model. The feedback module is configured to determine the noise-added image input at the previous time step based on the noise-added image input at the later time step and the predicted noise output.
[0158] The training module 1020 is configured to input the full text describing the object, the first time step representing the first true noise, the first reference image of the object from the first perspective, and the first noise-added image with the first true noise added into the first diffusion model to obtain the first predicted noise output by the first diffusion model; input the full text describing the object, the second time step representing the second true noise, the second reference image of the object from the second perspective, and the second noise-added image with the second true noise added into the second diffusion model to obtain the second predicted noise output by the second diffusion model; update the parameters of the first diffusion model according to the first predicted noise and the first true noise; update the parameters of the second diffusion model according to the second predicted noise and the second true noise; after updating the values of the time step and the perspective, iteratively train the diffusion model, where the time step includes the first time step and the second time step, the perspective includes the first perspective and the second perspective, and the diffusion model includes the first diffusion model and the second diffusion model. Wherein: the resolution of the object in the first reference image is lower than that in the first noise-added image; the resolution of the object in the second reference image is lower than that in the second noise-added image.
[0159] The reference image module 1040 is configured to input the partial text describing the object and the first perspective into the neural radiance field model to obtain the first reference image of the object from the first perspective output by the neural radiance field model; input the partial text describing the object and the second perspective into the neural radiance field model to obtain the second reference image of the object from the second perspective output by the neural radiance field model;
[0160] The generation module 1030 is configured to obtain a second 3D image generation model after iteratively training the diffusion model, and the second 3D image generation model includes the neural radiance field model and the first 3D image generation model.
[0161] Figure 11 The structural schematic diagram of a 3D image generation device showing some embodiments of the present disclosure.
[0162] As Figure 11 shown, the 3D image generation device 1100 of this embodiment includes modules 1110 - 1120; or / and, the device 1100 includes modules 1120 - 1140.
[0163] In some embodiments, the model loading module 1110 is configured to load the first 3D image generation model; or, load the second 3D image generation model, and the second 3D image generation model includes the neural radiance field model and the first 3D image generation model.
[0164] Correspondingly, the three-dimensional image generation module 1120 is configured to:
[0165] When the first three-dimensional image generation model is loaded, input the full text describing the object, any third perspective, the third reference image of the object in the third perspective, and the third noise image into the first three-dimensional image generation model to obtain the third denoised image of the object in the third perspective output by the first three-dimensional image generation model; synthesize the three-dimensional image of the object based on the third denoised images of the object in each third perspective; or,
[0166] When the second three-dimensional image generation model is loaded, input the partial text describing the object and any fourth perspective into the neural radiance field model to obtain the fourth reference image of the object in the fourth perspective output by the neural radiance field model; input the full text describing the object, the fourth perspective, the fourth reference image of the object in the fourth perspective, and the fourth noise image into the first three-dimensional image generation model to obtain the fourth denoised image of the object in the fourth perspective output by the first three-dimensional image generation model; synthesize the three-dimensional image of the object based on the fourth denoised images of the object in each fourth perspective.
[0167] In some embodiments, the mixed noise module 1130 is configured to: use the diffusion model in the three-dimensional image generation model to predict the mixed noise of the mixed image according to the full text describing the first object, the full text describing the second object, the mixed image of the first object and the second object in any fifth perspective, and the mask used to edit the mixed image.
[0168] The mixed noise module 1130 is configured to: based on the mixed image and the full text describing the first object, use the diffusion model in the three-dimensional image generation model to predict the third predicted noise; based on the mixed image and the full text describing the second object, use the diffusion model in the three-dimensional image generation model to predict the fourth predicted noise; mix the third predicted noise and the fourth predicted noise using the mask to obtain the mixed noise.
[0169] Correspondingly, the mixed text module 1140 is configured to: fix the diffusion model and optimize the mixed text variable to fit the mixed noise to obtain the optimal mixed text.
[0170] The mixed text module 1140 is configured to: use the diffusion model to construct a prediction function for predicting the mixed noise based on the mixed image and the mixed text variable; construct a loss function for the gap between the mixed noise and the predicted value of the mixed noise output by the prediction function; calculate the value of the mixed text variable when the value of the loss function is minimized as the optimal mixed text.
[0171] In the case where the partial image of the first object is edited using the partial image of the second object to obtain the mixed image, if the edited part in the image of the first object is less than a preset ratio, the loss function is the sum of the first dot product result and the second dot product result, where the first dot product result is obtained by performing a first dot product operation on the difference between the third predicted noise and the prediction function and the difference between 1 and the mask, and the second dot product result is obtained by performing a second dot product operation on the difference between the fourth predicted noise and the prediction function and the product of the mask and the loss weight, and the loss weight is greater than 1.
[0172] Correspondingly, the three-dimensional image generation module 1120 is configured to: according to the optimal mixed text and any sixth view angle, use the three-dimensional image generation model to obtain the image of the mixed image at any sixth view angle; and synthesize and edit the three-dimensional image based on the mixed image at the fifth view angle and the images of the mixed image at each sixth view angle.
[0173] Wherein, the three-dimensional image generation model is a first three-dimensional image generation model or a second three-dimensional image generation model.
[0174] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as methods, systems, or computer program products. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more non-transitory computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer program code.
[0175] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0176] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or boxes Figure 1 specified in the boxes or boxes.
[0177] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or boxes Figure 1 specified in the boxes or boxes.
[0178] The above are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for generating a three-dimensional image generation model, comprising: A feature interaction module is set between the first diffusion model and the second diffusion model; Through the feature interaction processing of the feature interaction module, the first diffusion model and the second diffusion model are jointly trained using image training data from different perspectives; A first 3D image generation model is obtained by adding a feedback module between the output end and the input end of the trained first diffusion model or second diffusion model, and the feedback module is configured to determine the noise-added image input at the previous time step based on the noise-added image input at the later time step and the predicted noise output; 2. According to the method described in claim 1, jointly training the first diffusion model and the second diffusion model using image training data from different perspectives includes: The full text describing the object, the first time step representing the first true noise, the first reference image of the object from the first perspective, and the first noise-added image with the first true noise added are input into the first diffusion model to obtain the first predicted noise output by the first diffusion model; The full text describing the object, the second time step representing the second true noise, the second reference image of the object from the second perspective, and the second noise-added image with the second true noise added are input into the second diffusion model to obtain the second predicted noise output by the second diffusion model; Update the parameters of the first diffusion model according to the first predicted noise and the first true noise; Update the parameters of the second diffusion model according to the second predicted noise and the second true noise; After updating the values of the time step and the perspective, the diffusion models are iteratively trained, where the time step includes the first time step and the second time step, the perspective includes the first perspective and the second perspective, and the diffusion models include the first diffusion model and the second diffusion model.
3. According to the method described in claim 2, wherein: The resolution of the first reference image of the object is lower than that of the first noise-added image; The resolution of the second reference image of the object is lower than that of the second noise-added image.
4. According to the method described in claim 2, further comprising: Input the partial text describing the object and the first perspective into the neural radiance field model to obtain the first reference image of the object from the first perspective output by the neural radiance field model; Input the partial text describing the object and the second perspective into the neural radiance field model to obtain the second reference image of the object from the second perspective output by the neural radiance field model; After iteratively training the diffusion models, a second 3D image generation model is obtained, and the second 3D image generation model includes the neural radiance field model and the first 3D image generation model.
5. According to the method described in any one of claims 1-4, The first diffusion model includes a plurality of first diffusion units, The second diffusion model includes a plurality of second diffusion units, The feature interaction module includes a plurality of feature interaction units, Wherein, The outputs of the first diffusion unit and the second diffusion unit at the upper level are fused by the feature interaction unit at the upper level and used as the inputs of the first diffusion unit and the second diffusion unit at the lower level.
6. According to the method described in claim 5, The outputs of the first diffusion unit at the upper level and the second diffusion unit are respectively weighted and summed using the first weighting coefficient and the second weighting coefficient to obtain a first calculation result, which is used as the input of the first diffusion unit at the lower level; The outputs of the second diffusion unit at the upper level and the first diffusion unit are respectively weighted and summed using the first weighting coefficient and the second weighting coefficient to obtain a second calculation result, which is used as the input of the second diffusion unit at the lower level.
7. According to the method described in claim 5, each first diffusion unit and each second diffusion unit include a convolution module and an attention module, and the output end of the convolution module is connected to the input end of the attention module.
8. A method for generating a three-dimensional image, comprising: Load the first 3D image generation model obtained by the method according to any one of claims 1-3, 5-7; Input the full text describing the object, any third perspective, the third reference image of the object from the third perspective, and the third noise image into the first 3D image generation model to obtain the third denoised image of the object from the third perspective output by the first 3D image generation model; Synthesize the 3D image of the object based on the third denoised images of the object from each third perspective.
9. A method for generating a three-dimensional image, comprising: Load the second 3D image generation model obtained by the method according to any one of claims 4-7, where the second 3D image generation model includes the neural radiance field model and the first 3D image generation model; Input the partial text describing the object and any fourth viewing angle into the neural radiance field model to obtain the fourth reference image of the object at the fourth viewing angle output by the neural radiance field model; Input the full text of the object, the fourth viewing angle, the fourth reference image of the object at the fourth viewing angle, and the fourth noise image into the first 3D image generation model to obtain the fourth denoised image of the object at the fourth viewing angle output by the first 3D image generation model; Synthesize the 3D image of the object based on the fourth denoised images of the object at each fourth viewing angle.
10. A method for generating a three-dimensional image, comprising: According to the full text of the first object, the full text of the second object, the mixed image of the first object and the second object at any fifth viewing angle, and the mask used to edit the mixed image, use the diffusion model in the 3D image generation model to predict the mixed noise of the mixed image; Fix the diffusion model and optimize the mixed text variable to fit the mixed noise to obtain the optimal mixed text; According to the optimal mixed text and any sixth viewing angle, use the 3D image generation model to obtain the image of the mixed image at any sixth viewing angle; Synthesize the edited 3D image based on the mixed image at the fifth viewing angle and the images of the mixed image at each sixth viewing angle, where the 3D image generation model is the first 3D image generation model obtained by the method according to any one of claims 1-3, 5-7 or the second 3D image generation model obtained by the method according to any one of claims 4-7.
11. The method according to claim 10, wherein the predicted mixed noise of the mixed image includes: Based on the mixed image and the full text of the first object, use the diffusion model in the 3D image generation model to predict the third predicted noise; Based on the mixed image and the full text of the second object, use the diffusion model in the 3D image generation model to predict the fourth predicted noise; Mix the third predicted noise and the fourth predicted noise using the mask to obtain the mixed noise.
12. The method according to claim 10 or 11, wherein the diffusion model is fixed, and the mixed text variable is optimized to fit the mixed noise to obtain an optimal mixed text, comprising: Use the diffusion model to construct a prediction function for predicting the mixed noise based on the mixed image and the mixed text variable; Construct a loss function for the gap between the mixed noise and the predicted value of the mixed noise output by the prediction function; Calculate the value of the mixed text variable when the value of the loss function is minimized as the optimal mixed text.
13. The method according to claim 12, wherein In the case where the mixed image is obtained by editing a partial image of the first object with a partial image of the second object, if the proportion of the edited part in the image of the first object is less than the preset ratio, the loss function is the sum operation of the first dot product result and the second dot product result, Wherein, the first dot product result is obtained by performing a first dot product operation on the difference between the third predicted noise and the prediction function and the difference between 1 and the mask, and the second dot product result is obtained by performing a second dot product operation on the difference between the fourth predicted noise and the prediction function and the product of the mask and the loss weight, and the loss weight is greater than 1.
14. A generating device for a three-dimensional image generation model, comprising: Memory; And A processor coupled to the memory, the processor being configured to execute the generation method of the three-dimensional image generation model according to any one of claims 1-7 based on instructions stored in the memory.
15. A generating device for a three-dimensional image, comprising: Memory; And A processor coupled to the memory, the processor being configured to execute the generation method of the three-dimensional image according to any one of claims 8-13 based on instructions stored in the memory.
16. A generating device for a three-dimensional image generation model, comprising: A module for executing the generation method of the three-dimensional image generation model according to any one of claims 1-7.
17. A generating device for a three-dimensional image, comprising: A module for executing the generation method of the three-dimensional image according to any one of claims 8-13.
18. A computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the steps of the method according to any one of claims 1-13.