Method for training and using diffusion model to edit images

WO2026164411A1PCT designated stage Publication Date: 2026-08-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-13
Publication Date
2026-08-06

Smart Images

  • Figure KR2026000774_06082026_PF_FP_ABST
    Figure KR2026000774_06082026_PF_FP_ABST
Patent Text Reader

Abstract

In an embodiment of the disclosure, a method for training a diffusion model to edit images. The method may include obtaining a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt, and a target output image. The method may include generating, for each triplet of training data, an initial noise image for the diffusion model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding. The method may include training the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by denoising, for each triplet, the generated initial noise image to generate an edited image, using a denoising network of the diffusion model; and updating one or more parameters of the diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR TRAINING AND USING DIFFUSION MODEL TO EDIT IMAGES

[0001] The present application generally relates to a method for training and using a diffusion model to edit images. In particular, the present techniques provide a method to train a diffusion machine learning, ML, model to edit images based on user prompts, and a method to use the trained diffusion ML to generate edited images.

[0002] Diffusion models have made immense progress in image generation, showing impressive results for conditional generation. Currently, all state-of-the-art (SOTA) conditional diffusion systems leverage Classifier Free Guidance (CFG) to yield diverse high-quality image generation, at the cost of the increased number of denoising passes of learned neural network models. Specifically, the instruction-based image editing task is conditioned on a given text and image and an instruction prompt. Applying CFG for image editing requires three denoising passes to yield plausible results.

[0003] Informally, the formulation of CFG is based on two premises. First, there exists conditional information that is not natively supported by the underlying generative modelling. Second, the score, i.e., the gradient of the log of the density of conditional generative modelling may be aided by the score of unconditional generative modelling. More formally, CFG can be seen as predictor-corrector. CFG boosts conditional image generation significantly across diffusion and flow models alike. This has led to a surge of research on different ways to combine not just conditional and unconditional models but more generally two or more models.

[0004] The main practical drawback of CFG lies in the need for multiple denoising passes, resulting in additional computation costs. To alleviate the computational cost, a student model is distilled to mimic in one pass the original teacher sampling mechanism requiring multiple passes. CFG distillation, however, imposes a cumbersome extra training stage and is prone to lag behind the teacher's performance.

[0005] The present applicant has therefore identified the need for an improved method for training diffusion models.

[0006] In an embodiment of the disclosure, a method for training a diffusion model to edit images. The method may include obtaining a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt to edit the input image, and a target output image that is an edited version of the input image according to the text prompt. The method may include generating, for each triplet of training data, (i) an input image embedding for the input image, using an image encoder of the diffusion model, (ii) a prompt embedding for the text prompt, using a text encoder of the diffusion model, and (iii) an initial noise image for the diffusion model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding. The method may include training the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by denoising, for each triplet, the generated initial noise image to generate an edited image, using a denoising network of the diffusion model; and updating one or more parameters of the diffusion model to improve alignment between the generated edited image and the output image of the triplet.

[0007] In an embodiment of the disclosure, a server for training a diffusion model to edit images. The server comprises at least one processor coupled to memory. The at least one processor may be coupled to memory to obtain a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt to edit the input image, and a target output image that is an edited version of the input image according to the text prompt. The at least one processor may be coupled to memory to generate, for each triplet of training data, (i) an input image embedding for the input image, using an image encoder of the diffusion model, (ii) a prompt embedding for the text prompt, using a text encoder of the diffusion model, and (iii) an initial noise image for the diffusion model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding. The at least one processor may be coupled to memory to train the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by denoising, for each triplet, the generated initial noise image to generate an edited image, using a denoising network of the diffusion model; and updating one or more parameters of the diffusion model to improve alignment between the generated edited image and the output image of the triplet.

[0008] In an embodiment, a method for using a trained diffusion model to edit images. The method includes obtaining an input image and a text prompt to edit the input image. The method includes generating (i) an input image embedding for the input image using an image encoder of the trained diffusion model, (ii) a prompt embedding for the text prompt, using a text encoder of the diffusion model, and (iii) an initial noise image for the diffusion model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding. The method may include denoising, using a denoising network of the trained diffusion model, the generated initial noise image to generate an edited image.

[0009] Implementations of the present techniques will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0010] Figures 1A is a flowchart showing the methods of the present techniques at training time;

[0011] Figure 1B is a flowchart showing the methods of the present techniques at inference time;

[0012] Figure 2 shows images generated using implicit conditioning and the explicit conditioning of the present techniques;

[0013] Figure 3 shows a schematic diagram of the present techniques during training;

[0014] Figure 4 shows a schematic diagram of the present techniques during inference;

[0015] Figure 5 shows quantitative results of the present techniques;

[0016] Figure 6 shows images generated using the explicit conditioning of the present techniques; and

[0017] Figure 7 is a block diagram of a system for training a diffusion model to performed image editing and for using the trained diffusion model.

[0018] Furthermore, each element illustrated in the drawings may be exaggerated, omitted, or schematically illustrated for convenience of explanation, and clarity. The illustrated size of each element does not substantially reflect its actual size. In each drawing, the same or corresponding elements are assigned with the same reference numerals.

[0019] The advantages and characteristics of the disclosure, and a method to achieve the same, will be clarified by referring to embodiments described below in detail with the accompanying drawings. However, the disclosure is not limited to embodiments described below, but may be implemented in various other forms. The disclosed embodiments are provided to perform the disclosure and to impart a person skilled in the art to which the disclosure belongs of a full scope of the disclosure. An embodiment of the disclosure may be defined according to the scope of claims. Throughout the disclosure, like reference numerals denote like elements. Furthermore, in the description of the disclosure, certain detailed explanations of related art are omitted when it is deemed that they may unnecessarily obscure the essence of the disclosure. The terms used in the disclosure have been selected from currently widely used general terms in consideration of the functions in the disclosure. However, the terms may vary according to the intention of a user or operator, case precedents, and the like. Thus, the definitions are determined based on the overall content of the disclosure.

[0020] In an embodiment of the disclosure, it will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable instruction execution apparatus, create a mechanism for implementing the functions / acts specified in the flowchart and / or block diagram block(s). These computer program instructions may also be stored in a computer readable medium that when executed can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions when stored in the computer readable medium produce an article of manufacture including instructions which when executed, cause a computer to implement the function / act specified in the flowchart and / or block diagram block(s). The computer program instructions may also be loaded onto a computer, other programmable instruction execution apparatus.

[0021] Furthermore, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.

[0022] In a first approach of the present techniques, there is provided a computer-implemented method for training a diffusion machine learning, ML, model to edit images, the method comprising: obtaining a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt to edit the input image, and a target output image that is an edited version of the input image according to the text prompt; generating, for each triplet of training data: an input image embedding for the input image, using an image encoder of the diffusion ML model, a prompt embedding for the text prompt, using a text encoder of the diffusion ML model, and an initial noise image for the diffusion ML model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding; and training the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by: denoising, for each triplet, the generated initial noise image to thereby generate an edited image using the diffusion ML model; and updating parameters of the diffusion ML model to improve alignment between the generated edited image and the output image of the triplet.

[0023] Advantageously, the present techniques provide a diffusion model for editing images according to a user's instructions, which is also less computationally intensive than existing techniques for doing the same. Not only are the resulting edited images produced in less time (because of the reduced number of computations), but the present techniques also result in high-quality images. The present techniques effectively alter the starting point for the denoising performed by the diffusion model, so that the starting point already contains some information about the required edit. This guides or 'conditions' the diffusion model to generate an edited image that corresponds to the user's instructions.

[0024] Diffusion model is a type of generative model, and consists of three main components: a forward process, a reverse process, and a sampling procedure. During the forward process (also known as forward diffusion, or simply, diffusion process), the diffusion model applies a sequence of transformations to diffuse samples in a training dataset (having a 'complex' distribution) until a desired simple data points distribution is reached. Each step in the process introduces more simplicity until all that is left is simple noise with original patterns obscured by this noise. An example of this noisy, simple, distribution may be a white noise distribution, or the noise may be distributed in any other way. During the reverse process (also known as reverse diffusion or denoising), the model generates a sample from the simple data points distribution, and then maps it back to a complex distribution by inverting the transformations. The diffusion model uses a conditioning, or prompt, to generate an image using the reverse process. That is, the conditioning is used to guide the denoising process and determine the content of the final image. In this way, diffusion models can generate new data samples by starting from a point in the simple distribution and diffusing it step-by-step to the desired complex data distribution. The whole training process can be thought of as destroying the training dataset samples through the successive addition of Gaussian noise, and learning to recover the data by reversing this noising process (i.e. denoising). Once trained, new samples can be generated by passing randomly sampled noise through the learned denoising process. This is the sampling process (i.e. new sample generation process).

[0025] An embedding is a representation of values or objects, like text, images or audio, that can be understood and processed by machine learning models. An embedding usually takes the form of a vector, and thus the terms "embedding" and "embedding vector" are used interchangeably herein. An embedding is therefore a mathematical representation of a data item (e.g. text, image, video, audio, etc.), and may represent some or all of the content of the data item. For example, an embedding may represent the semantic meaning of a data item. Embeddings make it possible for machine learning models to understand the relationships between different data items. Embeddings are normally analyzed within embedding space, i.e. a mathematical space in which similar items are positioned closer to one another than less similar items. For example, if embedding A for data item A is close to embedding B for data item B in embedding space, then data item A and data item B are similar in some way.

[0026] A vision / image encoder is a neural network model that is able to process an input image and output a single vector representing the visual content of the input image. An example vision encoder is the CLIP encoder.

[0027] The present techniques are concerned with image editing, and therefore, the conditioning or prompt referred to above must include both an image from which the editing starts, and a prompt that describes what should be edited. That is, ordinarily, a diffusion model may be conditioned on more than one input, and more than one type of input (i.e. both an image and a text prompt in this case). However, ordinarily, this creates the need for several "passes" i.e. a larger number of iterations to both denoise the noise image and generate an image that is adapted to all conditionings that are supplied to the diffusion model. This in turn means that the diffusion process takes longer and, importantly, requires more computational resources, making it unsuitable for use on resource constrained devices. That is, to apply all image editing conditions, conventional models need to run through three passes. One pass without any conditions, one pass with a conditioning, i.e. the input image, and one pass with both the conditioning and a prompt, i.e. text specifying what the output is meant to look like. This is computationally expensive and has also been shown not to give the most optimal output.

[0028] In contrast, the present techniques condition the noise (noisy) image that is used as the starting point in the denoising process, using the user-provided input image (which is to be edited) and a user-provided text prompt (which explains what changes / edits are required to the input image). This means that the information comprised in the conditionings that would normally need to be provided to the diffusion model during separate passes, can be incorporated into the initial noise image. Advantageously, there is then no need to perform separate passes for each conditioning that needs to be incorporated, which reduces the computational complexity of using a diffusion model for image editing. This noise image is generated by extracting relevant information from the input image and text prompt using specifically trained encoders. That is, an image encoder is used to extract salient information from the input image, and a text encoder is used to extract relevant information from the prompt.

[0029] The present techniques utilize various embeddings for the inputs and the generated edited image. Thus, the step of training the diffusion model may further comprise: generating, using the image encoder, an output image embedding for the generated edited image; and generating, using the image encoder, a target image embedding for the target output image. The step may further comprise: generating a first text caption for the input image and a second text caption for the target output image, wherein the text captions describe the images; and generating, using the text encoder, a first text embedding for the first text caption, and a second text embedding for the second text caption. The first text caption and the second text caption may be generated by any trained vision-language model that is capable of generating text captions for images. The first text embedding and the second text embedding may be generated using a text encoder, such as the CLIP text encoder.

[0030] In order to assess whether the training is successful or whether further training is required, the present techniques utilize one or more evaluation metrics.

[0031] One of these evaluation metrics is referred to as "Directional CLIP Similarity" or DCS. This measures the alignment between the performed edit (using a difference between the input image and the generated edited image) and the resulting change in text caption / descriptor (using a difference in captions for the input image and the original target output image).

[0032] Another of these evaluation metrics is referred to as "Directional Visual Similarity" or DVS. This measures the alignment between the performed edit (using a difference between the input image and the generated edited image) and the impact of the target output image (using a difference between the input image and the original target output image).

[0033] DCS and DVS can be used on their own to evaluate the present model. However, when they are used in combination they provide a better overall picture of the model's performance. In any case, both evaluation metrics compare the generated edited image and the input image. Thus, the step of training the diffusion model may further comprise: calculating a first difference between the output image embedding for the generated edited image and the input image embedding for the input image, i.e. , as explained in more detail below with reference to the Figures.

[0034] When DCS is used, the step of updating parameters of the diffusion ML model may comprise: calculating a first alignment measure between the generated edited image and the target output image of the triplet by: calculating a second difference between the first text embedding corresponding to the input image and the second text embedding corresponding to the target output image, i.e. , as explained in more detail below with reference to the Figures. Then, calculating the first alignment measure may comprise: calculating a cosine similarity between the first difference and the second difference. The first alignment measure means that it becomes possible to evaluate the performance of the diffusion model in terms of how similar an output image is to an input image, and in terms of how similar the prompt for which the output image has been created is to a description of the output image. Advantageously, image and text encoders for creating embeddings of, respectively, an image and a text prompt are used in the process to create the noise image (i.e. in the process to condition the initial noise image on the input image embedding and the prompt embedding), and thus, the same encoders can be used for this evaluation step also.

[0035] When DVS is used, the step of updating parameters of the diffusion model may comprise: calculating a second alignment measure between the generated edited image and the output image of the triplet by: calculating a third difference between the output embedding for the generated edited image and the target image embedding generated for the target output image, i.e. , as explained in more detail below with reference to the Figures. Then, calculating the second alignment measure may comprise: calculating a cosine similarity between the first difference and the third difference. The second alignment measure ensures that the generated output image is both sufficiently close to the input image and to the target image. Advantageously, this measure is robust to misalignment between image and text encoders. That is, the first and third differences are both based on obtaining embeddings for images, and therefore each measure should be "Weighted" equally, whereas when two different modalities are used (as for DCS), the measures may be skewed towards one modality if there is not perfect alignment between different modalities.

[0036] Updating parameters of the diffusion ML model may comprise: updating parameters of a neural network (e.g. a UNet, which is a type of convolutional neural network) of the diffusion ML model. That is, all other modules of the diffusion ML model may be frozen. This means, that, for example, the encoders used may not be updated, but only the neural network that is responsible for the denoising process may be trained. This means that pre-trained components may be used for the parts of the diffusion ML model. For example, pre-trained and highly accurate encoders are available for most modalities and using these is more resource efficient than retraining such encoders, or finetuning them during training of the ML model.

[0037] Generating an initial noise image may comprise: obtaining a sample from a Gaussian distribution having a mean and variance that is conditioned on the input image embedding and the prompt embedding, and generating an initial noise image using the obtained sample. That is, the parameters of the Gaussian may be adapted to the embeddings and to obtain noise from which to generate the noise image, samples are taken from this Gaussian. In this way, the input image and prompt are incorporated into the noise image. A Gaussian distribution is a well-known and conventional distribution but it will be appreciated that other probabilistic distributions that reflect the image and prompt embeddings may be used.

[0038] Preferably, the method may further comprise: generating, using the input image embedding for the input image, a mean and variance for the input image; and generating, using the prompt embedding for the text prompt, a mean and variance for the text prompt. As mentioned above, noise for the noise image may be drawn from, for example, a Gaussian distribution. In that case, the Gaussian distribution may be conditioned on the mean and variance of both the input image and prompt embeddings. That is, obtaining a sample from a Gaussian distribution may comprise: summing the generated mean for the input image and the generated mean for the text prompt to generate a total mean; summing the generated variance for the input image and the generated variance for the text prompt to generate a total variance; and using the total mean and total variance as the mean and variance of the Gaussian distribution. It will be appreciated that any other suitable way of combining the mean and variance of the image and prompt may be used. For example, the method does not necessarily have to draw noise from a single Gaussian, but could also draw noise from two separate Gaussians, which represent the mean and variance of the image and prompt conditionings each.

[0039] In a second approach of the present techniques, there is provided a server for training a diffusion machine learning, ML, model to edit images, the server comprising: at least one processor coupled to memory for: obtaining a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt to edit the input image, and a target output image that is an edited version of the input image according to the text prompt; generating, for each triplet of training data: an input image embedding for the input image, using an image encoder of the diffusion ML model, a prompt embedding for the text prompt, using a text encoder of the diffusion ML model, and an initial noise image for the diffusion ML model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding; and training the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by: denoising, for each triplet, the generated initial noise image to thereby generate an edited image using the diffusion ML model; and updating parameters of the diffusion ML model to improve alignment between the generated edited image and the output image of the triplet.

[0040] Advantageously, the training is performed on a server, rather than on resource-constrained user devices, which would likely not have the capabilities for the training. Furthermore, this enables a trained model that is centrally trained to be deployed post-training on to many user devices.

[0041] The features described above with respect to the first approach apply equally to the second approach and therefore, for the sake of conciseness, are not repeated.

[0042] In a third approach to the present techniques, there is provided a computer-implemented method for using a trained diffusion machine learning, ML, model to edit images, the method comprising: obtaining an input image and a text prompt to edit the input image; generating: an input image embedding for the input image using an image encoder of the trained diffusion ML model, a prompt embedding for the text prompt using a text encoder of the trained diffusion ML model, and an initial noise image for the trained diffusion ML model to use to generate edited images, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding; and denoising, using the trained diffusion ML model, the generated initial noise image to thereby generate an edited image.

[0043] Advantageously, using the present techniques means that generating an edited image can done in a much more resource efficient way. In particular, having to only perform one pass, with all required information incorporated in the noise image, means that the model may also be run on resource-constrained devices such as smartphones etc.

[0044] It can be seen that many features of the third approach are similar to the first approach and therefore, the descriptions of these features provided in relation to the first approach apply equally to the third approach.

[0045] Generating an initial noise image may comprise: obtaining a sample from a Gaussian distribution having a mean and variance that is conditioned on the input image embedding and the prompt embedding, and generating an initial noise image using the obtained sample.

[0046] That is, the parameters of the Gaussian may be adapted to the embeddings and to obtain noise from which to generate the noise image, samples are taken from this Gaussian. In this way, the input image and prompt are incorporated into the noise image. A Gaussian distribution is a well-known and conventional distribution but it will be appreciated that other probabilistic distributions that reflect the image and prompt embeddings may be used. Sampling from a probabilistic distribution is a process that is not computationally intensive, so being able to incorporate conditionings in this way means that the overall process requires significantly fewer computational resources.

[0047] The method of using the trained model may further comprise: generating, using the input image embedding for the input image, a mean and variance for the input image; and generating, using the prompt embedding for the text prompt, a mean and variance for the text prompt. Encoders used to generate these embeddings may be the same as those used during the training process described above. This is to ensure consistency in generation of embeddings, and may be necessary for best results, given that training was performed using these encoders. Noise for the noise image may be drawn from, for example, a Gaussian distribution. In that case, the Gaussian distribution may be conditioned on the mean and variance of both the input image and prompt embeddings. That is, obtaining a sample from a Gaussian distribution may comprise: summing the generated mean for the input image and the generated mean for the text prompt to generate a total mean; summing the generated variance for the input image and the generated variance for the text prompt to generate a total variance; and using the total mean and total variance as the mean and variance of the Gaussian distribution.

[0048] In a fourth approach of the present techniques, there is provided an electronic device for using a trained diffusion machine learning, ML, model to edit images, the electronic device comprising: at least one processor coupled to memory for: obtaining an input image and a text prompt to edit the input image; generating: an input image embedding for the input image using an image encoder of the trained diffusion ML model, a prompt embedding for the text prompt using a text encoder of the trained diffusion ML model, and an initial noise image for the trained diffusion ML model to use to generate edited images, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding; and denoising, using the trained diffusion ML model, the generated initial noise image to thereby generate an edited image.

[0049] The electronic device may further comprise a display for displaying the generated edited image.

[0050] The of the electronic device memory may store instructions that, when executed by the at least one processor individually or collectively, cause the at least one processor to perform the methods described herein.

[0051] The features described above with respect to the third approach apply equally to the fourth approach and therefore, for the sake of conciseness, are not repeated.

[0052] In a related approach of the present techniques, there is provided a computer-readable storage medium comprising instructions which, when executed by a processor, causes the processor to carry out any of the methods described herein.

[0053] As will be appreciated by one skilled in the art, the present techniques may be embodied as a system, method or computer program product. Accordingly, present techniques may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[0054] Furthermore, the present techniques may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.

[0055] Computer program code for carrying out operations of the present techniques may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise sub-components which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.

[0056] Embodiments of the present techniques also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to carry out any of the methods described herein.

[0057] The techniques further provide processor control code to implement the above-described methods, for example on a general purpose computer system or on a digital signal processor (DSP). The techniques also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD-ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as Python, C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. The techniques may comprise a controller which includes a microprocessor, working memory and program memory coupled to one or more of the components of the system.

[0058] It will also be clear to one of skill in the art that all or part of a logical method according to embodiments of the present techniques may suitably be embodied in a logic apparatus comprising logic elements to perform the steps of the above-described methods, and that such logic elements may comprise components such as logic gates in, for example a programmable logic array or application-specific integrated circuit. Such a logic arrangement may further be embodied in enabling elements for temporarily or permanently establishing logic structures in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using fixed or transmittable carrier media.

[0059] In an embodiment, the present techniques may be realized in the form of a data carrier having functional data thereon, said functional data comprising functional computer data structures to, when loaded into a computer system or network and operated upon thereby, enable said computer system to perform all the steps of the above-described method.

[0060] The method described above may be wholly or partly performed on an apparatus, i.e. an electronic device, using a machine learning or artificial intelligence model. The model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.

[0061] As mentioned above, the present techniques may be implemented using an AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning. Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.

[0062] The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.

[0063] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0064] Broadly speaking, embodiments of the present techniques provide a diffusion model for editing images according to a user's instructions, which is also less computationally intensive than existing techniques for doing the same. Not only are the resulting edited images produced in less time (because of the reduced number of computations), but the present techniques also result in high-quality images. The present techniques effectively alter the starting point for the denoising performed by the diffusion model, so that the starting point already contains some information about the required edit. This guides or 'conditions' the diffusion model to generate an edited image that corresponds to the user's instructions. However, it will be understood that this is one example use case and the present techniques may be used for other applications outside of image editing, including video editing.

[0065] The present techniques utilize Explicit Conditioning (EC). Rather than working on another candidate replacement for CFG or a distillation mechanism to cease the extra computation cost, the present techniques investigate generative modelling that natively supports conditional information. Instead of diffusing the image distribution into a noise distribution, and adding the condition as an additional input, the present techniques consider diffusing the data distribution into a purposefully designed distribution made of noise and condition. Explicit conditioning alleviate the computational burden of CFG: the proposed generative framework natively incorporates conditional information, resulting in a reduction of denoising passes.

[0066] By explicitly conditioning each term, the need for the corresponding denoising pass in CFG is diminished. The current conditioning approach for image editing by Brooks et al requires three denoising passes to achieve plausible results. By explicitly conditioning both context input and the instruction prompt, the image editing framework using the present EC approach outperforms CFG based sampling using only pass, being not only faster but also improving the performance.

[0067] To aid the understanding of the present techniques and the advantages, a brief discussion of existing related techniques is provided.

[0068] Conditional generative models: Image generation is dominated by diffusion and flow models which yield state-of-the-art (SOTA) diverse high-quality image generation, far surpassing GAN alternatives. There has been a surge in the development of diffusions and flows.

[0069] At a high level, diffusions and flows transform distributions into distributions. For example, unconditional image generation learns to transform the underlying image data distribution into a multivariate normal distribution. A key design choice, in this paradigm, is whether the transformation occurs in the original pixel space or in a latent space. SOTA approaches are dominated by the latter.

[0070] Another key design criteria is whether the transformation from data into noise is done in a discrete number of time steps, or in a continuous range of time steps. Diffusion models owe their success to rapid progress on the discrete time setting, formally modelled by hierarchical Variational Autoencoders, VAEs.

[0071] An important practical observation early on was that although diffusion models needed to be trained with a high number of discrete time steps, they could be sampled with a low number of time steps. Another key practical observation of diffusion models was that their training and sampling very much depended on the logarithm of the Signal-to-Noise Ratio (SNR) which was rigidly tied to the predefined discrete time steps. This led to the development of continuous time diffusion models, seen as the limit of hierarchical VAEs or as Stochastic Differential Equations (SDEs). Despite the success of diffusion models, faster training and more diverse and higher quality sampling was shown to depend on selecting data, noise or velocity parametrizations. Conditional Flow Matching (CFM) provides an alternative natural parametrization for diffusion, namely via optimal transport, that yields SOTA models. More broadly, general diffusion and flow models can be recast in a unified framework, viewed under Evidence Lower BOund (ELBO) objectives.

[0072] CFG and selected alternatives: Classifier Free Guidance (CFG) is a key pillar of SOTA diffusion and flow systems, defining a conditional framework to image generation. Great focus is on its theoretical underpinning and candidate alternatives such as Independent Condition Guidance and Time Step Guidance (TSG). Advantageously, the present techniques provide a new avenue of research for conditional image generation with a new Explicit Conditioning (EC) mechanism that is on par with CFG in terms of sample diversity and quality.

[0073] Single pass distillation: Previous work can distil CFG from three passes to one pass, sensibly reducing latency, while matching CFG on diversity and quality generation. More recently, D Ahn et al proposed to extend the distillation-based line of methods to learn a mapping from normal Gaussian to a guidance-free Gaussian distribution. The work of D Ahn et al shares similarities with the present techniques in terms of exploring the sampling distribution. However, D Ahn et still relies on guidance-dependent teachers to learn such a mapping. Advantageously, the present techniques achieve the same goal without the painstaking of distillation.

[0074] Image editing: The most notable example of diffusion based image editing causes an unedited image to be transformed into an edited image by passing the editing prompt to the denoising network. This is achieved by simple concatenation along the channel dimension on the first layer of Stable Diffusion (SD).

[0075] Editing via distributions rather than inputs: While the vast majority of diffusion and flow models transform image into noise through some process, there are exceptions that transform between more general data distributions. Notable instances include class transformations and super resolution. However, unlike the present techniques, no prior work transforms image and text based distributions into image distributions directly without text conditioning as an additional input. Although the present focus is on numerical experiments on image editing tasks, the present techniques are applicable to a much wider class of conditional generation problems such as video editing and text-to-image tasks. More generally, the present techniques are applicable to other conditional diffusion / flow models.

[0076] The present techniques are now explained in more detail.

[0077] In an embodiment of the disclosure, the operations of training and inference (use) of the diffusion model described with reference to FIGS. 1A and 1B may be performed by an electronic device or a server, and are not limited to the disclosed examples. Alternatively, some of the operations may be distributed and performed by a plurality of devices, such as the electronic device and the server. For convenience of description, the operations of FIGS. 1A and 1B may be described as being performed by a server or an electronic device. But, the present disclosure is not limited to the disclosed examples.

[0078] Figure 1A is a flowchart of example steps to train a diffusion model to edit images.

[0079] At step S100, the method may include obtaining a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt to edit the input image, and a target output image that is an edited version of the input image according to the text prompt.

[0080] At step S102 and S104, the method may include generating, for each triplet of training data: (i) an input image embedding for the input image, using an image encoder of the diffusion ML model, (ii) a prompt embedding for the text prompt, using a text encoder of the diffusion ML model, and (iii) an initial noise image for the diffusion ML model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding.

[0081] At step S106 and S108, the method may include training the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by: denoising, for each triplet, the generated initial noise image to thereby generate an edited image using a denoising network of the diffusion ML model , and updating one or more parameters of the diffusion ML model to improve alignment between the generated edited image and the output image of the triplet.

[0082] In step S100, training data is obtained. In particular, triplets of training data are obtained, with each triplet comprising an image that is to be edited, a prompt describing how the image is to be edited and a target image that shows what the ideal edited image would look like. The prompt may be a text prompt, i.e. a string of text describing desired edits. Meanwhile, a diffusion model may also be referred to as a diffusion ML (machine learning) model.

[0083] At step S102, input image embeddings and prompt embeddings are generated. Generating embeddings means inputting the input image and prompt into encoders that create a latent space embedding for each. The input image is input to the image encoder to generate the input image embeddings, and the prompt is input to the CLIP encoder to generate the prompt embeddings. These embeddings represent the content of both the input image and the prompt. Diffusion models require a noisy image to start with, and this noisy image is then iteratively denoised using a machine learning model(e.g., a denoising network like denoising Unet).

[0084] In an embodiment of the disclosure, the electronic device may generate an output image embedding for the generated edited image. The electronic device may generate, using the image encoder, a target image embedding for the target output image. The electronic device may generate a first text caption for the input image and a second text caption for the target output image, wherein the first text caption and the second text caption describe the images. The electronic device may generate, using the text encoder, a first text embedding for the first text caption, and a second text embedding for the second text caption.

[0085] At step S104, the electronic device may generate an initial noise image. This uses the embeddings of the input image and prompt generated in the previous step. That is, the noise image which the machine learning model denoises during a reverse diffusion process already contains the information needed to generate a target edited image. This is done by incorporating all required conditionings into the noise present in the noise image.

[0086] In an embodiment of the disclosure, the electronic device may obtain a sample from a Gaussian distribution having a mean and variance that is parameterized by the input image embedding and the prompt embedding. For example, the electronic device may generate, using the input image embedding for the input image, a mean and variance for the input image. The electronic device may generate, using the prompt embedding for the text prompt, a mean and variance for the text prompt. The electronic device may sum the generated mean for the input image and the generated mean for the text prompt to generate a total mean. The electronic device may sum the generated variance for the input image and the generated variance for the text prompt to generate a total variance. The electronic device may use the total mean and total variance as the mean and variance of the Gaussian distribution.

[0087] In an embodiment of the disclosure, the electronic device may generate an initial noise image using the obtained sample. Meanwhile, for ease of explanation, the term 'initial noise image' is used herein. However, the 'initial noise image' may refer to, or may be represented in, a latent space as an initial noise latent corresponding to the initial noise image. In an embodiment of the disclosure, the initial noise latent may be conditioned on input-image embeddings and prompt embeddings.

[0088] In an embodiment of the disclosure, the electronic device may train the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet.

[0089] At step S106, the electronic device may denoise using the machine learning model as like a denoising network. This step may use conventional denoising techniques that are used in other diffusion models. This is possible, because all information is present in the noise itself, so there is no need to adapt the denoising process itself.

[0090] In an embodiment of the disclosure, by explicitly reflecting conditioning information in a noise distribution (an initialization distribution), the denoising network may be configured so as not to necessarily require additional inputs related to the conditioning information, such as prompt embeddings. For example, an electronic device may perform denoising without inputting conditioning information (e.g., a prompt, a prompt embedding, etc.) to the denoising network. Alternatively, the electronic device may input the conditioning information to the denoising network. For example, the electronic device may input 77 CLIP tokens (prompt embeddings) to the denoising network. However, unlike the foregoing prior art documents, a high-quality edited image may be obtained by invoking the UNet only once, regardless of whether the conditioning information (e.g., a prompt, a prompt embedding, etc.) is provided to the denoising network.

[0091] At step S108, the electronic device may update one or more parameters of the diffusion model to improve alignment between the generated edited image and the output image of the triplet.

[0092] In an embodiment of the disclosure, the electronic device may calculate a first difference between the output image embedding for the generated edited image and the input image embedding for the input image. The electronic device may calculate a first alignment measure between the generated edited image and the target output image of the triplet by: calculating a second difference between the first text embedding corresponding to the input image and the second text embedding corresponding to the target output image, and calculating a cosine similarity between the first difference and the second difference. The electronic device may calculate a second alignment measure between the generated edited image and the output image of the triplet by: calculating a third difference between the output embedding for the generated edited image and the target image embedding generated for the target output image, and calculating a cosine similarity between the first difference and the third difference.

[0093] In an embodiment of the disclosure, the denoising network may obtain, as an input, a noise state generated by combining an embedding corresponding to a target image with a noise sample drawn from a Gaussian distribution whose mean and variance are determined based on an input image embedding and a prompt embedding. In addition, the denoising network may update parameters of the network according to a loss function based on an error between (i) a standard noise used to generate the noise sample and (ii) an estimated standard noise output based on the noise state.

[0094] Finally, to train the model, the edited image that is generated using the denoising process is compared to a target. This target can be the target output image of the triplet of training data. The parameters of the diffusion ML model can be updated using the comparison between the target and image output by the machine learning model.

[0095] Meanwhile, during the denoising process, the machine learning model tries to generate an image that is close to an original image, i.e. remove the noise, while adhering to conditionings given to the machine learning model. Conditionings act as a guide of what the output image should look like and allow the machine learning model to generate images that are quite different from an input image in this way. Conditionings also allow the machine learning model to generate an image that is edited relative to an input image. In this case, both the input image and the prompt are needed as conditionings. Using conventional techniques, this means that the machine learning model needs to repeat the denoising process three times - once using just the noisy image, once taking into account the conditioning provided by the input image, and a final time taking into account the conditioning provided by the prompt. However, the present techniques use a more efficient way to incorporate conditionings to allow the machine learning model to perform image editing in a single pass.

[0096] Figure 1B is a flowchart of example steps for using a trained diffusion model to edit images (i.e. inference time method). At inference time, a very similar process to that described above with respect to training is performed, but the model is not optimised - instead, the model weights and parameters are frozen. In particular, at step S200, the method may include obtaining an input image and a text prompt to edit the input image. At step S202 and step S204, the method may include generating an input image embedding for the input image using an image encoder of the trained diffusion ML model, a prompt embedding for the text prompt using a text encoder of the trained diffusion ML model, and an initial noise image for the trained diffusion ML model to use to generate edited images, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding. At step S206, the method may include denoising, using the trained diffusion ML model, the generated initial noise image to thereby generate an edited image.

[0097] Steps S200, S202, and S204 may correspond to steps S100, S102, and S104, respectively. Accordingly, redundant descriptions thereof will be omitted.

[0098] That is, at step S200, first an input image and a prompt explaining how the image should be edited are obtained. The electronic device may obtain an input image and a text prompt to edit the input image.

[0099] Then, embeddings of both the input image and prompt are generated at step S202. The electronic device may generate an input image embedding for the input image using an image encoder of the trained diffusion model, and a prompt embedding for the text prompt using a text encoder of the trained diffusion model. This may be done using the same encoders as those used during training, i.e. separate encoders for the image and the prompt. These encoders may be trained to perform well for each specific modality, i.e. trained to perform meaningful embeddings for images or text.

[0100] At step S204, an initial noise image which includes information on the input image and prompt is generated using the generated embeddings for both. The electronic device may generate an initial noise image for the trained diffusion model to use to generate edited images, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding.

[0101] At step S206, this input image is then denoised using the denoising network of the diffusion machine learning model to arrive at an edited output image. The electronic device may denoise, using a denoising network of the trained diffusion model, the generated initial noise image to generate an edited image.

[0102] Figure 2 shows images generated using implicit conditioning and the explicit conditioning of the present techniques. Specifically, a context image 210 may be an input image to be input to the diffusion model. And, the target image 220 may be an target output image. In an embodiment of the disclosure, the EC image 230 and the CFG images 240, 250, 260 may be results generated based on the instruction prompt, "Make her a bride." And, an EC image 230 may be generated using the explicit conditioning of the present techniques, and may be the edited result produced by the proposed Explicit Conditioning (EC) method using a single denoising pass per step. CFG images 240, 250, 260 show images generated using implicit conditioning. Furthermore, the EC image 230 shows the results of the present techniques using only a single pass, whereas the CFG images 240, 250, 260 may be the results of implicit conditioning are shown for 1, 2 and 3 passes, respectively. It is clear, the implicit conditioning produces good results needing 3 passes, whereas the present techniques produce good results using 1 pass. Meanwhile, depending on whether the prompt condition is directly incorporated into the sampling distribution, the conditioning scheme may be referred to as implicit conditioning or explicit conditioning.

[0103] Figure 3 shows a schematic diagram of the present techniques during training.

[0104] In an embodiment of the disclosure, the input image 320 may be converted into a normal distribution using an image encoder 322 which generates an input image embedding of the image which can in turn be related to the normal distribution. Meanwhile, the input image may be referred to as a context image. And, the image encoder may referred to as a SD(stable diffusion) encoder. A CLIP encoder 312 may be used to obtain prompt embeddings 314 from a text prompt 310. A variational autoencoder 316 maps the pooled embeddings to a latent distribution . Meanwhile, the variational autoencoder (VAE) 316 may be referred to as a prompt VAE encoder, and it takes the pooled prompt embeddings. This is done to ensure comparability with the embedding / distribution obtained for the image. These distributions (for the prompt and the image) are merged and fused together to model the noise which is added to the image that is later denoised. A noise vector is sampled from the joint distribution 340, representing noise to be applied to a noise image and this is denoised through a UNet 350 to obtain an output image. Finally, the denoised latent embedding is decoded into an image embedding. Different methods may be used to compare the decoded output image to a target output image 330 included in the triplet of training data. Meahwhile, the UNet may also be referred to as a denoising network or a denoising UNet. In an embodiment of the disclosure, the denoising UNet 350 receives, as input, a noisy latent representation 342 generated by combining (i) a target image embedding of a target output image and (ii) a noise vector (noise sample) drawn from the joint distribution (a conditional Gaussian distribution). Meanwhile, the noisy latent representation 342 may be referred to as an initial noise image.

[0105] The present techniques utilize various embeddings for the inputs and the generated edited image. Thus, the step (S108 in Figure 1A) of training the diffusion model may further comprise: generating an output image embedding for the generated edited image; and generating, using the image encoder 332, a target image embedding for the target output image 330. The step may further comprise: generating a first text caption for the input image 320 and a second text caption for the target output image 330 (using any trained vision-language model), wherein the text captions describe the images; and generating, using the text encoder 316, a first text embedding for the first text caption, and a second text embedding for the second text caption.

[0106] In order to assess whether the training is successful or whether further training is required, the present techniques utilise one or more evaluation metrics. The evaluation metrics generally compare the generated edited image and the input image. Thus, the step (S108 in Figure 1A) of training the diffusion model may further comprise: calculating a first difference between the output image embedding for the generated edited image and the input image embedding for the input image 320, i.e. , as explained in more detail below.

[0107] For example, training the diffusion model may comprise: calculating a first alignment measure between the generated edited image and the target output image 330 of the triplet by: calculating a second difference between the first text embedding corresponding to the input image 320 and the second text embedding corresponding to the target output image 330, i.e. , as explained in more detail below. Then, calculating the first alignment measure may comprise: calculating a cosine similarity between the first difference and the second difference. This approach is termed Directional CLIP Similarity Metric (DCS) and described in more detail below.

[0108] Additionally or alternatively, training the diffusion model may comprise: calculating a second alignment measure between the generated edited image and the output image of the triplet by: calculating a second alignment measure between the generated edited image and the output image of the triplet by: calculating a third difference between the output embedding for the generated edited image and the target image embedding generated for the target output image 330, i.e. , as explained in more detail below with reference to the Figures. Then, calculating the second alignment measure may comprise: calculating a cosine similarity between the first difference and the third difference. This approach is termed Directional Visual Similarities (DVS) and described in more detail below.

[0109] Meanwhile, the denoising UNet 350 may obtain 77 CLIP embeddings derived from a prompt embedding and use the obtained CLIP embeddings as additional inputs. Alternatively, the denoising UNet 350 may perform denoising without receiving any prompt-related input, by utilizing prompt-related information reflected in the joint distribution 340.

[0110] Figure 4 shows a schematic diagram of the present techniques during inference.

[0111] In an embodiment of the disclosure, the prompt 410 and input image 420 are encoded to obtain a prompt embedding and input image embedding. The same encoders 312, 322, 332 as during training of Figure 3 may be used during inference. That is, a CLIP encoder 412 may be used to obtain embeddings from a text prompt 410, and an image encoder 422 may be used to encode the input image 420. A variational text encoder 416 maps the pooled embeddings to a latent distribution . These prompt and image distributions are merged together to model the noise as a joint distribution(a conditional Gaussian distribution) 440. An electronic device may obtain an initial noisy latent representation by sampling from a joint distribution (conditional Gaussian distribution). Meanwhile, the noisy latent representation 342 may be referred to as an initial noise image. The electronic device may input the initial noisy latent representation to a denoising UNet to perform one or more denoising steps and generate an output latent representation. The output latent representation from the denoising UNet may be used to obtain an output / edited image from the noise image. The UNet's output is run through an image decoder to obtain an image from the denoised embedding generated by the UNet.

[0112] Meanwhile, the denoising UNet 450 may obtain 77 CLIP embeddings derived from a prompt embedding and use the obtained CLIP embeddings as additional inputs. Alternatively, the denoising UNet 450 may perform denoising without receiving any prompt-related input, by utilizing prompt-related information reflected in the joint distribution 440.

[0113] In the following, we first describe CFG and then provide a detailed description of EC of the disclosure.

[0114] The foundation of diffusion and flow image generation is in the unconditional setting i.e., generate an image from a latent through a process

[0115] (1)

[0116] with density

[0117]

[0118] The conditional setting is based on combining the aforementioned density with

[0119]

[0120] in a gamma-powered distribution fashion

[0121] (2)

[0122] Diffusions and flows are grounded on solid modelling theory that ensures they excel at approximating individual scores

[0123] (3)

[0124] (4)

[0125] with which noise can be transformed into data. However, that theory has no approximation guarantees about the actual quantity of interest, namely the joint score

[0126] (5)

[0127] Most critically, the approximation

[0128] (6)

[0129] does not guarantee that noise can be transformed into data. Nevertheless, this is exactly what Classifier Free Guidance (CFG) does, i.e. it approximates Eq. (5) with (6).

[0130] Continuous time diffusions and flows are trained by abstracting the schedule via the log signal-to-noise ratio and training a parametric model by minimizing for each

[0131] .

[0132] Critically, gradient descent may yield such that

[0133] (7)

[0134] (8)

[0135] In essence, CFG up-weights the probability of the samples where the conditional likelihood of an implicit classifier is higher:

[0136] . (9)

[0137] which is the result of shifting the score function with the gradient of the classifier at each sampling step.

[0138] (10)

[0139] It is noted that CFG approximates the gradient of the implicit classifier in Eq. 10 by which results in Eq. 6.

[0140] Intuitively, since and are sampled independently, is not guaranteed to be sampled from a subspace that corresponds to the given condition in the generative modeling transformation from noise to data. This is particularly problematic in the early sampling steps with where is close to noise, and contains minimal information about . CFG corrects the sampling at each time step by pushing it towards the sharper modes of the classifier. However, as seen in Eq. 6, it requires an extra denoising function evaluation.

[0141] In the following, CFG for instruction-based image editing is first reviewed, where two extra denosing function evaluations are required, and explicit conditioning is discussed, which up-weights the probability of samples with higher conditional likelihood effortlessly by embedding explicitly in .

[0142] CFG for Image Editing: Existing methods leverage CFG extended from single to multiple conditions. Given unedited image and text prompt , the image editing task involves generating a new image which is the result of applying on , posing it as an image generation model conditioned on both . Notable works use implicit conditioning (Eq. 1) and provides as an additional input via concatenation with . They use a pre-trained text-to-image model to provide a good initialization and incorporate as another input via cross-attention. As a result, they need to leverage CFG with respect to both conditionings to achieve plausible results where the denoising model is called three times at each iteration:

[0143] (11)

[0144] where and are the image and prompt guidance scales, respectively. For a result that takes both unedited image and text prompt into account, it is necessary to perform all three passes by setting . This conditioning mechanism is called "implicit" because are implicitly embedded in .

[0145] In the next section, explicit conditioning is discussed, which achieves CFG-like generation quality and diversity but with just a single pass.

[0146] Explicit Conditioning: In this section, the Explicit Conditioning (EC) method of the present techniques is discussed.

[0147] Figure 3 is a schematic diagram showing the training process for Explicit Conditioning. Here, and as explained above with reference to Figure 1A, the method comprises obtaining the mean and variance encoder of the context image via an image encoder, shown in blue. For the instruction prompt, a text encoder, shown in yellow, is used, which has the same latent dimension as the SD, e.g., 4 x 64 x 64. The text encoder takes the pooled CLIP embeddings, shown in red, as input and maps them to the corresponding mean and variance. The context and prompt mean and variances are fused to form a Gaussian used for sampling in the diffusion process (Eq. 15 below), shown in purple, as input to the diffusion model (e.g. a UNet) for the sake of consistency with the internalization.

[0148] Figure 4 is a schematic diagram showing the inference process for using a diffusion model trained to perform Explicit Conditioning. Here, and as explained above with reference to Figure 1B, the method comprises calling the denoising model recursively starting from a point sampled from a Gaussian that its mean and variance is a fusion of context image and instruction prompt.

[0149] The mathematical underpinnings of Explicit Conditioning are now explained.

[0150] Instead of diffusing or flowing data into noise, where the conditioning term is implicitly embedded in via an additional input, the present techniques diffuse or flow data into a purpose-built distribution that explicitly incorporates the conditioning information:

[0151] (12)

[0152] To be consistent with the diffusion and flow model frameworks, is defined to be sampled from Gaussian as follows:

[0153] (13)

[0154] where are functions that project a given conditioning term into the same dimension as . Intuitively, instead of sampling from a normal Gaussian as done in traditional diffusions and flows, the present techniques sample from a Gaussian with mean and variance conditioned on functions of .

[0155] Explicit Conditioning for Image Editing: The technical detail on how EC is used for an image editing task is now provided. The present techniques build on a widely used Stable Diffusion (SD) framework and develop the present method in two stages. First, the present techniques condition explicitly only on context, i.e. unedited image, and then the present techniques extend the explicit conditioning for both context and editing text. It is shown that by incorporating only context, the model requires guidance only for text during sampling (i.e. passes with and without text), while incorporating both results in guidance-free sampling (i.e., only pass).

[0156] In the following, we describe EC of this closure.

[0157] EC for Context: The diffusion model framework operates in the latent space of a Variational Autoencoder (VAE), that allows us to simply use it to obtain the mean and variance of the explicitly conditioned Gaussian on the context image, as shown in Figures 3 and 4

[0158] (14)

[0159] where are the pre-trained SD-VAE mean and variance estimators. Note that the same VAE is used to obtain edited image embeddings , as a result, , are of the same dimensions, i.e. 4 x 64 x 64. Here, only is explicitly conditioned, and is used as a score function that does not take as an additional input and . CFG is required in this case only to guide the sampling with respect to the prompt, resulting in sampling passes

[0160] .

[0161] EC for Context and Prompt: To incorporate the prompt, it is necessary to define a function that takes , as input and generates a tensor with the same dimension as . For this purpose, the best results are achieved via training a tiny prompt VAE that takes the pooled CLIP embeddings of the prompt and has the latent space to a latent distribution of the same dimension as . During the image editing training via EC, its frozen encoder is used to map pooled CLIP embeddings of the instruction prompt to a latent distribution of the same dimension as , resulting in the following two functions: . Then is sampled from the following

[0162] (15)

[0163] Figure 5 shows quantitative results of the present techniques. The corresponding score function presumably is independent of any auxiliary input, where and is free from guidance for high-quality sampling. The model is initialized with a text-to-image model where full 77 CLIP tokens are fed to the score function via cross attention. It was found that providing the model with the same clip token results in faster convergence and slightly better performance. Cases with and without an extra 77 CLIP tokens are compared in Figure 5.

[0164] Prompt VAE architecture: Prompt VAE is a tiny VAE with a latent space of dimension 4 x 64 x 64. The encoder takes the 1 x 768 dimensional pooled CLIP embeddings of the instruction prompt as input and it has three layers. First, a linear layer projects the input from 768 to 1024 reshaped to 1 x 32 x 32. Then the number of channels are increased to 16 via a 1 x 1 Conv layer, resulting in a 16 x 32 x 32 tensor. Then a parameter-efficient spatial upsampling operation of PixelShuffle is used to obtain a 4 x 64 x 64 tensor. The mean and variance are obtained via two separate 1 x 1 Conv layers. The decoder mirrors of the encoder. The encoder and decoder are trained with standard VAE losses, , reconstruction, and KL losses. A LeakyReLU and an instance normalization layer follow all Conv and projection layers. The rationale behind the present design choices here is to first, keep the number of changes to the original signal minimal and second, transform it to the desired dimension that has a Gaussian distribution.

[0165] Sampling: During inference, for a given , , their corresponding mean variances are found via image and prompt encoders. Then a starting point is sampled from Eq. 15 and 10 DDIM steps are performed.

[0166] Experiments

[0167] In this section, the proposed explicit conditioning method is evaluated for image editing task and compared extensively with Classifier Free Guidance (CFG). Moreover, a new metric - directional visual similarities (DVS) - is introduced, as an alternative to directional CLIP similarity. It is shown both quantitatively and qualitatively that explicit conditioning is more robust and faster and outperforms CFG in terms of both metrics. The results of Brooks et al are regenerated and 10 sampling steps are used for all methods throughout the experiment.

[0168] Experimental Setup: The Stable Diffusion V1.5 architecture is used due to its affordable computational cost and performance. Initialisation is with the text-to-image pre-training, and fine-tuning for the image editing task is performed on the training set of instruct-pix2pix dataset (Brooks et al) using v-parametrization. Evaluation is on 2000 samples from the test set. The model is trained using model using 8ХA100 GPUs.

[0169] Evaluation Metrics: Two evaluation metrics are used: (i) the widely used directional CLIP similarity metric (DCS), (ii) the proposed Directional Visual Similarities (DVS) that is discussed herein. A natural way to think of a metric is to measure how the difference between the generated image and the input image - and the change in captions - is aligned with the difference between the target image and the input image . This alignment is measured in the latent space through CLIP encoders. Having access to the entries in the form of triplets which corresponds to the context image, instruction prompt, and the target edited image, respectively. Moreover, the captions of , in the datasets is given, denoted by respectively. Having the generated image denoted by , directional clip similarity computes the following quantity:

[0170]

[0171] where is the cosine similarity and , are CLIP's vision and text encoders, respectively. This metric uses the text embeddings because it takes into account that the training dataset is synthetic where the input and target images are generated synthetically via text-to-images using their corresponding captions, making the target image truthfulness depend on the image synthesizer behavior. Note the image editing task training is purely based on the target image and the captions are not presented during the training process therefore this metric suffers from the misalignment injected by the image synthesizer. Moreover, the misalignment between the CLIP vision and text encoders imposes another bias. As a result, this metric is usually considered jointly with CLIP visual similarity between the target and the output.

[0172] The present techniques introduce directional CLIP visually similarity which is aligned with the training objective function and robust to the deviations between the CLIP vision and text encoders:

[0173]

[0174] where is the cosine similarity and is a vision encoder. A CLIP vision encoder is used to be consistent with DCS. This metric measures the alignment of the transformation context-to-generated image - to -, and the transformation context-to-target image - to . Therefore, DVS is more aligned with the training objective where the task is to generate the target image regardless of how they have been generated. As shown in Figure 5, DCS and DVS are correlated, however, DCS is more sensitive to the true output that might not be depicted in the pseudo-gt. On the other hand, DVS is more consistent with the present qualitative assessment.

[0175] Quantitative Evaluation: The following three cases are used for the comparisons. First, the implicit conditioning mechanism is used as a baseline that heavily relies on CFG. The pre-trained model of Brooks et al is used, and to determine the impact of CFG, evaluation is of pass, i.e. and passes, i.e. . Note that it is infeasible to find an optimal set of guidance scales for the entire dataset as it is sample dependent. Second, and for the purpose of analyses, the trained explicitly conditioned model of the present techniques is used, with conditioning only on the context image. Conditioning on the context image is independent of context guidance during the sampling but yet relies on prompt guidance. Sampling is with pass, i.e. , and passes, i.e. . Finally, the explicitly conditioned model of the present techniques is used, with conditioning on both context image and instruction prompt, which is guidance-free and requires only pass. Figure 5 summarises the results. It can be seen that the explicitly conditioned model outperforms the currently used implicit conditioning in terms of DCS and DVS metrics while being faster. Two cases of the present method are compared with and without the 77 CLIP tokens as additional input to the UNet. It was found that providing the UNet with the 77 extra CLIP tokens results in a faster convergence and better DCS score, which could be due to consistency with the text-to-image initialization model. The proposed EC method outperforms CFG in both cases in terms of DCS and DVS metrics while being Х3 faster.

[0176] Figure 6 shows images generated using the explicit conditioning of the present techniques. Figure 6 compares qualitatively the present single-pass explicit conditioning model with CFG.

[0177] In an embodiment of the disclosure, the first images 612, 614, 616, and 618 may include an input image 612 and edited results generated based on the instruction prompt, "Make the character wear a yellow dress." For example, the CFG images 614 and 616 may correspond to results obtained using CFG with 1 denoising pass and 3 denoising passes, respectively. Further, the EC image 618 may correspond to a result obtained using EC with 1 denoising pass.

[0178] In an embodiment of the disclosure, the second images 622, 624, 626, and 628 may include an input image 622 and edited results generated based on the instruction prompt, "Make the canyon look like a giant cave." For example, the CFG images 624 and 626 may correspond to results obtained using CFG with 1 denoising pass and 3 denoising passes, respectively. Further, the EC image 628 may correspond to a result obtained using EC with 1 denoising pass.

[0179] In an embodiment of the disclosure, the third images 632, 634, 636, and 638 may include an input image 632 and edited results generated based on the instruction prompt, "Make it a winter scene." For example, the CFG images 634 and 636 may correspond to results obtained using CFG with 1 denoising pass and 3 denoising passes, respectively. Further, the EC image 638 may correspond to a result obtained using EC with 1 denoising pass.

[0180] In an embodiment of the disclosure, the fourth images 642, 644, 646, and 648 may include an input image 642 and edited results generated based on the instruction prompt, "Make it a hotel." For example, the CFG images 644 and 646 may correspond to results obtained using CFG with 1 denoising pass and 3 denoising passes, respectively. Further, the EC image 648 may correspond to a result obtained using EC with 1 denoising pass.

[0181] Two cases for CFG (eg., CFG images 614, 616, 624, 626, 634, 636, 644, 646) are presented: with pass, i.e., in Eq. 11 which is computationally similar to EC (eg., EC images 618, 628, 638, 648), and with passes, i.e., . It can be seen from the results in Figure 6 that the quantitative results in Figure 5 are supported, where the proposed explicit conditioning significantly outperforms CFG while being faster. Although the present techniques are more efficient and robust compared with CFG, however, the impact of the input context can be only controlled by the instruction prompt. The guidance scale in CFG can be considered as a tool to control the impact of the context image. While this can be seen as an advantage, it makes CFG ill-conditioned to find the optimal scale and computationally inefficient. As shown in the experiments, explicit conditioning is superior in following the given instructions for image editing, as a result, the proposed method benefits more significantly from a higher-quality dataset where more diverse and detailed instructions are provided.

[0182] Figure 7 is a block diagram of a system for training a diffusion model to performed image editing and for using the trained diffusion model.

[0183] In an embodiment of the disclosure, the system may include at least one of a server 1000 for training a diffusion model 1008 into one that can perform explicit conditioning, using the methods described above (e.g. with reference to Figure 1A). Meanwhile, without being limited to the disclosed example, training and inference of the diffusion model may also be performed by the electronic device.

[0184] Referring to Figure 7, a server 1000 according to an embodiment may include a processor 1002 , a memory 1004, a training dataset 1006, a diffusion model 1008. However, the components of the server 1000 may not be limited to the example described above, and the server 1000 may include more components or fewer components than the components described above. In an embodiment of the disclosure, some or all of the processor 1002 and the memory 1004 may be implemented in the form of one chip, and the processor 1002 may include one or more processors.

[0185] The processor 1002 may be configured to control a series of processes for operating the server 1000 according to the embodiments described herein. The processor 1002 may include one or a plurality of processors. In this state, the one or a plurality of processors may include general purpose processors such as a central processing unit (CPU), an application processor (AP), a digital signal processor (DSP), and the like, graphics-dedicated processors such as a graphics processing unit (GPU), a vision processing unit (VPU), and the like, or artificial intelligence-dedicated processors such as a neural processing unit (NPU). For example, when the processor 1002 includes an artificial intelligence-dedicated processor, the artificial intelligence-dedicated processor may be designed in a hardware structure specified for processing a particular artificial intelligence model.

[0186] The memory 1004, which is a configuration for storing various programs or data, may include a storage medium, such as ROM, RAM, a hard disk, a CD-ROM, a DVD, and the like, or a combination of such storage media. In some embodiments, the memory 1004 may be included in the processor 1002 rather than existing separately. The memory 1004 may include a volatile memory, a non-volatile memory, or a combination thereof. The memory 1004 may store programs or instructions that, when executed by the processor 1002, cause the processor 1002 to perform operations according to the embodiments described herein (e.g., explicit conditioning training of a diffusion model). The memory 1004 may provide data stored therein to the processor 1002 according to a request of the processor 1002. And the memory 1004 may include one or more memory.

[0187] In an embodiment, the server 1000 may store, in the memory 1004, a training dataset 1006 and a diffusion model 1008 to be trained. The training dataset 1006 may include training data (e.g., triplets including an input image, a prompt, and a target output image) for training the diffusion model 1008 to perform image editing using explicit conditioning. The processor 1002 may train the diffusion model 1008 using the training dataset 1006 by executing instructions stored in the memory 1004, and may write updated model parameters and / or related training results to the memory 1004. The training uses a training dataset 1006.

[0188] The processor 1002 may write data to the memory 1004, read data stored in the memory 1004, and process data according to predefined operation rules or an artificial intelligence model by executing the programs or instructions stored in the memory 1004. Accordingly, the processor 1002 may perform the operations described in the embodiments described herein, and it can be said that the operations performed by the server 1000 in the embodiments described herein are performed by the processor 1002 unless otherwise stated.

[0189] The server 1000 may comprise at least one processor 1002 coupled to memory 1004 and arranged for implementing the methods described herein. The memory 1004 may store instructions that, when executed by the at least one processor 1002 individually or collectively, cause the at least one processor to perform the methods described herein.

[0190] Referring to Figure 7, the electronic device 2000 may be able to use a trained diffusion model. The user device may therefore comprise the trained diffusion model 2008 that has been fine-tuned by the server 1000.

[0191] In an embodiment of the disclosure, the electronic device 2000 according to an embodiment may include a processor 2002, a memory 2004, a storage 2006, and a display 2010. However, the components of the electronic device 2000 may not be limited to the example described above, and the electronic device 2000 may include more components or fewer components than the components described above. The electronic device 2000 may also be referred to as an user device. In an embodiment of the disclosure, some or all of the processor 2002 and the memory 2004 may be implemented in the form of one chip, and the processor 2002 may include one or more processors.

[0192] The memory 2004, which is a configuration for storing various programs or data, may include a storage medium, such as ROM, RAM, a hard disk, a CD-ROM, a DVD, and the like, or a combination of such storage media. In some embodiments, the memory 2004 may be included in the processor 2002 rather than existing separately. The memory 2004 may include a volatile memory, a non-volatile memory, or a combination thereof. The memory 2004 may store programs or instructions that, when executed by the processor 2002, cause the processor 2002 to perform operations according to the embodiments described herein (e.g., performing image editing using a trained diffusion model). The memory 2004 may provide data stored therein to the processor 2002 according to a request of the processor 2002.

[0193] The processor 2002 may be configured to control a series of processes for operating the electronic device 2000 according to the embodiments described herein. The processor 2002 may include one or a plurality of processors, including, for example, a CPU, an AP, a DSP, a GPU, a VPU, and / or an NPU.

[0194] In an embodiment, the electronic device 2000 may include, store, or otherwise maintain a trained diffusion model 2008. The trained diffusion model 2008 may be a model that has been fine-tuned by the server 1000 (e.g., from the diffusion model 1008) to support explicit conditioning for image editing. The processor 2002 may execute the trained diffusion model 2008 to generate edited images based on an input image and an instruction prompt, as described herein.

[0195] The storage 2006 may store input images to be edited and / or resulting generated edited images. The display 2010 may display, for example, generated edited images and / or related user interface elements.

[0196] The processor 2002 may write data to the memory 2004 and / or the storage 2006, read data stored in the memory 2004 and / or the storage 2006, and process data according to predefined operation rules or an artificial intelligence model by executing the programs or instructions stored therein. Accordingly, the processor 2002 may perform the operations described in the embodiments described herein, and it can be said that the operations performed by the electronic device 2000 in the embodiments described herein are performed by the processor 2002 unless otherwise stated. Meanwhile, the memory 2004 and the storage 2006 are described as being separate for convenience of explanation, but may, in practice, be the same.

[0197] The display 2010 may include an input interface (e.g., a touch screen, a hard button, etc.) for receiving a control command, information, and the like from a user, and an output interface for displaying a result of execution of an operation according to the user's control or a state of the electronic device 2000.

[0198] As explained above, the electronic device 2000 may comprise at least one processor 2002 coupled to memory 2004 and arranged for implementing the methods described herein (such as those described above with reference to Figure 1B). The electronic device 2000 may comprise storage 2006 for storing the input images to be edited and / or the resulting generated edited images, and a display 2100 for displaying the generated edited images.

[0199] As related art to the present application, the following documents may have been referenced. (i) Brooks et al - Brooks, T., Holynski, A., and Efros, A. A. Instructpix2pix: Learning to follow image editing instructions. arXiv preprint arXiv:2211.09800, 2022, (ii) Classifier Free Guidance (CFG) - J. Ho and T. Salimans, "Classifier-free diffusion guidance," in Arxiv, 2022 & Q. Zheng, M. Le, N. Shaul, Y. Lipman, A. Grover, and R. T. Q. Chen, "Guided flows for 279 generative modeling and decision making," in Arxiv, 2023, (iii) D Ahn et al - D. Ahn, J. Kang, S. Lee, J. Min, M. Kim, W. Jang, H. Cho, S. P. S. Kim, E. Cha, K. H. Jin, and 322 S. Kim, "A noise is worth diffusion guidance," 2024, (iv) Stable Diffusion - R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, "High-resolution image synthesis with latent diffusion models," in CVPR, 2022, (v) PixelShuffle - W. Shi, J. Caballero, F. Huszar, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang, "Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network," 2016.

[0200] In an embodiment of the disclosure, a method for training a diffusion model to edit images, the method may include obtaining a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt to edit the input image, and a target output image that is an edited version of the input image according to the text prompt. The method may include generating, for each triplet of training data: an input image embedding for the input image, using an image encoder of the diffusion model, a prompt embedding for the text prompt, using a text encoder of the diffusion model, and an initial noise image for the diffusion model to use to generate the output image of the triplet. The initial noise image may be conditioned on the input image embedding and the prompt embedding. The method may include training the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by denoising, for each triplet, the generated initial noise image to generate an edited image, using a denoising network of the diffusion model; and by updating one or more parameters of the diffusion model to improve alignment between the generated edited image and the output image of the triplet.

[0201] In an embodiment of the disclosure, the method may include generating an output image embedding for the generated edited image. The method may include generating, using the image encoder, a target image embedding for the target output image. The method may include generating a first text caption for the input image and a second text caption for the target output image, wherein the first text caption and the second text caption describe the images. The method may include generating, using the text encoder, a first text embedding for the first text caption, and a second text embedding for the second text caption.

[0202] In an embodiment of the disclosure, the method may include calculating a first difference between the output image embedding for the generated edited image and the input image embedding for the input image.

[0203] In an embodiment of the disclosure, the method may include calculating a first alignment measure between the generated edited image and the target output image of the triplet by calculating a second difference between the first text embedding corresponding to the input image and the second text embedding corresponding to the target output image; and by calculating a cosine similarity between the first difference and the second difference.

[0204] In an embodiment of the disclosure, the method may include calculating a second alignment measure between the generated edited image and the output image of the triplet by calculating a third difference between the output embedding for the generated edited image and the target image embedding generated for the target output image, and by calculating a cosine similarity between the first difference and the third difference.

[0205] In an embodiment of the disclosure, the method may include updating one or more parameters of a neural network of the diffusion model. The neural network may be referred to as a denoising network.

[0206] In an embodiment of the disclosure, the method may include obtaining a sample from a Gaussian distribution having a mean and variance that is parameterized by the input image embedding and the prompt embedding. The method may include generating an initial noise image using the obtained sample.

[0207] In an embodiment of the disclosure, the method may include generating, using the input image embedding for the input image, a mean and variance for the input image. The method may include generating, using the prompt embedding for the text prompt, a mean and variance for the text prompt.

[0208] In an embodiment of the disclosure, the method may include summing the generated mean for the input image and the generated mean for the text prompt to generate a total mean. The method may include summing the generated variance for the input image and the generated variance for the text prompt to generate a total variance. The method may include using the total mean and total variance as the mean and variance of the Gaussian distribution.

[0209] In an embodiment of the disclosure, a server for training a diffusion model to edit images. The server may include at least one processor coupled to memory. The at least one processor may obtain a training dataset comprising triplets of training data, each triplet comprising an input image, a text prompt to edit the input image, and a target output image that is an edited version of the input image according to the text prompt. The at least one processor may generating, for each triplet of training data: an input image embedding for the input image, using an image encoder of the diffusion model, a prompt embedding for the text prompt, using a text encoder of the diffusion model, and an initial noise image for the diffusion model to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding. The at least one processor may training the diffusion model to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by denoising, for each triplet, the generated initial noise image to generate an edited image, using a denoising network of the diffusion model, and by updating one or more parameters of the diffusion model to improve alignment between the generated edited image and the output image of the triplet.

[0210] In an embodiment of the disclosure, a method for using a trained diffusion model to edit images, the method may include obtaining an input image and a text prompt to edit the input image. The method may include generating an input image embedding for the input image using an image encoder of the trained diffusion model, a prompt embedding for the text prompt using a text encoder of the trained diffusion model, and an initial noise image for the trained diffusion model to use to generate edited images, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding. The method may include denoising, using a denoising network of the trained diffusion model, the generated initial noise image to generate an edited image.

[0211] In an embodiment of the disclosure, the method may include obtaining a sample from a Gaussian distribution having a mean and variance that is parameterized by the input image embedding and the prompt embedding. The method may include generating an initial noise image using the obtained sample.

[0212] In an embodiment of the disclosure, the method may include generating, using the input image embedding for the input image, a mean and variance for the input image. The method may include generating, using the prompt embedding for the text prompt, a mean and variance for the text prompt.

[0213] In an embodiment of the disclosure, the method may include summing the generated mean for the input image and the generated mean for the text prompt to generate a total mean. The method may include summing the generated variance for the input image and the generated variance for the text prompt to generate a total variance. The method may include using the total mean and total variance as the mean and variance of the Gaussian distribution.

[0214] In an embodiment of the disclosure, the method may include displaying the generated edited image.

[0215] Those skilled in the art will appreciate that while the foregoing has described what is considered to be the best mode and where appropriate other modes of performing present techniques, the present techniques should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognise that present techniques have a broad range of applications, and that the embodiments may take a wide range of modifications without departing from any inventive concept as defined in the appended claims.

Claims

1.A method for training a diffusion model (300) to edit images, the method comprising:obtaining (S100) a training dataset comprising triplets of training data, each triplet comprising an input image (320), a text prompt (310) to edit the input image (320), and a target output image (330) that is an edited version of the input image (320) according to the text prompt (310);generating (S102, S104), for each triplet of training data:an input image embedding for the input image (320), using an image encoder (322) of the diffusion model (300),a prompt embedding for the text prompt (310), using a text encoder of the diffusion model (300), andan initial noise image for the diffusion model (300) to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding; andtraining (S106, S108) the diffusion model (300) to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by:denoising, for each triplet, the generated initial noise image to generate an edited image, using a denoising network (350) of the diffusion model (300); andupdating one or more parameters of the diffusion model (300) to improve alignment between the generated edited image and the output image of the triplet.2.The method of claim 1, the method further comprises:generating an output image embedding for the generated edited image;generating, using the image encoder (332), a target image embedding for the target output image (330);generating a first text caption for the input image (320) and a second text caption for the target output image (330), wherein the first text caption and the second text caption describe the input image (320) and the target output image (330), respectively; andgenerating, using the text encoder, a first text embedding for the first text caption, and a second text embedding for the second text caption.3.The method of claim 2, wherein the training the diffusion model further comprises:calculating a first difference between the output image embedding for the generated edited image and the input image embedding for the input image (320).4.The method of claim 3, wherein the updating one or more parameters of the diffusion model comprises:calculating a first alignment measure between the generated edited image and the target output image (330) of the triplet by:calculating a second difference between the first text embedding corresponding to the input image (320) and the second text embedding corresponding to the target output image (330); andcalculating a cosine similarity between the first difference and the second difference.5.The method of claim 3 or 4, wherein the updating one or more parameters of the diffusion model comprises:calculating a second alignment measure between the generated edited image and the output image of the triplet by:calculating a third difference between the output embedding for the generated edited image and the target image embedding generated for the target output image (330); andcalculating a cosine similarity between the first difference and the third difference.6.The method of any one of claims 1 to 5, wherein the updating one or more parameters of the diffusion model comprises:updating one or more parameters of a denoising network (350) of the diffusion model (300).7.The method of any one of claims 1 to 6, wherein the generating the initial noise image comprises:obtaining a sample from a Gaussian distribution (340) having a mean and variance that is parameterized by the input image embedding and the prompt embedding, andgenerating an initial noise image using the obtained sample.8.The method of claim 7, the method further comprises:generating, using the input image embedding for the input image, a mean and variance for the input image; andgenerating, using the prompt embedding for the text prompt, a mean and variance for the text prompt.9.The method of claim 8, wherein the obtaining the sample from the Gaussian distribution comprises:summing the generated mean for the input image (320) and the generated mean for the text prompt (310) to generate a total mean;summing the generated variance for the input image (320) and the generated variance for the text prompt (310) to generate a total variance; andusing the total mean and total variance as the mean and variance of the Gaussian distribution (340).10.A server (1000) for training a diffusion model (300, 1008) to edit images, the server (1000) comprising:at least one processor coupled to memory for:obtaining (S100) a training dataset comprising triplets of training data, each triplet comprising an input image (320), a text prompt (310) to edit the input image (320), and a target output image (330)that is an edited version of the input image (320) according to the text prompt (310);generating (S102, S104), for each triplet of training data:an input image embedding for the input image (320), using an image encoder (322) of the diffusion model (300),a prompt embedding for the text prompt (310), using a text encoder of the diffusion model(300), andan initial noise image for the diffusion model (300) to use to generate the output image of the triplet, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding; andtraining (S106, S108) the diffusion model (300) to generate edited images using the triplets of training data and the generated initial noise image for each triplet, by:denoising, for each triplet, the generated initial noise image to generate an edited image, using a denoising network (350) of the diffusion model (300); andupdating one or more parameters of the diffusion model (300)to improve alignment between the generated edited image and the output image of the triplet.11.A method for using a trained diffusion model (400, 2008) to edit images, the method comprising:obtaining (S200) an input image (420) and a text prompt (410) to edit the input image (420);generating (S202, S204):an input image embedding for the input image (420) using an image encoder (422) of the trained diffusion model (400),a prompt embedding for the text prompt (410) using a text encoder of the trained diffusion model (400), andan initial noise image for the trained diffusion model (400) to use to generate edited images, wherein the initial noise image is conditioned on the input image embedding and the prompt embedding; anddenoising, using a denoising network (450) of the trained diffusion model (400), the generated initial noise image to generate an edited image.12.The method of claim 11, wherein the generating an initial noise image comprises:obtaining a sample from a Gaussian distribution (440) having a mean and variance that is parameterized by the input image embedding and the prompt embedding, andgenerating an initial noise image using the obtained sample.13.The method of claim 12, wherein the obtaining the sample from the Gaussian distribution comprises:generating, using the input image embedding for the input image (420), a mean and variance for the input image (420); andgenerating, using the prompt embedding for the text prompt (410), a mean and variance for the text prompt (410).14.The method of claim 13, wherein the obtaining the sample from the Gaussian distribution further comprises:summing the generated mean for the input image (420) and the generated mean for the text prompt (410) to generate a total mean;summing the generated variance for the input image (420) and the generated variance for the text prompt (410) to generate a total variance; andusing the total mean and total variance as the mean and variance of the Gaussian distribution (440).15.The method of any one of claims 11 to 14, the method further comprising:displaying the generated edited image.