Optimization-based image editing

By introducing an encoding-decoding process and a dual loss function into the image generative diffusion model, the controllability and cost issues of image editing in existing technologies are solved, enabling flexible and efficient image editing.

CN121844356APending Publication Date: 2026-04-10HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2023-09-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing generative diffusion models for images struggle to achieve precise controllability in image editing, especially for local modifications, and existing methods are typically costly or lack generalization ability.

Method used

An encoding-decoding process is employed, combining a first loss function and a second loss function to preserve the details of the input image and to modify it according to image editing instructions, respectively. Image editing is performed using a pre-trained generative diffusion model.

Benefits of technology

It achieves highly controllable image editing, accurately preserving areas that do not need editing, while allowing flexible modifications through text, doodles, and pose commands, reducing training costs and improving editing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844356A_ABST
    Figure CN121844356A_ABST
Patent Text Reader

Abstract

An image processing apparatus (800) is described, the apparatus comprising one or more processors (804) configured to: receive (701) image editing instructions (207, 208), where the image editing instructions (207, 208) indicate a desired modification to an input image (201); inputting (702) the input image to a coding-decoding process to form an edited image (213), where the coding-decoding process comprises: (i) inputting an intermediate image (205) obtained from the input image, (ii) inputting the image editing instructions (207, 208) to a pre-trained generative diffusion model, and (ii) inputting the image editing instructions (207, 208) to the pre-trained generative diffusion model to form an edited image (213). The image editing instruction is configured to generate the edited image (213) from respective outputs of a first loss function for preserving details in one or more portions of the input image and a second loss function for directing the input image to be modified in accordance with the image editing instruction. In this way, the pre-trained generative model can be endowed with highly controllable editing ability, an image area (such as a background area) which does not need to be edited can be accurately reserved, and controllability can be achieved through various instructions such as texts, graffiti and postures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to image processing, and more particularly to image editing using a generative diffusion model in conjunction with image editing instructions. Background Technology

[0002] Large-scale text-to-image generative diffusion models have revolutionized image generation capabilities. Fundamental models such as stable diffusion have achieved remarkable results in composition, image quality, scene diversity, and text controllability. While these models can generate complex images based on individual text conditions, the precise controllability of image structure, composition, and appearance remains limited. Achieving desired results typically requires time-consuming and tedious cueing engineering.

[0003] A natural approach to improving controllability is to give generative models the ability to edit images. Compared to iterative cue updates that cannot guarantee the preservation of the original image's appearance, editing allows for more precise iteration of the image's appearance by locally modifying image regions, while retaining the desired aesthetic effect.

[0004] Image editing using generative diffusion models can be approached from two angles: (1) fine-tuning the diffusion model; and (2) regulating the generative diffusion process on a frozen pre-trained model.

[0005] The former technique (e.g., "Instructpix2pix: Learning to follow image editing instructions" by Brooks, T., Holynski, A., and Efros, AA, published in the IEEE / CVF Conference Proceedings on Computer Vision and Pattern Recognition, pp. 18392-18402, 2023, and "Sine: Singleimage editing with text-to-image diffusion models" by Zhang, Z., Han, L., Ghosh, A., Metaxas, DN, and Ren, J., published in the IEEE / CVF Conference Proceedings on Computer Vision and Pattern Recognition, pp. 6027-6037, 2023) often achieves better results, but at the cost of high training costs, significant effort required for data acquisition, and reduced generalization ability. This model is specifically designed for editing and has a reduced ability to generate standard images. Furthermore, some training-based methods (such as the method described in "Imagic: Text-based real image editing with diffusion models" by Kawar, B. et al., published in the IEEE / CVF Conference Proceedings on Computer Vision and Pattern Recognition, pp. 6007-6017, 2023) require fine-tuning for each image, a process that can be extremely costly.

[0006] The latter technique focuses on regulating the generative diffusion process by introducing specific constraints. A well-known strategy, as described by Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, JY, and Ermon, S. in their 2021 paper "Sdedit: Guided image synthesis and editing with stochastic differential equations" (arXiv preprint, arXiv:2108.01073), is commonly referred to as image-to-image. Figure 1 As shown, noise is iteratively applied to the image 101 to be edited at time steps 1, t–1, t, and t, respectively. ENoisy images 102, 103, 104, and 105 are generated. Image 105 is then denoised using a diffusion model and an editing condition 107, which in this example is pose modification. This denoising process is implemented using a pre-trained diffusion model in the following manner: based on the current noisy images 111, 110, 109, and 108, the text prompt 106, and the editing condition 107, at multiple time steps 1, t–1, t, and t… E The noise vector to be removed is predicted at the specified location. The result is the edited image 112.

[0007] These methods offer significant practical advantages because they require no training or optimization. However, they struggle to accurately incorporate the required edits into the final image.

[0008] Therefore, there is a need to develop an image editing method that can at least overcome some of the problems mentioned above. Summary of the Invention

[0009] The present invention provides an image processing apparatus comprising one or more processors, the processors being configured to: receive an image editing instruction, wherein the image editing instruction indicates a desired modification to an input image; input the input image into an encoding-decoding process to form an edited image, wherein the encoding-decoding process includes: (i) inputting an intermediate image obtained based on the input image, and (ii) inputting the image editing instruction into a pre-trained generative diffusion model to form an edited image based on the corresponding outputs of a first loss function and a second loss function, the first loss function being used to preserve details in one or more portions of the input image, and the second loss function being used to guide the input image to be modified according to the image editing instruction.

[0010] This can give pre-trained generative models highly controllable editing capabilities, accurately preserving image regions that do not need editing (e.g., background regions), while also enabling controllability through various instructions such as text, doodles, and poses.

[0011] The one or more processors may be used to form a guide image that depicts the desired modified guide image in one or more local regions of the input image. The guide image may be formed by the one or more processors under more relaxed preservation constraints.

[0012] The one or more processors can be used to form the guide image from the input image based on the image editing instructions. This allows the guide image to guide modifications to the input image according to the image editing instructions. The guide image can serve as a reference image in the second loss function. The guide image can be used to guide the second loss function.

[0013] The image editing instructions may include text describing the content of the input image and / or the desired modifications to the input image. This textual prompt can be used to guide the modification of the input image.

[0014] The encoding-decoding process may include an encoding stage and a decoding stage. The encoding stage converts the input image into a noise vector, encoding the input image into a noisy intermediate image. During the encoding stage, the latent image features of the input image are encoded and fixed. The decoding stage decodes the noisy intermediate image. During the decoding stage, the latent image features of the intermediate image can be updated. This update can be performed at multiple time steps.

[0015] The encoding stage of the encoding-decoding process may include progressively converting the input image into a noise vector to form the intermediate image. The intermediate image may be a noisy image formed during the encoding stage. This allows the input image to be transformed into an edited image through a generative diffusion model.

[0016] The decoding stage of the encoding-decoding process may include inputting the intermediate image into the pre-trained generative diffusion model to denoise the intermediate image according to the image editing instructions, so as to form the edited image.

[0017] The input image can be a latent representation of the original image to be edited. The input image (in the latent space) can be formed by transforming the original image (in the image space). For example, the input image can be a latent representation of the original image to be edited obtained using a pre-trained Variational Autoencoder (VAE) model.

[0018] The input image may include latent image features. These latent image features can be updated based on the corresponding outputs of the first loss function and the second loss function. The latent image features can be updated during the encoding-decoding process. The features may remain fixed during the encoding phase but be updated during the decoding phase. The latent image features of the input image may be encoded and fixed, and used to guide the first loss function to preserve details in one or more portions of the input image. The decoding phase can update the image features in the latent space. This allows an edited image to be formed in the latent space. Subsequently, by transforming the edited image in the latent space, for example using a VAE decoder model, the edited image in the latent space can be converted to an image space.

[0019] During the encoding-decoding process, the latent image features can be updated at multiple time steps. Specifically, the latent image features can be updated at multiple time steps during the decoding phase. This can improve the accuracy of the process.

[0020] The first loss function can calculate the distance between the latent image features of the input image and the corresponding latent image features of the edited image. This can be performed at each of the multiple time steps in the process. This allows details in one or more portions of the image to be preserved with pixel-level consistency. For example, the first loss function can calculate the mean squared error between image features, maintaining consistency between the input image and the edited image. This can help ensure that original details are preserved in one or more portions of the image.

[0021] The second loss function can calculate the distance between the latent image features of the edited image and the corresponding latent image features of the guide image. For example, the second loss function can use a cosine similarity loss between the edited image and the guide image. This allows for a larger difference between the edited image and the guide image.

[0022] The image editing instructions may include the pose or doodles of the subject in the input image, or the edge map of the input image. This allows the process to be controlled by using a variety of instruction types.

[0023] The region to be modified in the input image can be defined by a mask. The mask can be integrated with the first loss function so that the first loss function is not applied to the region in the intermediate image corresponding to the region to be modified in the input image. This allows the first loss function to be calculated only for the image region in the input image that needs to be retained.

[0024] The effects of the first and second loss functions on the formation of the edited image can be configured by the user through adjustable parameters. This makes the process controllable, allowing for the discarding of guiding image components by adjusting parameters for simpler editing instructions, and the editing of the image through simple retention loss.

[0025] According to another aspect, this application provides an image processing method, the method comprising: receiving an image editing instruction, wherein the image editing instruction indicates a desired modification of an input image; inputting the input image into an encoding-decoding process to form an edited image, wherein the encoding-decoding process comprises: (i) inputting an intermediate image obtained based on the input image, and (ii) inputting the image editing instruction into a pre-trained generative diffusion model to form an edited image based on the corresponding outputs of a first loss function and a second loss function, the first loss function being used to preserve details in one or more portions of the input image, and the second loss function being used to guide the input image to be modified according to the image editing instruction.

[0026] This method can endow pre-trained generative models with highly controllable editing capabilities, accurately preserving image regions that do not need editing (such as background regions), and achieving controllability through the use of various instructions such as text, doodles, and poses.

[0027] According to another aspect, this application provides a computer program that, when executed by a computing device, causes the computing device to perform the above-described method.

[0028] In another aspect, this application provides a data carrier for storing the aforementioned computer program in a non-transient form. Attached Figure Description

[0029] This application will now be described by way of example with reference to the accompanying drawings.

[0030] In the attached diagram:

[0031] Figure 1 This illustration illustrates an existing image-to-image editing method.

[0032] Figure 2 This illustration demonstrates a decoupled inference time optimization method.

[0033] Figure 3 An exemplary flow of the method described herein is illustrated schematically.

[0034] Figure 4 An exemplary image editing result is shown, wherein the image editing instructions include the target pose of the image subject.

[0035] Figure 5 An example image editing result is shown, where the image editing instructions include doodles.

[0036] Figure 6 The illustration shows different values ​​of the adjustable parameter λ.

[0037] Figure 7 An example of an image processing method according to an embodiment of this application is shown.

[0038] Figure 8 A schematic diagram of an apparatus for performing the methods described herein and some of its associated components is shown. Detailed Implementation

[0039] The image editing method described in this paper employs inference time optimization (ITO), using a generative diffusion model to generate the edited image based on the image editing instructions. This method is essentially designed for both text-based and structural editing instructions (e.g., pose or doodle modifications).

[0040] The editing task is decomposed into two mutually constraining subtasks. The first subtask is to preserve the original image details in one or more parts of the input image. The second subtask is to modify one or more local image regions (wherein, the local image region to be modified is a part of the input image that differs from one or more parts of the original image where the preserved details are located). The editing task can be decoupled into the preservation and modification tasks by two separate optimization losses. The former is achieved by maintaining consistency between the latent features of the input image and the latent features of the edited image, while the latter utilizes an intermediate edit output (guide image) with looser preservation constraints to guide the editing process of the modified region.

[0041] In this process, latent image features are updated according to two loss functions. The first loss function helps ensure consistency (preservation) between the input image to be edited and the edited image, while the second loss function guides the editing process toward the desired aesthetic effect (modification) by generating intermediate outputs in the form of guide images. The guide image is an image that accurately depicts the appearance of the modifications to be made in one or more local regions of the image. This semantic image editing method can utilize layout control modules to allow editing using various types of editing instructions, including text, poses, doodles, or more complex image layouts.

[0042] This editing process can be advantageously adapted to the difficulty of the task by discarding or refining the guide image.

[0043] Given an input image I and semantic editing conditions C, the goal is to locally modify I according to C while accurately preserving the unaffected image content. Unlike existing techniques that typically restrict editing condition input to text, C can be provided as text and / or image spatial layout input.

[0044] See Figure 1The described image-to-image editing concept is built on a noise vector, and the input image can be gradually transformed into a noise vector, and then denoised at multiple time steps using the prediction results of a pre-trained latent diffusion model according to the editing condition C. In this encoding-decoding process, an optimization process is performed on the latent image features using the two different loss functions described above.

[0045] Figure 2 An overview of the method is schematically shown. For simplicity, Figure 2 the process in the image space is shown. But the process actually takes place in the latent space. 201 shows the input image. In this example, the input image 201 is a latent image including latent image features. The latent image features of image 201 can be formed by transforming the original image to be modified. In this example, the input image = fφ(I) is the latent representation of image I. This latent representation can be obtained from image I through a pre-trained variational autoencoder (VAE) model.

[0046] The denoising diffusion model learns to reverse a multi-step diffusion process in which the input image 201 is gradually transformed into a Gaussian noise vector by iteratively adding noise N(0, I) at multiple time steps. . Figure 2 Shows that the input image 201 is transformed into intermediate noisy images at intermediate time steps 1, t–1, t, and t E <T respectively, , , and (shown as 202, 203, 204, and 205 respectively).

[0047] In this example, it is assumed that one or more processors of the image processing device can access the text prompt or description S207. This text prompt 207 can be written manually or generated by an existing image captioning model. One or more processors of the image processing device can also pass through a pre-trained text-to-image diffusion model that can utilize additional layout conditions 208 (such as pose or scribble), for example, using the ControlNet module (as described in "Adding conditional control to text-to-image diffusion models" by Zhang, L. and Agrawala, M. (arXiv preprint, arXiv:2302.05543) published in 2023).

[0048] During the encoding stage, the input image shown in 201... The image is gradually converted into a noise vector, and then, during the decoding stage, the intermediate noisy image is denoised according to the new editing condition 208 (pose modification in this example). This denoising process is implemented using a pre-trained diffusion model in the following way: based on the current noisy image (for time step t)... E t, t-1, and 1 are the images shown in figures 209, 210, 211, and 212, respectively. , , and One of the methods is to predict the noise vector to be removed at multiple time steps, using text prompts 207 and editing conditions 208. To provide the ability to edit images based on conditions other than text, layout conditioning modules, such as ControlNet described in "Adding conditional control to text-to-image diffusion models" (arXiv preprint, arXiv:2302.05543) by Zhang, L. and Agrawala, M., published in 2023, can be integrated.

[0049] Figure 2 An intermediate image of the encoding stage is shown. 205 (along with image editing instructions) as an image 209 (i.e., in) Figure 2 middle, = The image is input to the decoding stage and denoised at multiple time steps. (Intermediate image) Denoising is performed at time steps t, t–1, and 1 to obtain the image. , and (As shown in 210, 211 and 212 respectively) and output the edited image at 213. .

[0050] As described above, an optimized process is employed in this encoding-decoding process. The editing task is decoupled into two mutually restraining subtasks through different content preservation losses and local image modification losses.

[0051] The first loss is "retention loss" ( Figure 2 In ).exist Figure 2 In the implementation shown, the loss is calculated for the input image at each time step (i.e., time steps 1, t–1, t, and t). EThe potential image features of the images 202, 203, 204, and 205 at [location] and the edited images at each time step (i.e., at time steps 1, t–1, t, and t E The mean square error (MSE) between the potential image features of the images 212, 211, 210, and 209 at [location]. Update the potential image features of the edited images at each time step. The potential image features of the input image are encoded and remain fixed to determine the first loss to preserve details of one or more parts (such as the background) of the input image. This loss is used to maintain consistency between the input image and the edited image. This helps ensure that the original details are retained. Other loss functions can also be used to calculate the distance between the potential image features of the input image and the corresponding potential image features of the edited image.

[0052] The second loss is the "guidance loss" ( Figure 2 in ). This loss utilizes an intermediate edited output (guidance image) that accurately depicts the desired modified appearance. This image is used to guide the editing process within the area to be edited towards the desired appearance. Different from the retention loss that requires pixel-level consistency in the preferred implementation, the distance between the edited image and the guidance image, such as the cosine similarity loss, can be adopted, resulting in a significant difference between the edited image and the guidance image. Other loss functions can also be used as the second loss function to calculate the distance between the potential image features of the edited image and the corresponding potential image features of the guidance images at each time step (i.e., at time steps 1, t–1, t, and t E of the guidance images 215, 216, 217, and 218 at [location]).

[0053] This optimization process is performed for the first N groups of denoising time steps (i.e., Figure 2 the time steps 1, t–1, t, and t shown in E <T).

[0054] The image modification loss provides visual guidance towards the expected image appearance within the editing area through the intermediate edited images obtained under the constraint of reducing the retention loss. The retention loss can be further enhanced by a mask 206 that delimits the local editing area, thus achieving a balance between content retention and local accurate editing. The use of the mask will be described in detail below.

[0055] Noise can be added to the input image 201 according to a specific variance schedule. Given the text prompt 207, the model (e.g., Unet as described in "U-net: Convolutional networks for biomedical image segmentation" published by Ronneberger et al. in Part III, J8, pp. 234-241 of the Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI) in 2015 (held in Munich, Germany from October 5th to 9th, 2015), Springer) can be trained by predicting the noise vector to be removed at a given time step t to reconstruct the input image. During inference, iteratively estimate the noise vector at T time steps to generate a new image from random noise. Using the deterministic sampling process of the Denoising Diffusion Implicit Model (DDIM) (as described in "Denoising diffusion implicit models" published by Song et al. in 2020 (arXiv preprint, arXiv:2010.02502)), the intermediate denoised image updated at time step t–1 can be estimated as:

[0056] (1)

[0057] where, is the embedding vector of the text prompt condition 207, and α t is a measure of the noise level, depending on the variance scheduler.

[0058] This forward-backward diffusion process can be regarded as an image encoding-decoding process. The forward diffusion process (encoding stage) stops at the intermediate time step t E <T, and then the image is reconstructed starting from t E with the editing condition 208. This strategy usually preserves some of the original image attributes but often leads to unexpected global modifications. With the trained diffusion model , combined with the deterministic reverse DDIM process for encoding the input image (as described in "Diffedit: Diffusion-based semantic image editing with mask guidance" by Couairon, G., Verbeek, J., Schwenk, H., and Cord, M. (arXiv preprint, arXiv:2210.11427) published in 2022), can improve image fidelity. In → direction (encoding process), the following reverse updates are performed:

[0059] (2)

[0060] This process is usually referred to as DDIM reversal and is a direct natural image reversal method (i.e., estimating the noise vector that can generate this exact image).

[0061] The diffusion models described as examples in this paper operate in the latent space (as described in "Resolution image synthesis with latent diffusion models" by Rombach et al. published in the Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 10684 - 10695 in 2022), = fφ(I) is the latent representation of the input image I, obtained through a pre-trained variational autoencoder (VAE) model.

[0062] Input latent image is encoded through DDIM reversal at multiple time steps 1, t – 1, t, and t E <T until reaching the intermediate encoding level t E to obtain the latent intermediate noisy image 205. Taking this latent noisy image as input, the image is decoded by first defining the editing condition 208 to obtain the latent edited image 213. The editing condition 208 can be a modified text prompt or a modified layout input (e.g., a modified pose, scribble, or edge map). Multiple edits can be considered simultaneously (e.g., a modified pose and text prompt as Figure 2 shown), without explicit method modification. In Figure 2 the latent images , and (As shown in 210, 211 and 212 respectively) are the denoised images at time steps t, t-1 and 1 respectively.

[0063] As mentioned earlier, a latent feature optimization process was employed to enhance the consistency of appearance and structure between the input and edited images. At time step t of the backdiffusion process, the latent image features... It will be updated to add with Similarity. Reconstruction task driven. Towards The direction of the update. Theoretically, The update is as follows:

[0064] (3)

[0065] Where the loss function =MSE for and The mean squared error between them; γ is the learning rate. This update is repeated k times with gradient updates, driving... Towards The approach approximates the solution while limiting the number of gradient updates to avoid trivial solutions. = Subsequently, the method continues with the standard diffusion process, As input to the next time step diffusion model.

[0066] The accuracy of the editing process can be further improved using a mask, which describes the region to be edited. This mask can be constructed manually or generated using existing techniques (e.g., as described in "Diffedit: Diffusion-based semantic image editing with maskguidance" by Couairon, G., Verbeek, J., Schwenk, H., and Cord, M, 2022, arXiv preprint arXiv:2210.11427). The mask can be integrated into the first loss, ensuring that the preservation loss only applies to areas outside the editing region.

[0067] Providing a binary mask to identify areas that editing conditions will modify can further facilitate the preservation of original content. For example... Figure 2 The mask m shown in Figure 206 can be integrated into the reconstruction loss, thereby giving greater flexibility to the target editing region of the image: =MSE That is, the reconstruction loss is calculated only in the masked area. Existing methods typically involve calculating using formula (1). After that, using m, via Update the latent image. By introducing mask constraints in the optimization framework, this strategy is more robust to mask quality (e.g., underestimating the editing region) and reduces the risk of introducing artifacts, while allowing non-edited regions (e.g., background regions) to be preserved.

[0068] While the retention loss focuses on faithfully maintaining the image content, the complementary guidance loss aims to enhance editing controllability and achieve precise local modifications by guiding the editing region to the desired appearance. It can be considered that there is an available guidance image G with a latent representation , which can accurately depict the desired editing region. For Perform the same encoding process as to obtain . At time step , update to minimize its cosine distance from the guidance image features:

[0069] , where (4)

[0070] The image at 214 is the original guidance image, while the images at 215, 216, 217, and 218 are the guidance images after adding random noise at time steps 1, t - 1, t, and t E <T respectively.

[0071] Since and are expected to have a large difference (see, is the reconstruction of ), a more conservative cosine distance can be used to replace the reconstruction MSE loss.

[0072] By combining the guidance loss and the retention loss, the fully decoupled ITO process can update the intermediate image features to:

[0073] (5)

[0074] where λ is a hyperparameter used to balance the influence of the editing subtasks. Therefore, the user can use λ as an adjustable parameter, which allows the user to balance between background retention and editing instruction accuracy according to their preference. The above feature update can be performed in the first t u steps of the reverse diffusion process.

[0075] The purpose of a guide image is to accurately describe the desired appearance of the area of ​​the image to be edited. This information cannot be obtained in advance and must be generated through intermediate steps. Encoding the input image with random noise (as described by Meng et al., 2021) facilitates image modification, but at the cost of reduced content preservation; while using DDIM reversal (as described by Couairon et al., 2022) can only provide more conservative modifications. Unlike the final edited output, which must accurately preserve the image background, the focus of guide image generation is on the accurate editing of local image regions (e.g., foreground).

[0076] Based on this observation, a process with inference time optimized using λ=0 (retained only) is used to generate the guiding image. Under this setting, the input image can be encoded using random noise instead of DDIM inversion, allowing for significant modifications to the input image.

[0077] A binary edit mask 206 can be generated using the mask generation method proposed by Couairon et al. (2022). This method involves measuring the difference between noise estimates using text hints and edit hints from the original image as conditions, thereby effectively inferring the regions most affected by different conditions. Furthermore, the mask can be estimated based on layout conditions.

[0078] Using the seed s, for the input Perform random noise coding to obtain Subsequently, based on two different conditions: the original image condition... (e.g., original prompts or gestures) and editing conditions Reverse diffusion → In settings using only layout conditions, images share the same text cue, and conditional modules such as ControlNet integrate layout input into the diffusion process (as described in "Adding conditional control to text-to-image diffusion models" by Zhang and Agrawala, 2023, arXiv preprint arXiv:2302.05543). The binary edit mask is estimated by comparing the noise estimates of the last time step.

[0079] (6)

[0080] m is then converted into a binary mask using a threshold τ. This mask can be averaged over n runs to increase the stability and accuracy of the noise estimation.

[0081] This modular editing program can adapt to the varying difficulty of different editing tasks (e.g., from minor local modifications to more significant image changes).

[0082] Figure 3 An exemplary editing workflow is shown, adaptable to varying task difficulty. 301 shows the input conditions for the input image 302. 303 shows the editing conditions indicating desired modifications to the input image. 304 shows a binary mask.

[0083] like Figure 3 As shown, three different settings can be considered as complexity increases. In the simplest form (i.e., for minor local modifications), the process can include a single generation step, namely mask reconstruction inference time optimization (rITO, λ=0). This setting focuses on preserving the background appearance, as foreground modifications can be easily implemented.

[0084] For simple tasks, the guiding loss can be discarded, and editing can be performed using only the retained loss. Figure 305 shows the edited image produced by retaining the loss.

[0085] Taking posture modification tasks as an example, typical examples of lower difficulty involve small posture adjustments (such as moving an arm), while more difficult examples involve the simultaneous movement of multiple limbs.

[0086] The intermediate task further refines this output 305, using the first output 305 as the guide image in the second step, combining the reconstruction loss and the guide loss (i.e., (r+g)ITO, the standard decoupled setting). Modifying the image in two steps, while balancing guide and reconstruction, provides greater flexibility and allows for larger local modifications.

[0087] Finally, for more difficult tasks (larger editing instructions, such as moving multiple limbs for pose modification), a guide image optimization step can be introduced. The goal is to use the input image as a guide image and, through a guide loss, readjust the appearance of the guide image to align it with the input image. An additional intermediate step optimizes the guide image before (r+g)ITO. This can be achieved by using the original input image as a guide through pure guide ITO (gITO, λ=1). This improves the quality of the guide image and corrects potential artifacts introduced by large editing instructions. Figure 306 shows the edited image for a medium / complex task.

[0088] Figure 4 and Figure 5 The visual comparison under different conditions is shown, where the input image is edited using ControlNet, ControlNet+DiffEdit, and this method.

[0089] exist Figure 4 and Figure 5In the diagram, column (a) shows the original images. Column (b) illustrates the target pose of the edited images. Column (c) shows the edited images formed using ControlNet. The images in column (d) are the edited images formed using ControlNet+DiffEdit (where the mask is calculated based on a layout-conditional update method). This method is shown in columns (e) and (f), where column (e) shows the results with λ=1 (formed solely through the preservation loss, without a guiding image or a second loss).

[0090] Figure 4 The results are shown with different pose conditions as part of image editing instructions. This method achieves correct pose modification while preserving image content. ControlNet successfully generates images with correct poses, but often fails to preserve image content. In contrast, while DiffEdit preserves image content due to its masking process, it often struggles to achieve correct pose changes, a phenomenon particularly noticeable when dealing with more complex instructions. Figure 5 The doodle-based editing results shown exhibit similar performance.

[0091] Figure 6 The effect of the adjustable balancing parameter λ is shown. This parameter controls the relative impact of retention loss and guiding loss.

[0092] Images 601 and 602 show the input image and its corresponding pose, respectively. Image 604 shows the target pose. Image 603 shows the guide image. Image 605 shows the output image. Images 606, 607, 608, 609, 610, and 611 show the edited images generated when λ is 0, 0.2, 0.4, 0.6, 0.8, and 1.0, respectively. When λ=0, the edited image is generated solely based on the guide. When λ=1.0, the edited image is generated solely based on the preserved image. It can be seen that smaller λ values ​​focus more on the foreground and accurate positioning, while larger λ values ​​increase background detail, but excessively high values ​​introduce additional artifacts. The output effect produced by extreme values ​​is either highly similar to the guide image 603 (λ=0), or similar to DiffEdit (λ=1, high background fidelity, but poor pose editing quality). This performance shows that for simpler editing instructions (e.g., ...), ... Figure 4 In the image of a bear or dancer (where the DiffEdit method enables pose modification), the guiding image component can be discarded and the image edited using a simple preservation loss.

[0093] Figure 7An example of an image processing method is shown. At step 701, the method includes receiving an image editing instruction, wherein the image editing instruction indicates a desired modification to an input image. At step 702, the method includes inputting the input image into an encoding-decoding process to form an edited image, wherein the encoding-decoding process includes: (i) inputting an intermediate image obtained based on the input image, and (ii) inputting the image editing instruction into a pre-trained generative diffusion model to form the edited image based on the corresponding outputs of a first loss function and a second loss function, the first loss function being used to preserve details in one or more portions of the input image, and the second loss function being used to guide the modification of the input image according to the image editing instruction.

[0094] Figure 8 An example of an image processing apparatus 800 for implementing the methods described herein is shown. The apparatus includes a device 801. Device 801 includes a processor 802 and a memory 803.

[0095] Device 801 may also include a transceiver 804 for communicating with other entities 810 and 811 via a network. These entities may be physically remote from device 801. The network may be a publicly accessible network such as the Internet. Entities 810 and 811 may be cloud-based. Entity 810 is a computing entity. Entity 811 is a command and control entity. All of these entities are logical entities. In practice, they may be provided by one or more physical devices (such as servers and data storage), and the functionality of two or more entities may be provided by a single physical device. The physical device implementing each entity includes a processor and memory. These devices may also include transceivers for sending data to and receiving data from transceiver 804 of device 801.

[0096] The memory stores code in a non-transient manner, which can be executed by a processor to implement the corresponding entity functions as described herein.

[0097] Therefore, this method can be deployed in various ways, such as cloud deployment, device deployment, or dedicated hardware deployment. As mentioned above, cloud facilities can perform training to develop new algorithms or improve existing ones. Depending on the computing power near the data corpus, training can be performed near the source data or in the cloud, for example, using an inference engine. This method can also be implemented on devices, dedicated hardware, or in the cloud.

[0098] Image editing instructions can include conditions such as text, pose, doodles, and Hough line graphs, allowing for complex editing through multiple conditions.

[0099] This method proposes an editing approach for a frozen-diffusion model that goes beyond text editing and can further leverage image structure conditions for modification. A decoupled inference-time editing strategy separates background preservation from foreground editing, allowing for flexible adjustment of the emphasis on either subtask. The method provides a multi-step adaptive workflow that adjusts the editing process according to task complexity, and offers detailed parameter analysis to aid in intuitive understanding of its operational mechanism.

[0100] Existing methods mostly focus on text-driven editing, requiring complex strategies (especially attention mechanisms) to achieve precise local modifications. This application constructs a method that utilizes explicit layout constraints, achieving precise editing control with only simple constraint optimization. Decoupling saving and modifying further enhances the method's controllability and flexibility.

[0101] Using this method, editing instructions other than text can be used to flexibly and precisely modify image content. Inference-time optimization (discriminative model weight updates) also supports flexible learning from image content without the need for training data. This method is independent of the specific architecture of the underlying latent diffusion model; by separating the save and modify tasks into two different loss functions, users can balance the impact of each task and adjust editing constraints according to their preferences.

[0102] This method can endow pre-trained generative models with highly controllable editing capabilities, accurately preserving image regions that do not need editing (such as background regions), and achieving controllability through the use of various instructions such as text, doodles, and poses.

[0103] The applicant hereby discloses each individual feature described herein, as well as any combination of two or more such features. With ordinary knowledge of those skilled in the art, such features or combinations can be implemented as a whole according to this specification, regardless of whether such features or combinations of features solve any problem disclosed herein; and without limiting the scope of the claims. The applicant notes that various aspects of the invention may include any such individual feature or combination of features. In view of the foregoing description, those skilled in the art will appreciate that various modifications can be made within the scope of the invention.

Claims

1. An image processing apparatus (800), the apparatus comprising one or more processors (801), the one or more processors (801) being configured to: Receive (701) image editing instructions (207, 208), where, The image editing instructions (207, 208) indicate the desired modifications to the input image (201); The input image (201) is input (702) into an encoding-decoding process to form an edited image (213), wherein the encoding-decoding process includes: (i) inputting an intermediate image (205) obtained based on the input image (201), and (ii) inputting the image editing instructions (207, 208) into a pre-trained generative diffusion model to form the edited image (213) based on the corresponding outputs of a first loss function and a second loss function, wherein the first loss function is used to preserve details in one or more parts of the input image, and the second loss function is used to guide the input image to be modified according to the image editing instructions.

2. The image processing apparatus (800) according to claim 1, characterized in that, The one or more processors (801) are used to form a guide image (214) that depicts the desired modification in one or more local regions of the input image.

3. The image processing apparatus (800) according to claim 2, characterized in that, The one or more processors (801) are configured to: form the guide image (214) from the input image (201) based on the image editing instructions (207, 208).

4. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The image editing instructions include text describing the content of the input image and / or the desired modifications to the input image.

5. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The encoding stage of the encoding-decoding process includes converting the input image into a noise vector to form the intermediate image.

6. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The decoding stage of the encoding-decoding process includes: inputting the intermediate image into the pre-trained generative diffusion model to denoise the intermediate image according to the image editing instructions to form the edited image.

7. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The input image (201) is a potential representation of the original image to be edited, which is formed by transforming the original image.

8. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The input image (201) includes latent image features, wherein the latent image features are updated according to the corresponding outputs of the first loss function and the second loss function.

9. The image processing apparatus (800) according to claim 8, characterized in that, The latent image features are updated at multiple time steps.

10. The image processing apparatus (800) according to claim 8 or 9, characterized in that, The first loss function calculates the distance between the latent image features of the input image and the corresponding latent image features of the edited image.

11. The image processing apparatus (800) according to claim 10, which is dependent on claim 2, is characterized in that, The second loss function calculates the distance between the latent image features of the edited image and the corresponding latent image features of the guide image.

12. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The image editing instructions include the pose or doodles of the subject of the input image, or the edge map of the input image.

13. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The region to be modified in the input image (201) is defined by a mask (206), wherein the mask is integrated with the first loss function such that the first loss function is not applied to the region in the intermediate image corresponding to the region to be modified in the input image.

14. The image processing apparatus (800) according to any one of the preceding claims, characterized in that, The effects of the first loss function and the second loss function on the formation of the edited image can be configured by the user through adjustable parameters.

15. An image processing method (700), the method comprising: Receive (701) image editing instructions (207, 208), wherein the image editing instructions (207, 208) indicate a desired modification to the input image (201); The input image (201) is input (702) into an encoding-decoding process to form an edited image (213), wherein the encoding-decoding process includes: (i) inputting an intermediate image (205) obtained based on the input image, and (ii) inputting the image editing instructions (207, 208) into a pre-trained generative diffusion model to form an edited image (213) based on the corresponding outputs of a first loss function and a second loss function, wherein the first loss function is used to preserve details in one or more parts of the input image, and the second loss function is used to guide the input image to be modified according to the image editing instructions.

16. A computer program, when executed by a computing device (800), causes the computing device to perform the method (700) according to claim 15.