Image correction method, device and product
By obtaining image representation and time step k in the diffusion model, using target and source prompt information for inversion and sampling, and determining the correction deviation, the problem of insufficient structural and behavioral similarity in image editing in the existing technology is solved, and higher image generation similarity is achieved.
Patent Information
- Application Number
- CN202510722073.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
Existing image editing methods based on diffusion models cannot effectively control the internal structure and similarity of images in appearance and behavior, resulting in error accumulation during the generation process.
By obtaining the image representation to be corrected and the corresponding time step k, inversion is performed based on the target prompt information to obtain the intermediate image representation or sequence, sampling is performed using the source prompt information, the auxiliary node image representation is determined, and the correction deviation is determined through the reference and source diffusion image representations to correct the image to be corrected.
We effectively utilize the structural information of source branches to constrain the image generation trajectory, avoid error accumulation, and improve the appearance and behavioral similarity of image editing.
Smart Images

Figure CN120672629A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image editing technology, and in particular to image correction methods, equipment and products. Background Art
[0002] Text-guided diffusion models have achieved excellent performance in image editing, generating highly realistic and controllable images. The editing process of text-guided diffusion models first inverts the true source image into a noisy representation, and then applies target cues to guide the generation of the edited image. This allows the image to conform to the desired guidance while maintaining key content.
[0003] Related art methods using diffusion models in the inversion process rely primarily on exchanging attention maps between the source image branch and the target branch (generated image branch) to achieve visual similarity between the edited image and the source image. However, this approach only ensures overall contour similarity and cannot effectively control the internal structure of the image or the similarity of the image content in appearance and behavior. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to provide an image correction method, device and product, aiming to improve the similarity of image editing in appearance and behavior.
[0005] To achieve the above objectives, an embodiment of the present application provides an image correction method, which is applied to a diffusion model. The method includes the following steps:
[0006] Obtaining a representation of an image to be corrected and a time step k corresponding to the representation of the image to be corrected;
[0007] Inverting the image representation to be corrected based on the target prompt information to obtain an intermediate image representation or a sequence of intermediate image representations, wherein the image representation to be corrected is obtained by sampling according to the target prompt information;
[0008] determining an auxiliary node image representation by means of the intermediate image representation or the sequence of intermediate image representations;
[0009] Sampling the auxiliary node image representation according to the source hint information to obtain a reference image representation or a reference image representation sequence;
[0010] A correction deviation is determined by using the reference image representation at time step k and the source diffusion image representation at time step k, and the image representation to be corrected is corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information.
[0011] In some embodiments, before obtaining the representation of the image to be corrected, the image correction method includes:
[0012] Acquire the source image, the source prompt information, and the target prompt information corresponding to the generated image;
[0013] Inverting the source image according to the source hint information to obtain a final source diffusion image representation and a source diffusion image representation sequence corresponding to source image branches;
[0014] Sampling the final source diffusion image representation according to the target prompt information to obtain a generated image representation sequence corresponding to a generated image branch;
[0015] The image representation to be rectified and the time step k are determined based on the generated image representation sequence.
[0016] In some embodiments, determining the auxiliary node image representation by using the intermediate image representation or the intermediate image representation sequence comprises:
[0017] If the time step corresponding to the intermediate image representation is the same as the time step corresponding to the final source diffusion image representation, the intermediate image representation is determined as the auxiliary node image representation.
[0018] In some embodiments, determining the auxiliary node image representation by using the intermediate image representation or the intermediate image representation sequence further comprises:
[0019] An intermediate image representation in the intermediate image representation sequence having the same time step as that corresponding to the final source diffusion image representation is determined as an auxiliary node image representation.
[0020] In some embodiments, determining the correction bias using the reference image representation at time step k and the source diffusion image representation at time step k comprises:
[0021] The correction deviation is determined by the difference between the source diffusion image representation at the time step k and the reference image representation at the time step k.
[0022] In some embodiments, correcting the image representation to be corrected based on the correction deviation includes:
[0023] A corrected image representation is determined according to the correction deviation, the hyperparameter, and the image representation to be corrected.
[0024] In some embodiments, inverting the representation of the image to be corrected based on the target prompt information includes:
[0025] The image representation to be corrected is inverted using a denoising diffusion implicit model (DDIM) inversion process based on target cue information.
[0026] In some embodiments, determining the image representation to be corrected based on the generated image representation sequence includes:
[0027] The to-be-corrected image representation is determined by generating the image representation at a predetermined number of consecutive time steps close to the time step of the final source-diffusion image representation.
[0028] To achieve the above objectives, another aspect of the present application provides an image correction device, comprising:
[0029] An acquisition module, configured to acquire a representation of an image to be corrected and a time step k corresponding to the representation of the image to be corrected;
[0030] an inversion module, configured to invert the image representation to be corrected based on the target prompt information to obtain an intermediate image representation or a sequence of intermediate image representations, wherein the image representation to be corrected is obtained by sampling according to the target prompt information;
[0031] a determination module, configured to determine an auxiliary node image representation by using the intermediate image representation or the intermediate image representation sequence;
[0032] a sampling module, configured to sample the auxiliary node image representation according to the source prompt information to obtain a reference image representation or a reference image representation sequence;
[0033] A correction module is configured to determine a correction deviation using the reference image representation at time step k and the source diffusion image representation at time step k, and correct the image representation to be corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information.
[0034] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned image correction method when executing the computer program.
[0035] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer program product, which, when executed in an electronic device, enables the electronic device to perform the above-mentioned image correction method.
[0036] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned image correction method is implemented.
[0037] The embodiments of the present application include at least the following beneficial effects:
[0038] The present application provides an image correction method, device and product. The solution obtains an image representation to be corrected and a time step k corresponding to the image representation to be corrected, and inverts the image representation to be corrected based on target prompt information to obtain an intermediate image representation or an intermediate image representation sequence, wherein the image representation to be corrected is obtained by sampling according to the target prompt information. Afterwards, an auxiliary node image representation is determined by the intermediate image representation or the intermediate image representation sequence, and the auxiliary node image representation is sampled by the source prompt information to obtain a reference image representation or a reference image representation sequence. Then, a correction deviation is determined by the reference image representation with a time step of k and the source diffusion image representation with a time step of k, and the image representation to be corrected is corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information. In an embodiment of the present application, the representation of the image to be corrected is inverted through the target prompt information to determine the auxiliary node image representation, and the auxiliary node image representation is sampled through the source prompt information to obtain the reference image representation, thereby constructing a symmetrical image editing process that is the reverse of the inversion and sampling of the source diffusion model, effectively utilizing the structural information contained in the source image branch; determining the correction deviation through the reference image representation and the source diffusion image representation, and correcting the image representation to be corrected through the correction deviation can constrain the image generation trajectory, thereby generating an image with higher structural similarity to the source image; by determining the auxiliary node image representation and using the auxiliary node image representation as the sampling starting point of the reverse image editing process, the erroneous accumulation of source noise during the image generation process can be avoided, and the generation process can be effectively adjusted to improve the similarity of image editing in appearance and behavior.
[0039] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0041] Figure 1 is a flowchart of an image correction method provided by some embodiments of the present application;
[0042] Figure 2 is a process diagram of an image correction method provided by some embodiments of the present application;
[0043] FIG3( a ) is a schematic diagram of the process of an image editing method in the related art;
[0044] FIG3( b ) is a schematic diagram of the process of image correction and editing methods provided by some embodiments of the present application;
[0045] Figure 4 This is a comparison chart of the effects of the image editing method of the related art and the image correction editing method provided by some embodiments of the present application;
[0046] Figure 5 is a schematic block diagram of modules of an image correction device provided by some embodiments of the present application;
[0047] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the reference to "embodiment" in this article means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0049] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0050] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0052] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0053] In order to make the inventive concept of the present application easy to understand, before explaining the embodiments of the present application in detail, the English abbreviations (terms) / related concepts involved in the embodiments of the present application are first explained. The English abbreviations (terms) / related concepts involved in the embodiments of the present application are subject to the following explanations.
[0054] Diffusion models: Diffusion models are a class of deep learning models used to generate data. They have attracted widespread attention in the field of generative models, particularly for image, audio, and text generation. Diffusion models simulate the process of gradually adding noise to data and then learn how to reverse this process to generate new data samples. Diffusion models typically involve two main processes: diffusion and denoising.
[0055] Diffusion process (Forward Process): Also known as the inversion process, the diffusion model gradually adds noise to the data until the data becomes random and loses its original characteristics. The amount of noise added at each step is gradually increased. This process is usually reversible and can be described as a series of transformation steps.
[0056] Denoising (Reverse Process): Also known as the sampling process, the denoising process is the reverse process of the diffusion process. In this stage, the model learns how to recover the original data from the noisy data.
[0057] Time steps (timesteps): The diffusion process is usually divided into a number of time steps, each step corresponding to a period of increasing noise in the data.
[0058] Conditional guidance (prompt information): In some types of diffusion models, conditional information (such as text, part of an image, etc.) can be introduced to guide the generation process so that the generated data meets specific conditions.
[0059] DDIM: Deep Implicit Diffusion Models (DDIM) is a variant of the diffusion model and has a faster sampling speed than the original diffusion model (such as DDPM).
[0060] Unlike image generation, which relies on limited input, image editing is based on the original image and editing instructions to modify the objects, colors, shapes, styles, and actions in the image. It has extremely wide application value in many fields, such as virtual try-on and design. Advances in image editing are driving many scenarios from labor-intensive to digitally driven processes. Imagine a scenario where you take a photo and then, with just a simple sentence, you can directly modify the background or remove unwanted objects. Editing prompts can come in a variety of forms, including text, images, masks, etc., but the basic principles remain the same: one is editing fidelity, ensuring that the editing results are accurate and consistent with the provided instructions; the other is basic content preservation, which involves reverse mapping the original image into the diffusion latent space, especially those areas that do not need to be modified, while ensuring accurate reconstruction during the editing process.
[0061] Diffusion models have been widely used in the field of image editing due to their ability to generate highly realistic and controllable images. Text-based image editing involves applying DDIM inversion to the source image based on the source cue to generate a noisy representation. The target cue information is then used as a condition on this representation to guide the generation of subsequent edited images. Methods used in the related art to improve editing fidelity include: merging the attention maps of the source and target branches through attention integration, or adjusting the latent representation of the source branch to move closer to the latent representation of the target branch through latent representation (latent feature) integration. In addition, some methods decompose the editing process and perform editing in multiple steps. Content preservation is also addressed by similar methods, such as replacing the attention map of the target branch with the attention map of the source image, or modifying the DDIM inversion process to enhance the influence of the source cue on the generation of the target cue.
[0062] However, in these existing methods, both the target editing branch and the source reconstruction branch start from a noisy image obtained by DDIM inversion of the source image. This ignores a key detail: the starting point is not pure noise, but contains the structural information of the source image. This detail leads to two problems: first, the existing inversion-based methods only replace the attention map that interacts with the source image cues, and fail to utilize the structural information contained in the source branch; second, in some scenarios, such as pose modification, the edited image and the original image should be noised to different representations. However, using a noisy representation containing source image information as the starting point for the target image (generated image) will lead to the accumulation of errors in the generation process. Both of these factors lead to a gradual deviation from a reasonable generation process.
[0063] In view of this, the present application proposes an image correction method, device and product, which are applied to a diffusion model. In an embodiment of the present application, an intermediate image representation or an intermediate image representation sequence is obtained by obtaining the image representation to be corrected and the time step k corresponding to the image representation to be corrected, and inverting the image representation to be corrected based on the target prompt information, wherein the image representation to be corrected is obtained by sampling according to the target prompt information. Afterwards, the auxiliary node image representation is determined by the intermediate image representation or the intermediate image representation sequence, and the auxiliary node image representation is sampled by the source prompt information to obtain a reference image representation or a reference image representation sequence. Then, the correction deviation is determined by the reference image representation with a time step of k and the source diffusion image representation with a time step of k, and the image representation to be corrected is corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information. In an embodiment of the present application, the image representation to be corrected is inverted through target prompt information to determine the auxiliary node image representation, and the auxiliary node image representation is sampled through source prompt information to obtain a reference image representation. In this way, a symmetrical image editing process that is the reverse of the inversion and sampling of the source diffusion model is constructed, which effectively utilizes the structural information contained in the source branch; the correction deviation is determined through the reference image representation and the source diffusion image representation, and the image representation to be corrected is corrected through the correction deviation to constrain the image generation trajectory, thereby generating an image with higher structural similarity to the source image; by determining the auxiliary node image representation and using the auxiliary node image representation as the sampling starting point of the reverse image editing process, the erroneous accumulation of source noise during the image generation process can be avoided, and the generation process can be effectively adjusted to improve the similarity of image editing in appearance and behavior.
[0064] The image correction method provided in the embodiments of the present application relates to the field of image editing technology. It can be applied to the electronic device provided in the present application. The electronic device can be a terminal or a server.
[0065] In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a car terminal, etc., but is not limited thereto.
[0066] The server side can be configured as an independent physical server, or as a server cluster or distributed system consisting of multiple physical servers. It can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in the blockchain network.
[0067] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0068] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0069] The following describes in detail the implementation steps of an image correction method provided by an embodiment of the present application with reference to the accompanying drawings.
[0070] Please refer to Figure 1 , Figure 1 Flowcharts of image correction methods provided for some embodiments of the present application. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system, such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be performed in a different order than shown.
[0071] The method of the embodiment of the present application includes the following steps:
[0072] Step 101: Obtain the image representation to be corrected and the time step k corresponding to the image representation to be corrected;
[0073] Step 102: Inverting the image representation to be corrected based on the target prompt information to obtain an intermediate image representation or a sequence of intermediate image representations, wherein the image representation to be corrected is obtained by sampling based on the target prompt information;
[0074] Step 103: Determine an auxiliary node image representation through an intermediate image representation or a sequence of intermediate image representations;
[0075] Step 104: sampling the auxiliary node image representation using the source hint information to obtain a reference image representation or a reference image representation sequence;
[0076] Step 105: Determine a correction deviation using a reference image representation at time step k and a source diffusion image representation at time step k, and correct the image representation to be corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information.
[0077] In steps 101 to 105 shown in the embodiment of the present application, the image representation to be corrected is inverted through the target prompt information to determine the auxiliary node image representation, and the auxiliary node image representation is sampled through the source prompt information to obtain the reference image representation, thereby constructing a symmetrical image editing process that is the reverse of the inversion and sampling of the source diffusion model, effectively utilizing the structural information contained in the source branch; determining the correction deviation through the reference image representation and the source diffusion image representation, and correcting the image representation to be corrected through the correction deviation can constrain the image generation trajectory, thereby generating an image with higher structural similarity to the source image; by determining the auxiliary node image representation and using the auxiliary node image representation as the sampling starting point of the reverse image editing process, the erroneous accumulation of source noise during the image generation process can be avoided, and the generation process can be effectively adjusted to improve the similarity of image editing in appearance and behavior.
[0078] The specific implementation methods of the above steps are introduced below.
[0079] Before performing step 101 of obtaining a representation of the image to be corrected, the image to be corrected may be determined first.
[0080] In some embodiments of the present application, a source image, source prompt information, and target prompt information corresponding to a generated image may be obtained, and then the source image may be inverted based on the source prompt information to obtain a final source diffusion image representation and a sequence of source diffusion image representations corresponding to source image branches. The final source diffusion image representation may be sampled based on the target prompt information to obtain a sequence of generated image representations corresponding to the generated image branches, and the image representation to be corrected and the time step k may be determined based on the generated image representation sequence.
[0081] Alternatively, the source image can be a real-world image. Source hints describe the content of the source image, informing the image editing system or algorithm about the image's existing features, content, and context. Examples include a cat wearing a pink hat, a dog lying on a white background, or a chair in front of a wall with a floral pattern.
[0082] The generated image (target image) can be generated based on the source image and the target prompt information. For example, given a source image and the source prompt information (a cat wearing a pink hat), a corresponding image can be generated based on the target prompt information (a tiger wearing a pink hat). The target prompt information is a description of the modification or edit that the user wishes to make to the image. It guides the system on how to modify the source image to achieve the desired effect.
[0083] After the source image and source hint information are obtained, the source image can be inverted according to the source hint information to obtain a final source diffusion image representation and a source diffusion image representation sequence corresponding to the source image branch.
[0084] Specifically, the original image can be converted into a latent representation (latent features) in a latent space. The latent space refers to the low-dimensional representation space within the data learned by the neural network. In this space, the complex features of the data are compressed and represented in a simpler form, which helps the generative model understand and generate data. The source image is then inverted based on the source hint information, which can be expressed as:
[0085]
[0086] in, is the source prompt condition, t represents the time step, ∈ θ Represents a deep learning network, which can be, for example, a Unet network, or a trained network such as Stable Diffusion-1.5, α t is a hyperparameter with time step t, Represents the source diffusion image representation corresponding to the source image branch (i.e., the potential representation of the source diffusion image), which can be understood as the noise image representation of the intermediate time step from the source image representation to the final source diffusion image representation. When t is 0, Represents the source image representation. When t takes the value T, Represents the final source diffusion image representation. Thus, when t is in the range between 0 and T, represents the sequence of source diffusion image representations corresponding to the source image branches.
[0087] Afterwards, the final source diffusion image representation is sampled according to the target cue information to obtain a sequence of generated image representations corresponding to the generated image branch.
[0088] Specifically, the denoising (sampling) process of the target image can be expressed as follows using a method without classifier guidance:
[0089]
[0090] in, Indicates target prompt information. Indicates an empty condition, represents the sequence of generated image representations corresponding to the generated image (target image) branch.
[0091] Therefore, the traditional target image denoising process can be expressed as:
[0092]
[0093] in, represents the final source diffusion image representation, and The image representations at the time steps between are the sequence of generated image representations corresponding to the generated image branch.
[0094] It is understandable that the first denoising (sampling) step can be achieved by The injection target prompt information is expressed by referring to formula (2).
[0095] During the image generation step, the image representation to be rectified and the time step k can be determined.
[0096] In some embodiments, determining the image representation to be corrected based on the sequence of generated image representations may be performed by generating image representations of a preset number of consecutive time steps close to the time step of the final source diffusion image representation.
[0097] Optionally, if the time step of the final source diffusion image representation is expressed as T, then generated images of a preset number of consecutive time steps close to the time step T may be selected to determine the image representation to be corrected.
[0098] For example, the preset number of time steps may be 1 / 4 of the total number of time steps. In other words, the generated image representation corresponding to the first 1 / 4 of the denoising time steps may be determined as the image representation to be corrected.
[0099] It should be understood that an image generated at any time step may also be determined as the image to be corrected, which can be flexibly set according to actual conditions.
[0100] By determining the generated image representations for a preset number of consecutive time steps close to the time step of the final source diffusion image representation as the image representation to be corrected, rather than determining all generated image representations as the image representation to be corrected, computing power can be saved. This is because, during the denoising process, the diffusion model prioritizes restoring low-frequency contour information, followed by high-frequency texture details. Since high-frequency information does not need to remain consistent before and after editing, while low-frequency semantics should be preserved as much as possible, the selection of corrected images in this application is mainly focused on the initial low-frequency recovery stage.
[0101] After the representation of the image to be corrected and the time step k are determined, the image to be corrected can be corrected.
[0102] In step 101, a representation of an image to be corrected and a time step k corresponding to the representation of the image to be corrected are obtained.
[0103] The image to be corrected may be the generated image in the initial low-frequency recovery stage of the diffusion model, or may be understood as the generated image of a preset number of consecutive time steps close to the time step of the final source diffusion image.
[0104] Image representation is a form of representation of an image that can be understood and processed by computers or other systems. It can be the potential representation (latent features) of an image in a latent space or an image encoding.
[0105] By obtaining the image representation to be corrected and the time step k corresponding to the image representation to be corrected, data support is provided for subsequent image correction.
[0106] In step 102 , the image representation to be corrected is inverted based on the target prompt information to obtain an intermediate image representation or a sequence of intermediate image representations, wherein the image representation to be corrected is obtained by sampling according to the target prompt information.
[0107] During the inversion of the image representation to be corrected based on the target hint information, if the time step k to be corrected is T-1, that is, the image to be corrected is the generated image at the first time step of the sampling (denoising) stage, then inversion is performed to obtain an intermediate image representation, whose time step is denoted as T. If the time step k to be corrected is close to time step T but not T-1, then inversion is performed to obtain a sequence of intermediate image representations. It should be understood that the image representation to be corrected is inverted until an intermediate image representation corresponding to time step T of the final source diffusion image representation is generated, where the time step of the intermediate image representation corresponding to time step T of the final source diffusion image representation is also T.
[0108] The inversion process with the time step k to be corrected being T-1 can be expressed as:
[0109]
[0110] To prompt information according to the target The intermediate image representation obtained by inverting the image to be corrected at time step k T-1.
[0111] In some embodiments, the image representation to be rectified may be inverted using a denoising diffusion implicit model (DDIM) inversion process based on target cue information.
[0112] By using the DDIM inversion process to invert the representation of the image to be rectified, the image sampling speed can be improved and the quality of the generated image samples can be guaranteed.
[0113] In step 103 , an auxiliary node image representation is determined by means of the intermediate image representation or a sequence of intermediate image representations.
[0114] Auxiliary node image representations can be determined from the generated intermediate image representations. These auxiliary node image representations can serve as the sampling starting point for the reverse image editing process and also as a bridge for the entire reverse text guidance process. This can avoid the erroneous accumulation of source noise during the image generation process, and effectively adjust the generation process to improve the similarity of image edits in appearance and behavior.
[0115] In some embodiments, if the time step corresponding to the intermediate image representation is the same as the time step corresponding to the final source diffusion image representation, the intermediate image representation is determined as the auxiliary node image representation; alternatively, the intermediate image representation in the intermediate image representation sequence that has the same time step corresponding to the final source diffusion image representation is determined as the auxiliary node image representation.
[0116] An intermediate image representation or an intermediate image representation sequence is obtained through the above step 102. In the case where only one intermediate image representation is obtained (ie, ), the intermediate image obtained can be represented as As an auxiliary node image representation; in the case of obtaining an intermediate image representation sequence, the intermediate image representation corresponding to the time step T of the final source diffusion image representation can be used as the auxiliary node image representation, which can also be expressed as
[0117] In step 104, the auxiliary node image representation is sampled using the source hint information to obtain a reference image representation or a reference image representation sequence;
[0118] Optionally, it can be based on auxiliary node image representation Injection source prompt information Then obtain the reference image representation for denoising through source hint information
[0119] It should be understood that if the time step k is T-1, then by sampling the auxiliary node image representation through the source prompt information, a reference image representation can be obtained If the time step k is not T-1, a reference image representation sequence can be obtained, and the time step of the reference image representation corresponding to the last time step of the sequence is k.
[0120] The time step k is T-1, and the process of sampling the auxiliary node image representation through the source prompt information can be expressed as:
[0121]
[0122] The meanings of the parameters in formula (5) can be referred to above and will not be repeated here.
[0123] The embodiment of the present application samples the auxiliary node image representation through source prompt information to obtain a reference image representation or a reference image representation sequence, which can provide data support for the final determination of the correction deviation.
[0124] In step 105, a correction deviation is determined by using a reference image representation at time step k and a source diffusion image representation at time step k, and the image representation to be corrected is corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information.
[0125] The correction deviation can be determined using the reference image and the source diffusion image corresponding to the time step k of the image to be corrected.
[0126] Here, the source diffusion image representation can be generated by inverting the source image according to the source prompt information through formula (1).
[0127] In some embodiments, the correction deviation is determined by the reference image representation at time step k and the source diffusion image representation at time step k, and the correction deviation can be determined by the difference between the source diffusion image representation at time step k and the reference image representation at time step k.
[0128] For example, the correction deviation of the image to be corrected at time step k T-1 can be expressed as:
[0129]
[0130] Ideally, Should be with Consistent, because arrive and from arrive For example, if the prompt information is "cat on the table" and "dog on the table", the basic content preservation will be better if the editing process from "cat" to "dog" and then back to "cat" is consistent.
[0131] It should be understood that ideal should be equal to This will be and Provide guidance and direction.
[0132] The correction deviation can also be called an offset, which itself carries the directional information used for correction. Therefore, correction can be achieved by simply adding the offset to the corresponding image representation to be corrected.
[0133] In some embodiments, correcting the image representation to be corrected based on the correction deviation may be performed by determining the corrected image representation based on the correction deviation, the hyperparameter, and the image representation to be corrected.
[0134] Since the corrected bias is an estimate, it can be multiplied by a hyperparameter before correction.
[0135] For example, the correction of the image representation to be corrected based on the correction deviation can be expressed as:
[0136]
[0137] The above example shows a round of correction process of the intermediate representation in the first step of the denoising process, which can be called consistent correction. As mentioned above, the example of this application only shows the correction process of the intermediate image with time step T-1, and the consistent correction method provided by this application can be extended to any time step t. It is understandable that before starting the reverse editing, it is necessary to add noise (inversion) according to the target prompt information until
[0138] The following is the pseudo code of the entire consistent correction process provided by the embodiment of the present application:
[0139]
[0140]
[0141] The pseudo code of the entire consistent correction process provided in the embodiment of the present application is divided into three parts.
[0142] Among them, the input of pseudocode includes: source prompt information Target prompt information Source image, source image representation Time steps T, time steps to be corrected K, trained denoising network ∈ θ .
[0143] The output of the pseudocode includes the target image and the target image representation
[0144] The first part of the pseudo code represents the noise addition process, namely the DDIM inversion process. You can refer to formula (1) to achieve the representation of the source image: Perform inversion to generate the inversion trajectory. The time step t is a positive integer from 1 to T.
[0145] Specifically, the first part is to continuously add noise to a given original image to obtain a latent representation of the image in the corresponding Gaussian noise distribution space. Ideally, the obtained noise can be perfectly reconstructed back to the original image after a given prompt text is given.
[0146] The second part of the pseudocode is the consistent correction process for time step k, which specifically involves calculating and injecting the correction bias.
[0147] First, refer to equation (1), i.e., DDIM_Inversion, to invert the image representation to be corrected at time step k in the target image branch to obtain the inversion trajectory, where the time step is a positive integer from k+1 to T.
[0148] The image representation with a time step T obtained by inversion is set as the auxiliary node image representation.
[0149] Afterwards, referring to formula (2), DDIM sampling is performed on the auxiliary node image representation, and the time step value is a positive integer from T to k+1.
[0150] Calculate the correction deviation and correct the image representation to be corrected.
[0151] Specifically, the second part of the pseudocode assumes that the generated image at step k needs to be corrected for deviation. This involves calculating the deviation for the reverse editing process: denoising the intermediate image at step k using the target cue information and then denoising it using the original cue information. This is symmetrical to the initial image editing process. After calculating the deviation, it is directly applied to the generated image at step k, ensuring that the generated image at step k meets expectations.
[0152] The third part of the pseudo code represents the denoising (sampling) process, namely the DDIM sampling (DDIM_Sampling) process. You can refer to formula (2) to achieve the final diffusion image representation Sampling is performed to obtain the generated image trajectory. The time step t is a positive integer from T to 1. If the time step is within the range of the number of time steps to be corrected K, the generated image is corrected using the consistent correction process in the second part of the pseudocode.
[0153] The third part of the pseudocode is the standard DDIM sampling process, which injects the target prompt text into the diffuse image (noise) obtained by adding noise to the original image for denoising. During this process, the second part, the consistency correction part, is called according to the execution conditions to correct the deviation of the generated image.
[0154] It should be understood that the above pseudocode is only one embodiment of the image correction method provided herein. Those skilled in the art can flexibly configure the parameters of the pseudocode based on actual applications, particularly the third portion, which determines the representation of the image to be corrected. For example, a piecewise function can be used to set the time step information of the image to be corrected. For example, the function can be set to a continuous sequence of time steps at the start of the denoising phase or any discrete time step during the denoising process.
[0155] Figure 2 This is a process diagram of the image correction method provided in some embodiments of the present application.
[0156] Figure 2 In the figure, the gray circle represents the source latent representation, the blue-green circle represents the auxiliary latent representation, the grass-green circle represents the latent representation to be corrected, and the yellow circle represents the latent representation after correction (it is not excluded that the actual display color does not match the description color due to color difference. Please refer to the attached figure for details). Figure 2 The legend in the figure shall prevail). Among them, z - and z + where represents the source image representation and the target (generated) image representation, respectively. +S represents image inversion under the condition of source cue information S, and +T represents image sampling under the condition of target cue information T. ①-⑤ represent the consistent correction process, where the first step is to directly inject the sampling of the target cue information (denoising), the second step is to inject the target cue information for inversion (noising), the third step is to inject the sampling of the source cue information (denoising), the fourth step is to calculate the correction deviation of the inverse editing, and the fifth step is to correct the deviation calculated in step 4 by applying it to the latent to be corrected. Figure 2 The process in can also be understood in conjunction with the pseudo code portion of this application.
[0157] Figures 3(a) and 3(b) illustrate the image editing process of the related art and the image editing process using the image correction method provided by the embodiments of the present application, respectively. In Figures 3(a) and 3(b), the upper red branch represents the DDIM inversion process, and the lower blue branch represents the DDIM sampling process.
[0158] As shown in Figure 3(a), the source image is a cat wearing a pink hat, and the source hint information describes the source image. The target hint information is a tiger wearing a pink hat. The source image needs to be edited according to the target hint information to generate the target image.
[0159] The image editing process of the related art generates a time step sequence corresponding to the source image branch by inverting the source image (diffusion or noise addition), and then directly converts the last noise image (the final source diffusion image) into the It uses the image as the starting node for denoising and generates a time step sequence corresponding to the generated image branch, and finally obtains the target generated image, which is a tiger wearing a pink hat.
[0160] As can be seen, the target generated image and the source image have an overall visual similarity, but their appearance and behavior still deviate from the source image. This is because the image editing process in related technologies only replaces the attention map that interacts with the source image cues, failing to utilize the potential structural information contained in the source branch. In addition, since the actual node (the last noise image) is not pure noise, this leads to the accumulation of errors in the image generation process. These factors lead to a gradual deviation in the generation process.
[0161] Figure 3(b) illustrates the image editing process using the image correction method provided by an embodiment of the present application. Given the same source image, source prompt information, and target prompt information as in Figure 3(a), the consistency correction method provided by an embodiment of the present application was used at the initial denoising stage to leverage the structural information of the source image to guide the denoising process. This results in the generated tiger image being closer in appearance and behavior to the source image.
[0162] The consistency correction method provided in the embodiment of the present application builds a bridge between the source branch and the target branch, improves the editing fidelity, and effectively retains important information during the generation process. Specifically, the consistency correction method constructs an auxiliary editing starting point (auxiliary node), and then performs reverse editing by conditionally restricting the auxiliary node through the source prompt information. The virtual noise representation (reference image representation) generated by reverse editing should be very similar to the representation of the current time step of the source image branch, satisfying reverse consistency. This effectively monitors whether the editing process of the target image (generated image) is over-edited and destroys important information. In addition, the consistency deviation can be calculated and incorporated into the denoising stage to further optimize the editing trajectory. This correction can be applied in multiple steps to enhance the generation of the entire editing content.
[0163] The consistency correction method provided in the embodiment of the present application effectively optimizes the editing trajectory and effectively utilizes the potential structural information of the source image. In addition, by using the auxiliary node as the starting noise representation of the editing process, the generated image branch can be decoupled from the source image branch, and the generation process can be further adjusted to improve the similarity of image editing in appearance and behavior.
[0164] Figure 4 This is a comparison chart of the effects of image editing methods in related technologies and image correction editing methods provided by some embodiments of this application. The source image is on the far left. From left to right, the methods compared are DDIM inversion, NULL-Text inversion (NT inversion), Negative-Prompt inversion (NP inversion), StyleDiffusion (StyleD inversion), Instruct-Pix2Pix (Instruct inversion), direct inversion (DI inversion), and the method provided by the embodiments of this application (consistency correction method) combined with the general editing method Prompt-to-Prompt / Plug-and-Play.
[0165] Figure 4 , when editing a dog into a lion, the relevant methods either changed the lion's morphology or introduced unreasonable details. Similarly, when editing a cat into a tiger, the method provided in the embodiment of the present application can keep the tiger's behavior consistent with that of the cat, and the outline is more similar. In the third row, the pattern was edited by replacing the floral pattern with spots. Direct inversion over-retained the background and failed to effectively replace the floral pattern with spots. Other methods even changed the shape of the chair. In contrast, the method provided in the embodiment of the present application achieved successful editing without such artifacts.
[0166] The above is an introduction to the image correction method provided in the embodiments of the present application.
[0167] The implementation of the image correction device provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0168] Regarding the image correction method provided in the above embodiment, the present application also provides an image correction device for implementing the above method, such as Figure 5 As shown, Figure 5 This is a schematic block diagram of the modules of the image correction device according to an embodiment of the present application. The image correction device 500 includes:
[0169] An acquisition module 501 is used to acquire the image representation to be corrected and the time step k corresponding to the image representation to be corrected;
[0170] an inversion module 502 for inverting the image representation to be corrected based on the target prompt information to obtain an intermediate image representation or a sequence of intermediate image representations, wherein the image representation to be corrected is obtained by sampling based on the target prompt information;
[0171] A determination module 503 is configured to determine an auxiliary node image representation through an intermediate image representation or a sequence of intermediate image representations;
[0172] A sampling module 504 is configured to sample the auxiliary node image representation using the source hint information to obtain a reference image representation or a reference image representation sequence;
[0173] The correction module 505 is used to determine a correction deviation using a reference image representation at a time step k and a source diffusion image representation at a time step k, and correct the image representation to be corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information.
[0174] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0175] Reference Figure 6 The embodiment of the present application further provides an electronic device 600, which includes a memory 601, one or more processors 602 ( Figure 6 Only one is shown) and a computer program stored in memory 601 and executable by processor 602. Memory 601 is used to store software programs and units. Processor 602 executes the software programs and units stored in memory 601 to perform various functional applications and data processing to obtain resources corresponding to the aforementioned preset events. Optionally, processor 602 implements the aforementioned image correction method by executing the computer program stored in memory 601.
[0176] The memory 601 is a non-transitory computer-readable medium that can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory 601 may include a high-speed random access memory and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 601 may optionally include a memory remotely located relative to the processor 602, and these remote memories may be connected to the processor 602 via a network.
[0177] It can be understood that the contents of the above method embodiments are all applicable to the electronic device embodiment. The functions specifically implemented by the electronic device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0178] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned image correction method is implemented.
[0179] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0180] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by one or more processors, it can implement the steps of the above-mentioned image correction method.
[0181] It can be understood that the contents of the above method embodiments are all applicable to this computer program product, the functions specifically implemented by this computer program product embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0182] The image correction method, device, electronic device, medium and computer program product provided in the embodiments of the present application utilize the symmetry of inversion and sampling in the diffusion model to construct an inverse consistency deviation, and inject it into the denoising process of the target image to intervene in and constrain the editing trajectory. Specifically, an auxiliary node is constructed to separate the noise addition points of the source image and bridge the source image and the target image to calculate the consistency deviation. This deviation reflects the modification of the structural information of the source image by the editing process, and injecting it into the target image effectively corrects the editing trajectory. Extensive experimental results on various editing types have demonstrated the effectiveness of the image correction method provided by the embodiments of the present application.
[0183] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0184] Although specific embodiments are described herein, those skilled in the art will recognize that many other modifications or alternative embodiments are also within the scope of the present disclosure. For example, any of the functions and / or processing capabilities described in conjunction with a particular device or component may be performed by any other device or component. In addition, although various exemplary implementations and architectures have been described in accordance with the embodiments of the present disclosure, those skilled in the art will recognize that many other modifications to the exemplary implementations and architectures described herein are also within the scope of the present disclosure.
[0185] Some aspects of the present disclosure have been described above with reference to the block diagrams and flow charts of the systems, methods, systems and / or computer program products according to the exemplary embodiments. It should be understood that the combination of one or more blocks in the block diagram and the flow chart and the blocks in the block diagram and the flow chart can be realized by executing computer executable program instructions respectively. Equally, according to some embodiments, some blocks in the block diagram and the flow chart may not need to be executed in the order shown, or may not need to be executed in full. In addition, additional components and / or operations beyond those components and / or operations shown in the blocks in the block diagram and the flow chart may be present in certain embodiments.
[0186] Therefore, the blocks in the block diagrams and flow charts support combinations of means for performing the specified functions, combinations of elements or steps for performing the specified functions, and program instruction means for performing the specified functions. It should also be understood that each block in the block diagrams and flow charts, and combinations of blocks in the block diagrams and flow charts, can be implemented by a dedicated hardware computer system that performs the specific functions, elements, or steps, or a combination of dedicated hardware and computer instructions.
[0187] The program modules, applications, etc. described herein may include one or more software components, including, for example, software objects, methods, data structures, etc. Each such software component may include computer-executable instructions that, in response to execution, cause at least a portion of the functionality described herein (e.g., one or more operations of the illustrative methods described herein) to be performed.
[0188] Software component can be encoded with any one in various programming languages.A kind of exemplary programming language can be low-level programming language, such as the assembly language associated with specific hardware architecture and / or operating system platform.Comprise that the software component of assembly language instruction may need to be converted to executable machine code by assembler before being executed by hardware architecture and / or platform.Another exemplary programming language can be a more advanced programming language, and it can be transplanted across multiple architectures.Comprise that the software component of more advanced programming language may need to be converted to intermediate representation by interpreter or compiler before execution.Other examples of programming language include but are not limited to macro language, shell or command language, job control language, script language, database query or search language or report writing language.In one or more exemplary embodiments, the software component that comprises the instruction of one in the above-mentioned programming language example can be directly executed by operating system or other software component, without first being converted into another form.
[0189] Software components can be stored as files or other data storage structures. Software components of similar types or related functions can be stored together, such as in a specific directory, folder, or library. Software components can be static (e.g., preset or fixed) or dynamic (e.g., created or modified at execution time).
[0190] The embodiments of the present application are described in detail above in conjunction with the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by ordinary technicians in the relevant technical field without departing from the purpose of the present application.
Claims
1. An image correction method, characterized in that: Applied to the diffusion model, the method comprises the following steps: Obtaining a representation of an image to be corrected and a time step k corresponding to the representation of the image to be corrected; Inverting the image representation to be corrected based on the target prompt information to obtain an intermediate image representation or a sequence of intermediate image representations, wherein the image representation to be corrected is obtained by sampling according to the target prompt information; determining an auxiliary node image representation by means of the intermediate image representation or the sequence of intermediate image representations; Sampling the auxiliary node image representation according to the source hint information to obtain a reference image representation or a reference image representation sequence; A correction deviation is determined by using the reference image representation at time step k and the source diffusion image representation at time step k, and the image representation to be corrected is corrected based on the correction deviation, wherein the source diffusion image representation is obtained by inverting the source image according to the source prompt information.
2. The image correction method according to claim 1, wherein: Before obtaining the representation of the image to be corrected, the image correction method includes: Acquire the source image, the source prompt information, and the target prompt information corresponding to the generated image; Inverting the source image according to the source hint information to obtain a final source diffusion image representation and a source diffusion image representation sequence corresponding to source image branches; Sampling the final source diffusion image representation according to the target prompt information to obtain a generated image representation sequence corresponding to a generated image branch; The image representation to be corrected and the time step k are determined according to the generated image representation sequence.
3. The image correction method according to claim 1, wherein: The determining of the auxiliary node image representation by using the intermediate image representation or the intermediate image representation sequence comprises: If the time step corresponding to the intermediate image representation is the same as the time step corresponding to the final source diffusion image representation, the intermediate image representation is determined as the auxiliary node image representation.
4. The image correction method according to claim 1, wherein: The determining of the auxiliary node image representation by using the intermediate image representation or the intermediate image representation sequence further includes: An intermediate image representation in the intermediate image representation sequence having the same time step as the final source diffusion image representation is determined as an auxiliary node image representation.
5. The image correction method according to claim 1, wherein: The determining of the correction deviation by using the reference image representation at time step k and the source diffusion image representation at time step k comprises: The correction deviation is determined by the difference between the source diffusion image representation at the time step k and the reference image representation at the time step k.
6. The image correction method according to claim 1, wherein: The correcting the image representation to be corrected based on the correction deviation includes: A corrected image representation is determined according to the correction deviation, the hyperparameter, and the image representation to be corrected.
7. The image correction method according to claim 1, wherein: The inverting the representation of the image to be corrected based on the target prompt information includes: The image representation to be corrected is inverted using a denoising diffusion implicit model (DDIM) inversion process based on target cue information.
8. The image correction method according to claim 2, wherein: The determining the image representation to be corrected according to the generated image representation sequence includes: The to-be-corrected image representation is determined by generating the image representation at a predetermined number of consecutive time steps close to the time step of the final source-diffusion image representation.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the image correction method according to any one of claims 1 to 8 when executing the computer program.
10. A computer program product, characterized in that When the computer program product runs in an electronic device, the electronic device is enabled to perform the image correction method according to any one of claims 1 to 8.