Image editing implementation method and system based on large model feature injection
By using the large model feature injection method in image editing and calculating the difference value in KL divergence, the problem of identity feature retention in image editing is solved, and a high-fidelity image editing effect is achieved.
Patent Information
- Application Number
- CN202510056462.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
Smart Images

Figure CN119941927A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image editing, and in particular to an image editing implementation and system based on large model feature injection. Background Art
[0002] Large-model-based text-guided image editing (image editing is to replace the subject's features, including actions, style, texture, etc., under the guidance of text given a reference image, but retain the subject's identity characteristics) requires generating a target image with the subject unchanged but the style and action changed given a reference image.
[0003] There are two types of similar existing technologies. One is image style transfer based on a large model. This solution can only control the background style of the image, such as changing the style from modern to anime, but cannot modify the main action of the character. The second type is image generation based on a large model and edge extraction technology. This technology mainly uses the edge features of the reference image to control the approximate shape of the generated image, and cannot complete editing tasks with action modification. Neither of them can guarantee that the identity characteristics of the reference subject remain unchanged. Summary of the invention
[0004] The technical problem to be solved by the embodiments of the present invention is to provide an image editing implementation method and system based on large model feature injection to solve the problem of retaining identity consistency in image editing tasks.
[0005] In order to solve the above technical problems, an embodiment of the present invention proposes an image editing implementation method based on large model feature injection, including: Reference feature extraction step: perform image inversion based on the real reference image to obtain a reference intermediate feature map; Target feature extraction step: extract features from the target text and align them with the image feature space of the reference intermediate feature map. Input the extracted features into the image generation model to guide the model to perform reasoning, and obtain a target intermediate feature map that matches the text semantics. KL divergence calculation steps: In the reasoning process of each step of the large model, the KL divergence is used to calculate the difference between the reference intermediate feature map and the target intermediate feature map. If the difference is greater than the preset threshold, it is considered a valid change and features are extracted from the target text; if the difference is less than the preset threshold, it is considered an invalid change and features are extracted from the reference image to retain the identity of the source image; Image generation step: Merge the obtained features, input the merged features into the large model for inference, generate and output the target image.
[0006] Accordingly, an embodiment of the present invention further provides an image editing implementation system based on large model feature injection, comprising: Reference feature extraction module: performs image inversion based on the real reference image to obtain the reference intermediate feature map; Target feature extraction module: extracts features from the target text and aligns them with the image feature space of the reference intermediate feature map. The extracted features are input into the image generation model to guide the model to perform reasoning, and a target intermediate feature map that matches the text semantics is obtained. KL divergence calculation module: In the reasoning process of each step of the large model, the KL divergence is used to calculate the difference between the reference intermediate feature map and the target intermediate feature map. If the difference is greater than the preset threshold, it is considered a valid change and features are extracted from the target text; if the difference is less than the preset threshold, it is considered an invalid change and features are extracted from the reference image to retain the identity of the source image; Image generation module: Merge the acquired features, input the merged features into the large model for inference, generate and output the target image.
[0007] The beneficial effects of the present invention are as follows: the present invention applies KL divergence calculation difference on each pair of attention maps of the reference image and the target image, so that the large model selectively introduces the identity features of the reference image in the inference process of target image generation and discards defects, which can well preserve the identity consistency before and after image editing and make the edited image have higher fidelity; the present invention solves the problem of identity consistency preservation in image editing tasks through a simple and efficient method, and realizes input of a picture to be edited, corresponding reference text and target text, and output of a new picture that conforms to the semantics of the target text and is similar to the main body of the picture to be edited. Compared with the existing solutions, the present invention does not need to introduce fine-tuning, and can have higher image fidelity and text image similarity when performing image editing and video generation, and is more suitable for application in the field of image editing, and has higher flexibility for users. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a flowchart of an image editing implementation method based on large model feature injection according to an embodiment of the present invention.
[0009] Figure 2 It is a structural block diagram of the KL divergence screening feature graph according to an embodiment of the present invention.
[0010] Figure 3 This is an example diagram of an image editing implementation method based on large model feature injection according to an embodiment of the present invention.
[0011] Figure 4 (a) is an example diagram of a reference image according to an embodiment of the present invention, and (b) is an example diagram of a target image. DETAILED DESCRIPTION
[0012] It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention is further described in detail below in conjunction with the drawings and specific embodiments.
[0013] In the embodiments of the present invention, if there are directional indications (such as up, down, left, right, front, back, etc.), they are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0014] In addition, in the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of the features.
[0015] Please refer to Figure 1 The image editing implementation method based on the injection of large model features (in vision, it refers to the attention-related parameters in the Transformer structure) of an embodiment of the present invention includes a reference feature extraction step, a target feature extraction step, a KL divergence (a method for measuring the difference between distributions, the larger the value, the greater the difference) calculation step, and an image generation step.
[0016] Reference feature extraction step: Perform image inversion based on the real reference image to obtain a reference intermediate feature map.
[0017] As an implementation mode, an image inversion model is used in the reference feature extraction step to perform image inversion to obtain intermediate features of the reference image, and the image inversion model is composed of DDIM inversion or Null-Text Inversion. The image inversion model is preferably composed of DDIM inversion, which utilizes the mathematical laws of DDIM deduction to treat the real image as the generated image and obtain the intermediate feature map of the generation process.
[0018] The process of generating an image starts with Gaussian noise, and gradually obtains a complete image after many steps of denoising operations. Each step of the denoising process inputs a noise image of the previous time step, and calculates the value to be subtracted (denoising) through the self-attention structure and cross-attention structure of each layer of the Unet structure algorithm model. The noise image of the previous step is subtracted from the calculated value to obtain the noise image of the next step.
[0019] Since the reference image actually input by the user is not generated by the algorithm model (it is a real-world picture taken or drawn by the user), there are not so many intermediate noise images (image features) in this step. If you want to use the existing algorithm model for image editing, you need to first assume that it is a picture generated by the algorithm model, and then reverse the noise value that needs to be subtracted at each time step, the value calculated by self-attention, and the value calculated by cross-attention. These three values represent the image features of the generation process. The value calculated by self-attention contains more features such as the shape and identity of the image, and the value calculated by cross-attention contains more semantic features of the prompt text (that is, what the subject in the picture is doing). This reverse process is the process of obtaining the intermediate features of the reference image. There are many methods, and the more classic ones are DDIM inversion and Null-Text Inversion.
[0020] Specific process: The real reference image is reversely calculated through the noise calculation formula to obtain the denoising value of the previous time step. This value is added to obtain the intermediate noise map of this time step (self-attention, cross-attention, and the noise map before the current image denoising). This cycle is repeated for T time steps (the same number of steps is used for editing the large model image, and the same number of steps is used for reverse deduction here), so that the final output is a Gaussian noise map, that is, a series of reference intermediate feature maps are finally output.
[0021] The formula for calculating noise is: ; x t−1 represents the noise map at time step t−1. α t It is a preset constant, such as 0.7, which is used to control the degree of noise addition. is α t The cumulative value of the continuous multiplication, from the initial value, multiplied, α1, α2... until α t . ϵ θ ( x t , t ) is the trained diffusion model. t is the variance of a Gaussian distribution, usually set to a constant. ϵ represents a random noise sampled from a standard normal distribution.
[0022] Target feature extraction step: extract features from the target text and align them with the image feature space (text must be extracted by a text encoder before it can be injected into the model, and the image features used in computer vision are extracted by an image encoder. These two encodings cannot be directly one-to-one corresponding, so they need to be aligned. The alignment method is to use the CLIP model. This is a model obtained by comparative training on a large data set consisting of text and image pairs. It can effectively convert text features and image features into each other), input the extracted features into the image generation model to guide the model to perform reasoning, and obtain a target intermediate feature map that is consistent with the text semantics. As an implementation method, the CLIP model is used for feature extraction.
[0023] The target text often describes a subject's specific behavior and expression. For example, a shy girl, a running girl. In a normal image generation process, the text is extracted through the CLIP model to obtain an embedding aligned with the image. The embedding contains the features such as the action, demeanor, behavior, and expression required to generate the image, and will be input into the image generation model to guide the model for reasoning (the image generation model here is the same model as the image generation model used for image inversion mentioned in the reference feature extraction step, and the model can be implemented based on the principle of the diffusion model. Reasoning refers to denoising step by step so that an initial Gaussian noise image eventually becomes a complete picture). Finally, an image that matches the text semantics is obtained.
[0024] KL divergence calculation steps: Use KL divergence to calculate the difference between the reference intermediate feature map and the target intermediate feature map. If the difference is greater than the preset threshold, it is considered a valid change and features are extracted from the target text; if the difference is less than the preset threshold (generally one thousandth, statistically it is generally believed that less than one thousandth is considered to be an insignificant difference, so it is considered an invalid change), it is considered an invalid change, features are extracted from the reference image, and the identity of the source image is retained. KL divergence is generally used to measure the difference between two distributions. In the large model reasoning process, TransFormer's multi-head attention mechanism consists of multiple heads, each head consists of three parameters Q, K, and V. After Q and K are multiplied and softmaxed, a square matrix with a sum of 1 can be obtained to represent the area of attention of the attention mechanism. This matrix can be approximately regarded as a distribution. If the difference between the reference feature and the target feature is not large, then there is no need to introduce interference from the target image, so reference feature injection is selected, otherwise target feature injection is selected.
[0025] Figure 2The upper and lower parts represent the values of the reference image and the target image before the multi-head self-attention calculation. There are multiple such values in each time step to pay attention to the features of different areas of the image. The value is passed through three mapping modules ( Figure 2 (not shown) to obtain Q, K, and V. Only Q and K are needed to calculate the difference in KL divergence.
[0026] The multiplication of Q and K gives the attention map, which is a graph where all values add up to 1, representing a sampling of data distribution. The orange Q and K calculate the attention map of the reference image; the gray Q and K calculate the attention map of the target image. The two are calculated using the formula A value is calculated, which indicates the degree of difference between the two distributions. If the difference is less than the threshold, it is considered that the difference is large and the modification is invalid, so the identity feature from the reference image is selected.
[0027] P and Q represent two probability distributions (p(x), Q(x) are P and Q). In the present invention, since the sum of the two feature maps of the attention mechanism is 1, they can be regarded as probability distributions.
[0028] The present invention uses KL divergence to calculate the difference between feature maps. If the difference is greater than a threshold, it is considered a valid change, and features are extracted from the target text. If the difference is less than a threshold, it is considered an invalid change, and features are extracted from the reference image, retaining the identity of the source image and avoiding defects caused by meaningless changes. In addition to being applied in the field of image editing, the present invention can also be used in the field of video generation, so that the difference between video frames is reduced, the subject identity is consistent, and higher fidelity is obtained.
[0029] Image generation step: Merge the obtained features, input the merged features into the large model for inference, generate and output the target image.
[0030] The main body of the present invention is divided into two parts. The first is to obtain the generation path of the real image through the image inversion mechanism, which is equivalent to obtaining an intermediate feature from Gaussian noise to the reference image. The second is to modify the target text to generate the target image. In each step of the reasoning process, the generated attention map and the obtained intermediate feature map are KL divergence calculated. If the requirements are met, it is screened and injected into the reasoning process, otherwise the attention of the reference image is maintained unchanged. Finally, after dozens of time steps, a completed edited target image can be obtained. The subject identity between the target image and the reference image is consistent, but the action, style, etc. are modified by modifying the descriptive text. Figure 4As shown, (a) is a real reference image a shy girl actually input by the user, which is modified into a cool girl by the present invention, as shown in (b).
[0031] The large model of the present invention is a large-scale text-generated image model based on a diffusion model, such as stable diffusion. Taking it as an example, it accepts text and Gaussian noise as input to generate pictures that conform to the target text. The model is a unet structure, and each layer has multiple attention blocks. Each block contains a multi-head attention structure, and each multi-head attention structure contains multiple self-attentions. Therefore, an image generation model contains many self-attention structures, and their values represent the identity and appearance of the image. My approach is to first use DDIM inversion to obtain the values of the self-attentions in the intermediate steps of the reference image, and then use the target text to generate the image. During the generation process, the self-attention at the same time step and the same position is calculated one-to-one with the self-attention obtained by inversion. The attention pair less than the threshold is replaced by the self-attention value obtained by selecting DDIM, otherwise it remains unchanged. In this way of replacement, the features of the reference image are injected into the image generation process, and the image editing task is actually realized. The above is the process of generating pictures from features. The replaced self-attention is called head self-attention. The example figure is as follows Figure 3 If the large model itself does not have sufficient generation capabilities or text understanding capabilities, it will not be able to generate the corresponding images. In specific implementation, auxiliary features such as edge maps, depth maps, and subject mask maps can be introduced to help the large model better complete image editing and obtain higher image fidelity.
[0032] The image editing realization system based on large model feature injection of the present invention comprises a reference feature extraction module, a target feature extraction module, a KL divergence calculation module and an image generation module.
[0033] Reference feature extraction module: performs image inversion based on the real reference image to obtain the reference intermediate feature map.
[0034] Target feature extraction module: extracts features from the target text and aligns them with the image feature space of the reference intermediate feature map. The extracted features are input into the image generation model to guide the model to perform reasoning, and a target intermediate feature map that is consistent with the text semantics is obtained.
[0035] KL divergence calculation module: In the reasoning process of each step of the large model, the KL divergence is used to calculate the difference between the reference intermediate feature map and the target intermediate feature map. If the difference is greater than the preset threshold, it is considered a valid change and features are extracted from the target text; if the difference is less than the preset threshold, it is considered an invalid change and features are extracted from the reference image to retain the identity of the source image.
[0036] Image generation module: Merge the acquired features, input the merged features into the large model for inference, generate and output the target image.
[0037] As an implementation manner, the reference feature extraction module uses an image inversion model to perform image inversion to obtain intermediate features of the reference image, and the image inversion model is composed of DDIM inversion or Null-Text Inversion.
[0038] The reference feature extraction module performs image inversion according to the following steps: The real reference image is reversely calculated through the following formula to get the denoising value of the previous time step, and this value is added to get the intermediate noise map of this time step. This cycle is repeated for T time steps, so that the intermediate features of the reference image are finally output: ; x t−1 represents the noise map at time step t−1, α t is a preset constant, is α t The cumulative value of the multiplication of ϵ θ ( x t , t ) is a trained diffusion model used to predict the noise removed in step t, which depends on the noise map in step t. t is the variance of a Gaussian distribution, and ϵ represents a random noise sampled from a standard normal distribution.
[0039] As an implementation method, the KL divergence calculation module calculates the difference value using the following formula: ; Among them, P and Q represent two probability distributions.
[0040] As an implementation method, the target feature extraction module uses the CLIP model to extract features.
[0041] The present invention measures the difference in attention maps between the reference image and the target image at each time step of the image generated by the large model inference through KL divergence. When the target image is generated, features are selectively extracted from the reference image according to the difference. That is, features that help to ensure that the identity and background remain unchanged are extracted from the reference image, and features such as action and expression are extracted from the target text. The two are combined to obtain a complete feature map, and then the image editing task is completed by modifying the text without fine-tuning the large model or manual intervention.
[0042] The present invention has low time cost, is simple and easy to implement, and has high flexibility, so that the large model can complete the image editing task of retaining identity features in most scenarios.
[0043] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for implementing image editing based on large model feature injection, characterized in that: include: Reference feature extraction step: perform image inversion based on the real reference image to obtain a reference intermediate feature map; Target feature extraction step: extract features from the target text and align them with the image feature space of the reference intermediate feature map. Input the extracted features into the image generation model to guide the model to perform reasoning, and obtain a target intermediate feature map that matches the text semantics. KL divergence calculation step: In the reasoning process of each step of the large model, the KL divergence is used to calculate the difference between the reference intermediate feature map and the target intermediate feature map. If the difference is greater than the preset threshold, it is considered a valid change and features are extracted from the target text; If the difference is less than a preset threshold, it is considered an invalid change, and features are extracted from the reference image, preserving the identity properties of the source image; Image generation step: Merge the obtained features, input the merged features into the large model for inference, generate and output the target image.
2. The image editing implementation method based on large model feature injection as claimed in claim 1, characterized in that: In the reference feature extraction step, an image inversion model is used to perform image inversion to obtain intermediate features of the reference image, and the image inversion model is composed of DDIM inversion or Null-Text Inversion.
3. The image editing implementation method based on large model feature injection as claimed in claim 2 is characterized in that: The image inversion process in the reference feature extraction step is: The real reference image is reversely calculated through the following formula to get the denoising value of the previous time step, and this value is added to get the intermediate noise map of this time step. This cycle is repeated for T time steps, so that the intermediate features of the reference image are finally output: ; x t−1 represents the noise map at time step t−1, α t is a preset constant, is α t The cumulative value of the multiplication of ϵ θ ( x t , t ) is the trained diffusion model, σ t is the variance of a Gaussian distribution, and ϵ represents a random noise sampled from a standard normal distribution.
4. The image editing implementation method based on large model feature injection as claimed in claim 1, characterized in that: The KL divergence calculation step uses the following formula to calculate the difference value: ; Among them, P and Q represent two probability distributions.
5. The image editing implementation method based on large model feature injection as claimed in claim 1, characterized in that: In the target feature extraction step, the CLIP model is used for feature extraction.
6. An image editing implementation system based on large model feature injection, characterized in that: include: Reference feature extraction module: performs image inversion based on the real reference image to obtain the reference intermediate feature map; Target feature extraction module: extracts features from the target text and aligns them with the image feature space of the reference intermediate feature map. The extracted features are input into the image generation model to guide the model to perform reasoning, and a target intermediate feature map that matches the text semantics is obtained. KL divergence calculation module: In the reasoning process of each step of the large model, the KL divergence is used to calculate the difference between the reference intermediate feature map and the target intermediate feature map. If the difference is greater than the preset threshold, it is considered a valid change and features are extracted from the target text; If the difference is less than a preset threshold, it is considered an invalid change, and features are extracted from the reference image, preserving the identity properties of the source image; Image generation module: Merge the acquired features, input the merged features into the large model for inference, generate and output the target image.
7. The image editing implementation system based on large model feature injection as claimed in claim 6, characterized in that: The reference feature extraction module uses an image inversion model to perform image inversion to obtain intermediate features of the reference image, wherein the image inversion model is composed of DDIM inversion or Null-Text Inversion.
8. The image editing implementation system based on large model feature injection as claimed in claim 7, characterized in that: The reference feature extraction module performs image inversion according to the following steps: The real reference image is reversely calculated through the following formula to get the denoising value of the previous time step, and this value is added to get the intermediate noise map of this time step. This cycle is repeated for T time steps, so that the intermediate features of the reference image are finally output: ; x t−1 represents the noise map at time step t−1, α t is a preset constant, is α t The cumulative value of the multiplication of ϵ θ ( x t , t ) is the trained diffusion model, σ t is the variance of a Gaussian distribution, and ϵ represents a random noise sampled from a standard normal distribution.
9. The image editing implementation system based on large model feature injection as claimed in claim 6, characterized in that: The KL divergence calculation module uses the following formula to calculate the difference value: ; Among them, P and Q represent two probability distributions.
10. The image editing implementation system based on large model feature injection as claimed in claim 6, characterized in that: The target feature extraction module uses the CLIP model for feature extraction.