Image editing method and system based on text graph large model
By using a large model based on literary image in image editing, the latent space noise map is optimized and the constraints of attention maps are calculated, the problem of difficult to balance fidelity and editability in the prior art is solved, and efficient image editing effect is achieved.
Patent Information
- Application Number
- CN202510412838.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The prior art is difficult to balance the fidelity and editability of image editing within a unified framework, resulting in the problem of over- or insufficient editing.
The image editing method based on the Wensheng Picture Big Model is adopted, and the latent space noise map is optimized through an adaptive time scheduler, combining the cross attention map and the self-attention map, cross attention consistency constraints and self-attention retention constraints are calculated, and the latent space noise map is updated to achieve balance.
It effectively balances the fidelity and editability of image editing within a unified framework, ensuring that the editing image achieves the target effect and maintains the overall image structure.
Smart Images

Figure CN119941928A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image editing under the field of artificial intelligence technology, and more specifically, to an image editing method and system based on a large model of a Vincent graph. Background Art
[0002] Image editing has always been an important means for people to process images and express aesthetics, but traditional manual editing methods require a lot of time and effort. With the continuous development of computer technology, deep learning technology and large-scale data sets, text-generated images have become a research field that has attracted much attention due to their advantages such as efficiency, flexibility and scalability, providing the possibility for automated image editing.
[0003] In text-based image editing (TIE), it is crucial to balance fidelity and editability, and failure to balance the two often leads to over-editing or under-editing problems. Existing methods usually rely on injecting cross-attention maps or self-attention maps to maintain structure and exploit the inherent text alignment capabilities of pre-trained text-based image models to achieve editability, but they lack a clear and unified mechanism to properly balance these two goals.
[0004] Therefore, how to achieve a balance between fidelity and editability within a unified framework is one of the current research directions in the field of artificial intelligence image editing. Summary of the invention
[0005] In view of the above-mentioned problem that current image editing methods cannot balance editable lines and fidelity within a unified framework, the purpose of the present invention is to provide an image editing method and system, which uses an adaptive time scheduler to optimize the latent space noise map so that the edited image can achieve the target effect while maintaining the overall image structure.
[0006] The image editing method based on the Wensheng graph large model provided by the present invention comprises the following steps: S110: Acquire a latent space noise map of a source image based on the Vincent map input data; wherein the Vincent map input data includes the source image, a mask of a target editing area, a text description of the source image, and a text description of an edited image; S120: performing denoising processing on the latent space noise map, and first performing optimization processing on the latent space noise map within a preset time step of the denoising processing; S130: decoding the latent space noise map obtained in the last step of the denoising process to obtain an edited image; The optimization processing of the latent space noise map includes: S121: embedding the latent space noise map of the current time step and the text of the text description of the edited image into a preset noise prediction network for noise prediction; S122: Extracting the cross attention map and the self-attention map generated in the noise prediction network; S123: Calculate the mask-based cross-attention consistency constraint and the self-attention retention constraint respectively using the cross-attention map and the self-attention map; S124: Calculate the gradients of the cross-attention consistency constraint and the self-attention retention constraint on the latent space noise map respectively through back propagation, and update the latent space noise map using an adaptive time step scheduler and the gradient.
[0007] Among them, an optional solution is that the cross-attention map includes the cross-attention map corresponding to the editing word that describes the edited image, and the self-attention map includes all self-attention maps of a preset resolution generated in the denoising network.
[0008] Among them, an optional solution is that the denoising process of the latent space noise map includes a reconstruction branch and an editing branch; wherein, The reconstruction branch is used to perform noise prediction using a text embedding of a text description of a source image and a latent space noise map at each time step of the denoising process; The editing branch is used to guide the latent space noise map within a preset time step of the denoising process; wherein, In a certain time step of the denoising process, a mask is used to perform a mixing operation on the latent space noise map of the reconstruction branch and the latent space noise map of the editing branch.
[0009] Among them, an optional solution is that the loss function of the self-attention retention constraint is as follows: , in, is the self-attention preservation constraint, is the current time step, is the self-attention map generated in the reconstruction branch, Self-attention map generated in the editing branch; mask The loss function used to calculate the local self-attention preserving constraint to preserve high-frequency details: , in, is the local self-attention preservation constraint, , for The flattened resulting vector.
[0010] Among them, the optional solution is that the loss function of the cross-attention consistency constraint is calculated as follows: , in, is the cross attention consistency constraint, represents the ratio of the cross attention map corresponding to the edited word inside and outside the mask, Represents the total number of cross-attention maps with a resolution of 16X16; where, , in, represents the spatial coordinates of the cross attention map, Indicates the layer number of the cross attention map, Indicates the index corresponding to the edited word. Represents the cross-attention map produced in the editing branch.
[0011] Among them, an optional scheme is to use an adaptive time step scheduler and the gradient to update the latent space noise map, including: based on the gradient of the latent space noise map with respect to the cross-attention consistency constraint and the gradient of the latent space noise map with respect to the self-attention retention constraint, using the adaptive time step scheduler to calculate the total gradient, and updating the latent space noise map according to the total gradient.
[0012] Among them, an optional solution is that the calculation of updating the potential noise map according to the total gradient is as follows: , in, is the updated latent space noise map, is the latent space noise map before updating, is the total gradient; , in, and are adaptive time step schedulers acting on the self-attention preservation constraint and the cross-attention consistency constraint, respectively, Preserve the gradient of the latent space noise map in the edit branch for the self-attention constraint, Gradient of the latent space noise map for cross-attention consistency constraints; , Where T is the total number of steps in the denoising process, is the current time step, , , , Hyperparameters for controlling the adaptive time step scheduler.
[0013] Among them, an optional solution is to use a mask to mix the latent space noise map of the reconstruction branch and the latent space noise map of the editing branch. The formula is as follows: , in, is the latent space noise map after the mixing operation, is the latent space noise map of the editing branch, For the mask, is the latent space noise map of the reconstruction branch.
[0014] The present invention also provides an image editing method based on the Wensheng graph large model, and the image editing is performed using the image editing method based on the Wensheng graph large model as described above, comprising: A latent space noise map acquisition unit, configured to acquire a latent space noise map of a source image based on a Vincent map input data; wherein the Vincent map input data includes the source image, a mask of a target editing area, a text description of the source image, and a text description of an edited image; A denoising unit, used for denoising the latent space noise map; An optimization unit, used for first performing optimization processing of the latent space noise map within a preset time step in the denoising process; A decoding unit, used for decoding the latent space noise map obtained in the last step of the denoising process to obtain an edited image; Wherein, the optimization unit comprises: A noise prediction unit, used for embedding the latent space noise map of the current time step and the text of the text description of the edited image into a preset noise prediction network for noise prediction; A cross-attention map and self-attention map extraction unit, used to extract the cross-attention map and self-attention map generated in the noise prediction network; A constraint calculation unit, used to calculate a mask-based cross-attention consistency constraint and a self-attention preservation constraint using the cross-attention map and the self-attention map respectively; The updating unit calculates the gradients of the cross-attention consistency constraint and the self-attention preservation constraint on the latent noise map, and updates the latent space noise map using the adaptive time scheduler and the gradients.
[0015] The present invention further provides an electronic device, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the steps in the image editing method based on the Wenshengtu model as described above.
[0016] From the above technical scheme, it can be seen that the image editing method and system based on the Vincent graph large model provided by the present invention uses the Vincent graph large model to perform image editing through the user inputting text describing the source image and the edited image and the mask of the target editing area; wherein, the present invention first performs optimization processing of the latent space noise map within a preset time step of the latent space noise map denoising process, and the optimization processing includes: first, inputting the latent space noise map and text prompt of the current step into the noise prediction network; secondly, extracting the cross-attention map and self-attention map generated in the noise prediction network, and using the cross-attention map and the self-attention map to respectively calculate the mask-based cross-attention consistency constraint and the self-attention retention constraint; Finally, the gradients of the cross-attention consistency constraint and the self-attention preservation constraint on the latent space noise map are calculated by back propagation, and the latent space noise map is updated using an adaptive time step scheduler and the gradient. In addition, within a certain time step of the denoising process, the latent space noise map of the reconstruction branch and the latent space noise map of the editing branch can be mixed using a mask, so that the edited image achieves the target effect and maintains the image structure of the background. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By referring to the following description in conjunction with the accompanying drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more clear and easy to understand. In the accompanying drawings: Figure 1 is a flow chart of an image editing method based on a Wenshengtu macro model according to an embodiment of the present invention; Figure 2 A schematic diagram of the principle of an image editing method based on a large model of a Wensheng graph according to an embodiment of the present invention; Figure 3 A schematic diagram of a framework of a system based on a large model of a cultural graph according to an embodiment of the present invention; Figure 4 is a schematic diagram of an electronic device according to an embodiment of the present invention; Figure 5 A schematic diagram showing the comparison of the effects of applying the present invention and other existing image editing solutions; Figure 6 This is an example diagram of the effect of image editing using the image editing method and system based on the Wensheng graph model of the present invention. DETAILED DESCRIPTION
[0018] In the following description, for the purpose of illustration, in order to provide a comprehensive understanding of one or more embodiments, many specific details are set forth. However, it is apparent that these embodiments may also be implemented without these specific details. In other examples, for ease of describing one or more embodiments, known structures and devices are shown in the form of block diagrams.
[0019] In view of the problem that existing image editing solutions cannot unify modeling fidelity and editability, the present invention provides an image editing method and system based on a large Wensheng graph model.
[0020] In order to better illustrate the technical solution of the present invention, some technical terms involved in the present invention are briefly explained below.
[0021] Text-to-image model, a model that generates corresponding images from simple text descriptions (in English). Existing text-to-image models include the large text-to-image model DALL-E2 and Stable Diffusion. Stable Diffusion is a generative large model in the field of computer vision that can perform image generation tasks such as text-to-image (txt2img) and image-to-image (img2img).
[0022] Null-text Inversion is the sequel to prompt-to-prompt, which is used to solve the problem of how to use prompt-to-prompt on real images.
[0023] The noise prediction network is used to process inputs of different modalities using a transformer-structured backbone network to predict noise, such as using a U-net network to predict the noise at each step. U-Net is a semantic segmentation algorithm based on a fully convolutional network that uses a symmetrical U-shaped structure and mirror operations to process boundary pixels.
[0024] Cross attention is a mechanism used in the noise prediction U-net network structure of Stable Diffusion. The idea is to enable one sequence to "pay attention" to another sequence.
[0025] The back-propagation (BP) algorithm calculates the gradient of the loss function of each parameter in the neural network through the derivative chain rule, and updates the parameters with the optimization method to reduce the loss function.
[0026] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0027] It should be noted that the following description of the exemplary embodiments is merely illustrative and is not intended to limit the present invention and its application or use. Technologies and devices known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, the technologies and devices should be considered as part of the specification.
[0028] In order to illustrate the image editing method and system provided by the present invention, Figure 1 , Figure 2 The process and principle of the image editing method based on the Wenshengtu large model according to the embodiment of the present invention are respectively exemplarily indicated; Figure 3 The logical structure of the image editing system based on the Wenshengtu macro model according to the embodiment of the present invention is exemplarily indicated.
[0029] like Figure 1 and Figure 2 As shown in the figure, the image editing method based on the Wenshengtu model provided by the present invention mainly includes the following steps: S110: Acquire a latent space noise map of a source image based on the Vincent map input data; wherein the Vincent map input data includes the source image, a mask of a target editing area, a text description of the source image, and a text description of an edited image; S120: performing denoising processing on the latent space noise map, and first performing optimization processing on the latent space noise map within a preset time step in the denoising processing; S130: Decoding the latent space noise map obtained in the last step of the denoising process to obtain an edited image.
[0030] Specifically, as an example, the large model of the Wensheng graph based on the present invention can adopt the large model provided by Stable-Diffusion, such as stable-diffusion-v1-4. For the convenience of description, in the following embodiments, Stable-Diffusion is used as an implementation example of the large model of the Wensheng graph.
[0031] The above steps of the image editing method based on the Wensheng graph model will be described in more detail below in conjunction with specific embodiments.
[0032] For step S110, the specific method of obtaining the latent space noise map of the source image based on the Vincent map input data is as follows: For the source image input by the user, the mask of the target editing area , a text description of the source image , edit the text description of the image First, the source image can be perceptually compressed by a Stable-Diffusion encoder to obtain a latent space representation of the source image. ; Then, based on the obtained latent space representation of the source image, an image inversion method such as Null-text Inversion is used to obtain the latent space noise map of the source image .
[0033] After obtaining the latent space noise map of the source image, the U-net network of Stable-Diffusion can be used to predict the noise, so as to denoise the latent space noise map of the source image, and optimize the latent space noise map within the preset time step of the denoising process. Among them, the U-net network is a neural network consisting of three parts: decoder, encoder, and bottleneck layer. The encoder extracts features by shrinking the feature map through convolution and pooling. The decoder is symmetrical with the encoder part. After upsampling, the feature map is cascaded with the encoding part to retain the features and then convolution. The bottleneck layer contains two 3 × 3 convolution layers. Finally, a 1×1 convolution layer is passed to obtain the final output.
[0034] In one embodiment of the present invention, the time step of the denoising process is 50 steps, and the preset time step of the optimization process of the latent space noise map is 25 steps. That is, through the 50-step cyclic denoising process, and in the [50,25] time step of the denoising process, the optimization process of the latent space noise map is added; the optimization process is performed first, and the denoising process is performed after the optimization process is completed, so as to better remove the noise in the latent space noise map and better edit the image to achieve the target effect.
[0035] The optimization process proposed in the present invention is equivalent to saying that the denoising process at t = [50, 0] is a large framework of stable diffusion, and the latent space noise map is optimized before the denoising process at t = [50, 25] within this large framework. Specifically, as an example, the optimization process of the latent space noise map includes: S121: embedding the latent space noise map of the current time step and the text of the text description of the edited image into a preset noise prediction network for noise prediction; S122: Extracting the cross attention map and the self-attention map generated in the noise prediction network; S123: Calculate the mask-based cross-attention consistency constraint and the self-attention retention constraint respectively using the cross-attention map and the self-attention map; S124: Calculate the gradients of the cross-attention consistency constraint and the self-attention retention constraint on the latent space noise map respectively through back propagation, and update the latent space noise map using an adaptive time step scheduler.
[0036] The present invention does not specifically limit the number of times the above optimization process is executed, and it can be executed multiple times or once depending on the optimization effect. In a specific embodiment of the present invention, the above optimization process is executed once.
[0037] From the overall framework of the denoising process, the entire denoising process of the present invention includes two branches: a reconstruction branch and an editing branch. In the following embodiments, the latent space noise map of the reconstruction branch is represented by the non-asterisk z Denote that the latent space noise map in the editing branch is expressed as For each step of the denoising process, the reconstruction branch is used to use the text description of the source image Text embedding and latent space noise map of Input U-net network prediction noise for noise prediction, and edit branch is used to loop the potential noise map within a certain time step of the denoising process. Optimization processing.
[0038] In other words, the reconstruction branch and the editing branch are two separate 50-step denoising processes, and the two branches are performed simultaneously. The cyclic optimization process is applied in the editing branch to make the latent space noise map features in the mask more consistent with the edited text, and the structural features of the latent noise map and the structural features of the reconstruction branch more similar.
[0039] More specifically, as an example, the optimization process of the latent space noise map includes the following operations: In the reconstruction branch, the latent space noise map of the current step is and edit the text description of the image The text embedding input U-net network; Extract the resolution generated in the U-net network as Five cross-attention maps with a resolution of , , , The self-attention map or resolution is , , the self-attention map; since the resolution is It can present the most significant semantic features. Therefore, the extraction resolution in this embodiment is Five cross attention maps.
[0040] Calculating a cross-attention consistency constraint and a self-attention preservation constraint using the cross-attention map and the self-attention map; The gradients of the cross-attention consistency constraint and the self-attention preservation constraint on the latent space noise map are calculated by back-propagation, and the latent space noise map is updated using an adaptive time step scheduler. .
[0041] After the denoising process is completed, the latent space noise map obtained in the last step is converted into Decoded into edited image I * .
[0042] exist Figure 1 , Figure 2 In the embodiment shown together, the cross-attention consistency constraint and the self-attention preservation constraint guidance are integrated into the Vincent graph large model, and the optimization of the latent space noise map is introduced in the denoising process without additional training.
[0043] When the denoising process introduces the attention adjustment constraint guidance, at each time step t, the potential noise mask is first Input the noise prediction network and extract all cross-attention maps with a resolution of 16×16. Then use the cross-attention maps corresponding to all edited words to calculate the cross-attention consistency constraint; extract all self-attention maps with a resolution of 8×8, 16×16, 32×32, 64×64 or 8×8, 16×16 to calculate the self-attention preservation constraint.
[0044] Specifically, as an example, given the text description of the source image and edit the text description of the image , define the editing process in The target edit word set that appears in but not in P is ;definition and P The set of words that appear together in S ={ s 1, s 2,... s J In order to alleviate the inaccurate positioning problem of the cross-attention map, the present invention introduces a mask to assist in the editing process.
[0045] The cross-attention map includes the cross-attention map corresponding to the edited word describing the edited image, and the self-attention map includes the self-attention map generated in the denoising network. Accordingly, the cross-attention map is used to calculate the cross-attention consistency constraint, and the self-attention map is used to calculate the self-attention preservation constraint.
[0046] In a specific embodiment of the present invention, the edit word is defined exist t The cross attention map corresponding to the moment is , the attention consistency constraint is used to adjust the cross attention map The value of .
[0047] Specifically, the cross-attention map reflects the consistency of text features and image features. The larger its value, the higher the alignment between text and visual features, thereby improving the editability of the target area. Therefore, the cross-attention map aims to increase the value of the cross-attention map corresponding to the edited word in the mask to enhance the editing effect. Cross-attention consistency constraint The loss function is as follows: , in, is the cross attention consistency constraint, represents the ratio of the cross attention map corresponding to the edited word inside and outside the mask, represents the total number of cross-attention graphs with a resolution of 16X16; in this embodiment, ; , in, represents the spatial coordinates of the cross attention map, Indicates the layer number of the cross attention map, Indicates the index corresponding to the edited word. Represents the cross-attention map produced in the editing branch.
[0048] In order to maintain the structural fidelity, the present invention also proposes a self-attention preservation constraint to reduce the self-attention map generated by the reconstruction branch. Self-attention map generated with the editing branch , and is more flexible than directly injecting the self-attention map. Specifically, as an example, the loss function of the self-attention preservation constraint is as follows: , in, is the self-attention preservation constraint, is the current time step, is the self-attention map generated in the reconstruction branch, is the self-attention map generated in the editing branch.
[0049] In addition, you can further use masks Compute the loss function with local self-attention preserving constraint to preserve high-frequency details:
[0050] in, is the local self-attention preservation constraint, , for Flatten the resulting vector. This approach of preserving high-frequency details is particularly useful when editing smaller objects in complex scenes.
[0051] Finally, in order to strike a balance between editability and fidelity, the present invention also proposes to calculate the total gradient through an adaptive time scheduler, and then the latent space noise map of the mask area content is updated using the total gradient. The adaptive time scheduler is calculated as follows: , The total gradient is calculated as follows: , The potential noise map is updated using the following formula: , In the above formula, the total gradient is subtracted from the masked area of the latent space noise map at a preset learning rate to obtain the updated latent space noise map of the masked area content. is the latent space noise map of the editing branch, is the noise map of the latent space after performing the gradient update in the edit branch, is the total gradient; and are adaptive time step schedulers acting on the self-attention preservation constraint and the cross-attention consistency constraint, respectively, Preserve the gradient of the latent space noise map in the edit branch for the self-attention constraint, is the gradient of the cross-attention consistency constraint on the noise map in the latent space; T is the total number of steps in the denoising process, is the current time step, , , , Hyperparameters for controlling the adaptive time step scheduler.
[0052] The above optimization process uses an adaptive time scheduler to control the balance between self-attention retention constraints and cross-attention consistency constraints, so that the editing effect is balanced between fidelity and editability, and the optimal editing effect is obtained. And by adjusting the parameters, different editing types can be achieved.
[0053] The proportional coefficients β1 and β2 are usually set to 5. In a specific embodiment of the present invention, k1 and k2 are set as follows: color change is 0.05, texture modification or background editing is 0.08, object replacement is 0.15, global style transfer is 0.1, and face attribute editing is 0.25.
[0054] In order to ensure that the image outside the editing area is consistent with the source image, in this embodiment, within a certain time step (e.g., 50 steps) of the denoising process, a mask is further used to perform a blending operation (blend) on the latent space noise map of the reconstruction branch and the latent space noise map of the editing branch as shown in the following formula: in, is the latent space noise map after the mixing operation, is the latent space noise map of the editing branch, For the mask, is the latent space noise map of the reconstruction branch.
[0055] like Figure 3 As shown, the present invention also provides an image editing system based on the Vincent image model, which uses the image editing method based on the Vincent image model as described above to perform image editing. The image editing system 300 based on the Vincent image model mainly includes three parts: a latent space noise map acquisition unit 310, an optimization unit 320, a denoising unit 330 and a decoding unit 340.
[0056] The latent space noise map acquisition unit 310 is used to acquire the latent space noise map of the source image based on the Wensheng map input data; wherein the Wensheng map input data includes the source image, the mask of the target editing area, the text description of the source image and the text description of the editing image; A denoising unit 330, configured to perform denoising on the latent space noise map; An optimization unit 320, configured to first perform optimization processing on the latent space noise map within a preset time step in the denoising processing performed by the denoising unit; The decoding unit 340 is used to decode the latent space noise map obtained in the last step of the denoising process to obtain an edited image.
[0057] The optimization unit 320 includes: A noise prediction unit 321 is used to embed the latent space noise map of the current time step and the text of the text description of the edited image into a preset noise prediction network for noise prediction; A cross-attention map and self-attention map extraction unit 322, used to extract the cross-attention map and the self-attention map generated in the noise prediction network; A constraint calculation unit 323 is used to calculate the mask-based cross-attention consistency constraint and the self-attention preservation constraint respectively using the cross-attention map and the self-attention map; The updating unit 324 is used to calculate the gradients of the cross-attention consistency constraint and the self-attention preservation constraint on the latent space noise map through back propagation, and update the latent space noise map using an adaptive time step scheduler and the gradient.
[0058] The above-mentioned image editing method is an implementation method corresponding to the above-mentioned image editing system. Its specific execution steps can refer to the specific embodiment of the above-mentioned image editing system, and will not be described in detail here.
[0059] It can be seen from the above embodiments that the image editing method and system based on the Vincent graph large model proposed by the present invention uses the Vincent graph large model to perform image editing by inputting text describing the source image and the edited image and the mask of the target edited area by the user. The latent space noise map is first optimized within the preset time step of the latent space noise map denoising process; including: first, the latent space noise map and text prompt of the current step are input into the noise prediction network; secondly, the cross attention map and self-attention map generated in the noise prediction network are extracted, and the cross attention map is used to calculate the cross attention consistency constraint, and the self-attention map is used to calculate the self-attention retention constraint; finally, the gradients of the two constraints on the latent space noise map are calculated respectively through back propagation, and the adaptive time step scheduler calculates the total gradient, and then the latent space noise map is updated using this total gradient. In addition, within a certain time step of the denoising process, the latent space noise map of the reconstruction branch and the latent space noise map of the editing branch are mixed using a mask. In this way, the edited image achieves the target effect and maintains the overall image structure.
[0060] Figures 5 and 6 The image editing effects of the image editing method based on the Wensheng graph model of the present invention are shown respectively. Figure 5 This is a schematic diagram comparing the effects of applying the present invention and other existing image editing solutions. Figure 6 The present invention provides an example of the effect of image editing using the image editing method and system based on the Wensheng graph model of the present invention.
[0061] like Figure 5 As shown, the colored part of the text on the leftmost image is the edit word (only appears in the text describing the edit image). ), the black part is the edited object (in the text describing the edited image and text describing the original image ).
[0062] The first column of images is the original image, the last column of images "ours" represents the result of image editing using the image editing method and system based on the Wensheng graph model of the present invention, and the second to ninth columns are the result of other comparison schemes. Other comparison schemes are as follows: DiffEdit: Paper Diffusion-based semantic image editing with maskguidance; P2P: Paper Prompt-to-prompt image editing with cross attention control; P2P+Bend: Combines hybrid operation with P2P; PnP: Plug-and-play diffusion features for text-driven image-to-imagetranslation; PnP+Blend: Combines blending operations with PnP.
[0063] SPD Inv: Paper Source Prompt Disentangled Inversion for Boosting ImageEditability with Diffusion Models; GR: Paper Guide-and-Rescale: Self-Guidance Mechanism for EffectiveTuning-Free Real Image Editing; MAG-Edit: Paper MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted Guidance.
[0064] exist Figure 6 In the figure, every 5 images form a group. The leftmost image is the original image, and the five images on the right are the editing effect images of the image editing method and system based on the Wensheng graph model of the present invention. The text below each image on the right is the editing word (only appears in the text describing the edited image). ) and edit objects (in the text describing the edited image and text describing the original image Phrases composed of (appear in ).
[0065] like Figure 4 As shown, the present invention also provides an electronic device, the electronic device comprising: at least one processor; and, a memory communicatively connected to at least one processor; wherein, The memory stores a computer program executable by at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the steps in the aforementioned image editing method based on the Wenshengtu large model.
[0066] Those skilled in the art will appreciate that the structure shown in the figure does not limit the electronic device 1 and may include fewer or more components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0067] For example, although not shown, the electronic device 1 may also include a power source (such as a battery) for supplying power to various components. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0068] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0069] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.
[0070] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0071] The image editing program 12 based on the Wenshengtu model stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, the following steps can be implemented: S110: Acquire a latent space noise map of a source image based on the Vincent map input data; wherein the Vincent map input data includes the source image, a mask of a target editing area, a text description of the source image, and a text description of an edited image; S120: performing denoising processing on the latent space noise map, and first performing optimization processing on the latent space noise map within a preset time step in the denoising processing; S130: decoding the latent space noise map obtained in the last step of the denoising process to obtain an edited image; The optimization processing of the latent space noise map includes: S121: embedding the latent space noise map of the current time step and the text of the text description of the edited image into a preset noise prediction network for noise prediction; S122: Extracting the cross attention map and the self-attention map generated in the noise prediction network; S123: Calculate the mask-based cross-attention consistency constraint and the self-attention retention constraint respectively using the cross-attention map and the self-attention map; S124: Calculate the gradients of the cross-attention consistency constraint and the self-attention preservation constraint to the latent space noise map through back propagation, and update the latent space noise map using an adaptive time step scheduler and the gradient.
[0072] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figure 4 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0073] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0074] As described above, the image editing method and system based on the Vincent map model proposed in the present invention are described by way of example with reference to the accompanying drawings. However, those skilled in the art should understand that various improvements can be made to the image editing method and system based on the Vincent map model proposed in the present invention without departing from the content of the present invention. Therefore, the protection scope of the present invention should be determined by the content of the attached claims.
Claims
1. An image editing method based on a large model of Wensheng graph, characterized in that: The steps include: S110: Acquire a latent space noise map of a source image based on the Vincent map input data; wherein the Vincent map input data includes the source image, a mask of a target editing area, a text description of the source image, and a text description of an edited image; S120: performing denoising processing on the latent space noise map, and first performing optimization processing on the latent space noise map within a preset time step of the denoising processing; S130: decoding the latent space noise map obtained in the last step of the denoising process to obtain an edited image; The optimization processing of the latent space noise map includes: S121: embedding the latent space noise map of the current time step and the text of the text description of the edited image into a preset noise prediction network for noise prediction; S122: Extracting the cross attention map and the self-attention map generated in the noise prediction network; S123: Calculate the mask-based cross-attention consistency constraint and the self-attention retention constraint respectively using the cross-attention map and the self-attention map; S124: Calculate the gradients of the cross-attention consistency constraint and the self-attention retention constraint on the latent space noise map respectively through back propagation, and update the latent space noise map using an adaptive time step scheduler and the gradient.
2. The image editing method based on the Wensheng graph model as claimed in claim 1, characterized in that: The cross-attention map includes the cross-attention map corresponding to the edited word-units describing the edited image, and the self-attention map includes all self-attention maps of a preset resolution generated in the denoising network.
3. The image editing method based on the Wensheng graph model as claimed in claim 2, characterized in that: The denoising process for the latent space noise map includes a reconstruction branch and an editing branch; wherein, The reconstruction branch is used to perform noise prediction using a text embedding of a text description of a source image and a latent space noise map at each time step of the denoising process; The editing branch is used to guide the latent space noise map within a preset time step of the denoising process; wherein, In a certain time step of the denoising process, a mask is used to perform a mixing operation on the latent space noise map of the reconstruction branch and the latent space noise map of the editing branch.
4. The image editing method based on the Wensheng graph model as claimed in claim 3, characterized in that: The loss function of the self-attention preservation constraint is as follows: , in, is the self-attention preservation constraint, is the current time step, is the self-attention map generated in the reconstruction branch, Self-attention map generated in the editing branch; mask The loss function used to calculate the local self-attention preserving constraint to preserve high-frequency details: , in, is the local self-attention preservation constraint, , for The flattened resulting vector.
5. The image editing method based on the Wensheng graph model as claimed in claim 4, characterized in that: The loss function of the cross-attention consistency constraint is calculated as follows: , in, is the cross attention consistency constraint, represents the ratio of the cross attention map corresponding to the edited word inside and outside the mask, Represents the total number of cross-attention maps with a resolution of 16X16; where, , in, represents the spatial coordinates of the cross attention map, Indicates the layer number of the cross attention map, Indicates the index corresponding to the edited word. Represents the cross-attention map produced in the editing branch.
6. The image editing method based on the Wensheng graph model as claimed in claim 5, characterized in that: The latent space noise map is updated using an adaptive time step scheduler and the gradient, including: based on the gradient of the latent space noise map with respect to the cross-attention consistency constraint and the gradient of the latent space noise map with respect to the self-attention retention constraint, the total gradient is calculated using the adaptive time step scheduler, and the latent space noise map is updated according to the total gradient.
7. The image editing method based on the Wensheng graph model as claimed in claim 6, characterized in that: The calculation to update the latent space noise map according to the total gradient is as follows: , in, is the updated latent space noise map, is the latent space noise map before updating, is the total gradient; , in, and are adaptive time step schedulers acting on the self-attention preservation constraint and the cross-attention consistency constraint, respectively, Preserve the gradient of the latent space noise map in the edit branch for the self-attention constraint, Gradient of the latent space noise map for cross-attention consistency constraints; , Where T is the total number of steps in the denoising process, is the current time step, , , , Hyperparameters for controlling the adaptive time step scheduler.
8. The image editing method based on the Wensheng graph model as claimed in claim 6, characterized in that: The formula for mixing the latent space noise map of the reconstruction branch and the latent space noise map of the editing branch using a mask is as follows: , in, is the latent space noise map after the mixing operation, is the latent space noise map of the editing branch, For the mask, is the latent space noise map of the reconstruction branch.
9. An image editing system based on a large model of a Vincent graph, which performs image editing based on the image editing method based on a large model of a Vincent graph as claimed in any one of claims 1 to 8, comprising: A latent space noise map acquisition unit, configured to acquire a latent space noise map of a source image based on a Vincent map input data; wherein the Vincent map input data includes the source image, a mask of a target editing area, a text description of the source image, and a text description of an edited image; A denoising unit, used for denoising the latent space noise map; An optimization unit, used for first performing optimization processing of the latent space noise map within a preset time step in the denoising process; A decoding unit, used for decoding the latent space noise map obtained in the last step of the denoising process to obtain an edited image; Wherein, the optimization unit comprises: A noise prediction unit, used for embedding the latent space noise map of the current time step and the text of the text description of the edited image into a preset noise prediction network for noise prediction; A cross-attention map and self-attention map extraction unit, used to extract the cross-attention map and self-attention map generated in the noise prediction network; A constraint calculation unit, used to calculate a mask-based cross-attention consistency constraint and a self-attention preservation constraint using the cross-attention map and the self-attention map respectively; An updating unit is used to calculate the gradients of the cross-attention consistency constraint and the self-attention preservation constraint on the latent space noise map respectively through back propagation, and update the latent space noise map using an adaptive time step scheduler and the gradient.
10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the steps in the image editing method based on the Vincent map model as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Image fine-grained editing method and system based on text graph large model
CN117808926A
Language tracking image editing method based on text and graph generation model
CN117934657A
Dynamic image editing method based on diffusion model
CN118608660A
Video generation method and system based on large-scale text video model
CN119031209A
Image editing method and device based on text guidance, equipment and medium
CN119516038A