A text-guided image editing method, device, and medium based on a diffusion model

CN122574141APending Publication Date: 2026-08-14SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

第一,现有反演方法在采样步数受限的条件下容易产生潜在表示偏差,进而导致重建结果与输入图像之间在纹理细节、局部结构和色彩一致性方面出现误差,难以同时兼顾反演精度与编辑效率

Benefits of technology

(1)本发明在反演阶段引入再噪声化和阶段性结构化结构损失的联合机制,有效提升了少步条件下潜变量轨迹的准确性与稳定性。相较于传统反演方式,可显著改善细节保留与全局结构一致性问题,降低反演误差累积,获得更高质量的图像重建结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574141A_ABST
    Figure CN122574141A_ABST
Patent Text Reader

Abstract

This invention relates to a text-guided image editing method, device, and medium based on a diffusion model, comprising: acquiring an original image, an original prompt word, and an editing prompt word; encoding the original image, the original prompt word, and the editing prompt word to obtain a first image latent variable, a prompt word text embedding, and an editing prompt word text embedding, respectively; based on the prompt word text embedding, diffusing and adding noise to the first image latent variable, and optimizing the diffused and denoised image latent variable by combining re-noiseing and staged structuring loss to obtain a second image latent variable; denoising the second image latent variable based on the editing prompt word text embedding to obtain a third image latent variable; and decoding the third image latent variable to obtain the edited output image. Compared with the prior art, this invention, while ensuring inversion efficiency and accuracy, achieves effective injection of editing semantics, stable protection of non-target regions, and natural transition of region boundaries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image editing, and in particular to a text-guided image editing method, device, and medium based on a diffusion model. Background Technology

[0002] With the rapid growth of AI-generated technologies and the demand for visual content production, text-guided diffusion models have gained widespread attention in the field of image generation and editing, and are gradually being applied to scenarios such as film and television post-production, game art, advertising design, and digital creation. This type of technology can control the semantics of images based on text prompts, generating new images or performing local replacements, attribute modifications, and style adjustments on existing images. In image editing tasks, to maintain the structural and visual consistency of the original image while achieving the target semantic modification, existing methods typically employ an "inversion-regeneration" approach. This involves first mapping the input image to the latent space of the diffusion model, and then introducing textual conditions for editing control during the denoising process. Simultaneously, related research has also incorporated mechanisms such as classifier-free guidance, attention constraints, and mask control to improve editing controllability and result quality, thus promoting the development of text-guided image editing technology.

[0003] Nevertheless, this technology currently has the following drawbacks: First, existing inversion methods are prone to potential representation bias under the condition of limited sampling steps, which leads to errors between the reconstruction results and the input image in terms of texture details, local structure and color consistency, making it difficult to simultaneously achieve inversion accuracy and editing efficiency.

[0004] Second, in the condition-guided phase, some methods typically require the introduction and optimization of empty text embedding to enhance generation stability and editing controllability. However, this strategy significantly increases the additional optimization burden and time cost, raising the overall process complexity and hindering practical deployment and interactive applications.

[0005] Third, during the editing process, existing methods still lack the ability to distinguish between the target editable area and the non-target area, which can easily lead to semantic disturbance spillover, causing the background or the main non-editable area to be mistakenly modified, thus destroying the original structural stability and visual coherence.

[0006] Therefore, how to achieve effective injection of editing semantics, stable protection of non-target regions, and natural transition of region boundaries while ensuring inversion efficiency and accuracy remains a key technical problem that urgently needs to be solved in this field. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a text-guided image editing method, device and medium based on a diffusion model. This method achieves effective injection of editing semantics, stable protection of non-target areas and natural transition of area boundaries while ensuring inversion efficiency and accuracy.

[0008] The objective of this invention can be achieved through the following technical solutions: According to a first aspect of the present invention, a text-guided image editing method based on a diffusion model is provided, comprising: S1. Obtain the original image, original prompt, and edit prompt; S2. Encode the original image, the original prompts, and the edit prompts to obtain the latent variables of the first image, the text embeddings of the prompts, and the text embeddings of the edit prompts, respectively. S3, Inversion Stage: Based on the text embedding of prompt words, the first image latent variables are diffused and denoised, and the diffused and denoised image latent variables are optimized by combining re-noiseing and structural loss to obtain the second image latent variables; wherein, the stage is a staged structural loss, which is only introduced in the previous denoising time step; S4. Denoising stage: Based on the text embedding of editing prompt words, the second image latent variables are denoised to obtain the third image latent variables; S5. Decode the latent variables of the third image to obtain the edited output image.

[0009] Preferably, fixed-point iteration is used for re-noiseing, and the calculation expression is: , in, , These represent the times within time step t. Second and third The latent variables of the image output at the next iteration; Let be the latent variables of the image output after the (t-1)th time step iteration; To introduce noise; c is for conditional embedding of cue text; The latent variable signal retention coefficient at time step t is used to control the influence of the latent variable from the previous time step on the current latent variable update result. This represents the noise injection coefficient at time step t, used to control the intensity of random noise during the diffusion process; This represents the noise scale adjustment coefficient at time step t, used to adjust the amplitude of the random noise term.

[0010] Preferably, the structured loss is calculated using the following expression: , In the formula: For structured loss; This represents the reconstructed image obtained after denoising and decoding of the latent variables of the image at the current time step; Original image; It is a feature mapping function; This is a structural similarity function.

[0011] Preferably, in step S4, the second image latent variable is denoised based on the text embedding of the editing prompt word to obtain the third image latent variable, specifically including: S401. Calculate the cross-attention map based on the pixel query Q corresponding to the latent variables of the second image and the text embedding of the editing prompt word K. ,in, The feature dimension of the second image latent variable; S402. Optimize the cross-attention map A based on the edge-aware attention mechanism to obtain the latent variables of the third image.

[0012] Preferably, in step S402, the optimization of the cross-attention map A based on the edge-aware attention mechanism specifically involves: Cross-attention map Perform gradient calculation to obtain the horizontal gradient. and vertical gradient ; Based on the horizontal gradient corresponding to each pixel and vertical gradient Calculate the gradient magnitude ; Calculate the local difference information for each pixel. ; Based on gradient magnitude and local difference information Attention map Divided into enhanced areas and inhibition region , represented as: , ,in, The gradient strength threshold, For pixels The corresponding gradient magnitude, For pixels Corresponding local difference information; In the expansion edge mask Under constraints, for cross-attention graphs Modulation is performed, resulting in a modulated cross-attention map. Represented as: , Among them, the first coefficient Used to enhance the area of ​​heightened attention, second coefficient Used to suppress areas of decreased attention.

[0013] Preferably, the calculation of local difference information corresponding to each pixel... The calculation expression is: , In the formula: For pixels Corresponding local difference information; For pixels The corresponding value; In pixels The average value of the local neighborhood centered on the center.

[0014] Preferably, the dilated edge mask The acquisition process is as follows: From the original image Edge information is extracted, and the horizontal and vertical gradients are calculated using the Sobel operator. ,in and These represent Sobel convolution kernels; Based on the horizontal gradient and vertical gradient Calculate edge amplitude ; After thresholding and dilution of the edge amplitudes, the resolution is dynamically adjusted according to the number of denoising steps t to match the attention map size, resulting in the dilated edge mask. The dilation scale gradually decreases as the number of denoising steps increases.

[0015] Preferably, in the noise reduction stage, the empty text embedding is replaced with the text embedding of the editing prompt word.

[0016] Compared with the prior art, the present invention has the following advantages: (1) This invention introduces a joint mechanism of re-noiseing and staged structuring loss in the inversion stage, which effectively improves the accuracy and stability of latent variable trajectories under the condition of few steps. Compared with the traditional inversion method, it can significantly improve the problems of detail preservation and global structure consistency, reduce the accumulation of inversion error, and obtain higher quality image reconstruction results.

[0017] (2) The present invention replaces the optimization of empty text embedding in the inversion stage with text embedding, which reduces the time and memory overhead caused by stepwise optimization and improves the efficiency of image editing while ensuring the quality of image editing.

[0018] (3) An edge-aware attention mechanism is introduced into the editing process, using edge information and attention gradients to finely constrain the editing area. Compared with traditional methods, this reduces computational costs, improves the accuracy of target area localization, suppresses boundary blurring and structural shift, and enhances the stability and naturalness of the editing results. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a detailed diagram of the text-guided editing process based on the diffusion model in the embodiment. Figure 3 This is a structural diagram of the model of the present invention.

[0020] Figure 4 Flowchart for re-noiseing combined with structured loss.

[0021] Figure 5 A schematic diagram illustrating the specific implementation of optimizing the cross-attention graph based on the edge-aware attention mechanism.

[0022] Figure 6 Flowchart for replacing empty text embeddings with text embeddings. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] Example like Figure 1 and Figure 2 As shown, this embodiment provides a text-guided image editing method based on a diffusion model, which includes: S1. Obtain the original image, original prompt, and edit prompt; S2. Encode the original image, the original prompts, and the edit prompts to obtain the latent variables of the first image, the text embeddings of the prompts, and the text embeddings of the edit prompts, respectively. S3. Inversion Stage: Based on the prompt word text embedding, the first image latent variables are diffused and noise-added, and then optimized using re-noiseing and structure loss to obtain the second image latent variables, such as... Figure 3 As shown; where the structured loss is a staged structured loss, introduced only at the initial noisy time step. Through multiple iterations and fusion of intermediate results, the accumulation of linear approximation error is reduced.

[0025] (1) Re-noiseing of image latent variables The general form of diffusion inversion is: , Where c represents the conditional embedding of the prompt text; The latent variable signal retention coefficient at time step t is used to control the influence of the latent variable from the previous time step on the current latent variable update result. This represents the noise injection coefficient at time step t, used to control the intensity of random noise during the diffusion process; This represents the noise scale adjustment coefficient at time step t, used to adjust the amplitude of the random noise term.

[0026] Due to the introduction of noise rely The above equation is implicit.

[0027] Traditional approximation is adopted: , In this embodiment, to improve the accuracy of the few-step inversion, fixed-point iteration is used for re-noiseing processing, and the calculation expression is as follows: , in, , These represent the times within time step t. Second and third The latent variables of the image output at the next iteration; Let be the latent variables of the image output after the (t-1)th time step iteration; To introduce noise; c is for conditional embedding of prompt text;

[0028] (2) Staged structured loss like Figure 4 As shown, structural loss is introduced in the early critical step of noise addition. In this embodiment, the expression for calculating the staged structural loss is as follows: , In the formula: For structured loss; This represents the reconstructed image obtained after denoising and decoding of the latent variables of the image at the current time step; Original image; It is a feature mapping function; This is a structural similarity function.

[0029] S4. Denoising Stage: Based on the text embedding of editing prompts, the second image latent variables are denoised to obtain the third image latent variables, specifically including: S401. Calculate the cross-attention map based on the pixel query Q corresponding to the latent variables of the second image and the text embedding of the editing prompt word K. ,in, represents the feature dimension of the second image's latent variables.

[0030] S402. Optimize the cross-attention map A based on the edge-aware attention mechanism to obtain the latent variables of the third image.

[0031] In this embodiment, the cross-attention graph A is optimized based on an edge-aware attention mechanism, such as... Figure 5 As shown, specifically: 1) Cross-attention map Perform gradient calculation to obtain the horizontal gradient. and vertical gradient .

[0032] 2) Based on the horizontal gradient corresponding to each pixel and vertical gradient Calculate the gradient magnitude .

[0033] 3) To characterize the local variation trend of the attention map, define the local difference information corresponding to each pixel. The calculation expression is: , In the formula: For pixels Corresponding local difference information; For pixels The corresponding value; In pixels The average value of the local neighborhood centered on the center.

[0034] 4) Based on gradient magnitude and local difference information Attention map Divided into enhanced areas and inhibition region , represented as: , , In the formula: This is the gradient strength threshold, used to filter out noise or weak response regions, ensuring that modulation is applied only to regions with significant attention changes. For pixels The corresponding gradient magnitude, For pixels The corresponding local difference information.

[0035] 5) In the expansion edge mask Under constraints, for cross-attention graphs Modulation is performed, resulting in a modulated cross-attention map. Represented as: , Among them, the first coefficient Used to enhance the area of ​​heightened attention, second coefficient Used to suppress areas of decreased attention.

[0036] In this embodiment, the dilated edge mask The acquisition process is as follows: 1) From the original image Edge information is extracted, and the horizontal and vertical gradients are calculated using the Sobel operator. ,in and These represent Sobel convolution kernels; 2) Based on the horizontal gradient and vertical gradient Calculate edge amplitude .

[0037] 3) After thresholding and dilution of the edge amplitude, the resolution is dynamically adjusted according to the number of denoising steps t to match the size of the attention map, thus obtaining the dilated edge mask. The expansion scale is gradually reduced as the number of denoising steps increases. In the early stages of denoising, a larger expansion scale is used to cover a wider context area. As the number of denoising steps increases, the expansion scale is gradually reduced so that the control area gradually concentrates on the local boundary, thus ensuring the accuracy of local modifications.

[0038] To ensure the quality of reconstruction and editing, text embedding is used in the denoising stage of this embodiment. Avoid optimizing empty text embeddings ,like Figure 6 As shown, this avoids a complex embedding optimization process. Next, to explain the rationale behind this alternative strategy, this embodiment will analyze it from the mathematical form of the diffusion process.

[0039] The inversion process of the denoising diffusion implicit model DDIM can be expressed as: , In the formula: The latent variable of the image output at the t-th step of denoising is represented as the inversion trajectory; The cumulative product of the noise scheduling coefficients at step t of the denoising process.

[0040] In empty text inversion NTI, the sampling process introduces an optimized empty text embedding, which takes the following form: , In the formula: The optimized image latent variable trajectory.

[0041] If we approximate the optimized image latent variable trajectory to the inverted trajectory, that is... Then we can obtain: , In the formula: The coefficient term is determined by noise scheduling.

[0042] Within the inversion framework, the accuracy of noise prediction and the stability of image latent variable trajectories are significantly improved by introducing an iterative noise addition mechanism and a staged structured loss constraint. Therefore, the additional correction effect of empty text embedding on noise prediction is significantly weakened. Based on this, the approximation is: .

[0043] Furthermore, since the diffusion process exhibits smoothness and continuity between adjacent time steps, it can be further approximated as follows: , Under the above approximation conditions, the residual term approaches zero, thus yielding: , This derivation shows that, given sufficiently accurate inversion trajectories, the impact of optimizing empty text embeddings on updating latent image variables is negligible. Therefore, replacing optimized empty text embeddings with edit prompt text embeddings between reconstruction and editing will not disrupt the consistency of latent image variable trajectories. This invention no longer optimizes empty text embeddings. In contrast, text embedding is used directly during the reconstruction / editing stage to reduce latency and memory usage.

[0044] S5. Decode the latent variables of the third image to obtain the edited output image.

[0045] To verify the performance of this invention, experiments were conducted on the publicly available PIE-Bench dataset, and quantitative and qualitative comparisons were performed with several mainstream image editing methods. To ensure the representativeness and fairness of the evaluation, 100 sets of samples were randomly selected from the PIE-Bench dataset for testing. Each set of samples contained an initial image, a corresponding initial prompt, an editing prompt, and an editing mask; all images had a uniform resolution of 512×512. During the inversion and sampling stages, the same total diffusion steps and consistent sampling settings were used for all methods to ensure the fairness and repeatability of the comparison process.

[0046] In comparative experiments, this invention is compared with several mainstream image editing and inversion methods, including DDIMInversion, Null-Text Inversion (NTI), Negative Prompt Inversion (NPI), StyleDiffusion (SD), and Direct Inversion (DI). For the comparison methods, reproduction and parameter configuration are strictly performed according to their original papers or official implementations to ensure that their performance is fully utilized. Among them, DDIMInversion serves as the basic inversion method, while NTI and NPI correct the inversion trajectory through embedding optimization or negative prompt word mechanisms. StyleDiffusion emphasizes style-guided editing control, while Direct Inversion employs a direct latent variable mapping strategy. These methods cover representative technical routes in the current image editing field, enabling the verification of the advantages of the proposed method in terms of inversion accuracy and editing stability from different perspectives.

[0047] Table 1 Comparison of Image Reconstruction Performance Table 1 compares and analyzes the various methods based on multiple metrics, including structural consistency, semantic alignment, pixel-level fidelity, and perceptual quality. Results for Structure Dis and LPIPS are presented as ×10⁻¹⁰. 3 Formal report (lower values ​​are better, ↓); CLIP Sim (CLIP similarity) and PSNR are better the higher the values ​​(↑). Bold indicates the best performance in that column. Overall, different methods show significant differences in reconstruction capabilities.

[0048] As shown in the table, DDIM's insufficient utilization of structural and edge information results in a low structural consistency index and limited spatial detail recovery capabilities. NTI balances semantic and structural recovery to some extent by optimizing empty text embedding, but due to the lack of explicit structural awareness constraints, local details still exhibit degradation. SD emphasizes global style consistency and lacks structural constraint mechanisms, thus its performance in detail recovery and reconstruction stability is average. NPI improves editing effects through negative cue words, but its capabilities in fine-grained texture and boundary recovery are limited. DI performs well in semantic preservation, but still falls short in the fine reconstruction of complex boundaries and structural regions.

[0049] To address the aforementioned issues, this invention improves both the inversion and editing stages. In the inversion stage, an iterative noise-adding mechanism and structured loss are introduced to enhance the latent representation's ability to preserve details and structural information. In the editing stage, text embedding replaces empty text embedding, and an edge-guided attention control mechanism strengthens the constraint capability of structural regions. As shown in Table 1, this invention achieves superior performance across all evaluation metrics, validating the effectiveness of the proposed improvement strategy.

[0050] Table 2 Comparison of running times for different methods Table 2 presents a comparison of the runtime of different methods in image reconstruction tasks. The numerical values ​​represent the average actual runtime (in seconds) per image for the complete process (inversion + sampling) under the same hardware and experimental settings; lower values ​​are better. The comparison methods include NPI, SD, DI, NTI, and the method of this invention. While NPI has the shortest runtime in the inversion stage, its reconstruction accuracy in detailed regions is relatively low. DI has certain advantages in reconstruction quality, but its computational cost is relatively high.

[0051] This invention, due to the introduction of iterative noise addition, structural loss, and edge-guided attention optimization mechanisms, has a slightly higher computation time than DI. However, by employing text embedding instead of empty text embedding optimization strategies, the additional optimization overhead is effectively reduced, keeping the overall efficiency within an acceptable range. Simultaneously, significant improvements are achieved in detail recovery and boundary preservation, realizing a good balance between reconstruction accuracy and computational efficiency.

[0052] Table 3 Ablation Experiment Setup and Evaluation Indicators Table 3 shows the reconstruction performance results under different inversion steps, denoising steps, re-noiseing iterations, and structure loss step settings. Inversion steps and inference steps represent the number of DDIM steps used for inversion and sampling, respectively; re-noiseing steps represent the number of re-noiseing iterations; structure loss steps represent the number of steps in which structure loss is applied in the early stages of noise addition; “–” indicates that it was not used. The right-hand column of the table shows the PSNR (↑), LPIPS (↓), and L2 (↓) metrics, with bold text indicating the best results for that metric.

[0053] Experimental results show that an appropriate number of re-noiseing iterations can significantly improve image reconstruction accuracy, while excessive iterations lead to image quality degradation, indicating that the number of iterations needs to be reasonably controlled, otherwise performance degradation may occur. Furthermore, introducing an appropriate structural loss term can further enhance image quality and detail recovery capabilities, especially in textured regions and structural boundaries. It should be noted that the number of inversion steps and denoising steps remained fixed in all experimental configurations. The structural loss is applied only in the early part of the noise addition stage, triggering a small amount of additional U-Net computation, thus slightly increasing the overall computational cost, but without changing the definition of the number of inversion steps or denoising steps. Despite the slight time overhead, the structural loss effectively guides the latent representation to converge to the original image structure in the early stages of inversion, thereby significantly improving reconstruction quality and edge detail preservation.

[0054] Furthermore, experimental results show that appropriately reducing the number of denoising steps within a certain range does not significantly affect the reconstruction quality, but can speed up the image editing process. This characteristic is particularly important when the same inversion result is reused multiple times for different editing tasks, and can significantly improve overall efficiency. Therefore, a reasonable combination of inversion steps, denoising steps, re-noiseing iterations, and structural loss steps can achieve a more efficient image editing workflow while ensuring high reconstruction accuracy and perceptual consistency.

[0055] Table 4 Module Ablation Experiment To further validate the above visual observations, Table 4 reports quantitative metrics such as PSNR, LPIPS, and L2. Each row shows the baseline (DDIM) and variations with progressively introduced re-noiseing, structured loss, and edge-aware attention modulation. Each column provides PSNR (↑), LPIPS (↓), and L2 (↓) metrics, where higher PSNR and lower LPIPS / L2 indicate better reconstruction quality. The results show that the DDIM baseline performs poorly across all metrics, reflecting a large inversion error and limited reconstruction fidelity. Introducing the re-noiseing mechanism significantly improves overall reconstruction fidelity and reduces inversion error. Further adding structured loss constraints further improves texture restoration and local structural consistency. Finally, combining this with the edge-aware attention optimization module, the reconstructed image achieves optimal edge and detail restoration, with the overall effect closest to the original image. Therefore, the ablation experiments clearly validate the independent contributions of each module: re-noiseing significantly improves inversion accuracy, structured loss enhances local texture and structural consistency, and edge-aware attention modulation further strengthens boundary and detail restoration capabilities. The three elements work together to achieve a good balance between structural consistency and detail fidelity, resulting in high-quality image reconstruction results.

[0056] Table 5 User Perception Experiment Furthermore, user perception experiments were conducted on 30 high-resolution natural images. Thirty participants rated the editing results of six methods in random order. Evaluation metrics included semantic fidelity, which measures the consistency between the edited image and the target text cue while preserving non-target content; and visual naturalness, which assesses the realism, coherence, and absence of obvious artifacts in the edited image. The average scores for all methods are summarized in Table 5, using a 5-point scale (1 – worst, 5 – best). The columns in the table represent semantic fidelity, visual naturalness, and the overall score (the average of the two). Experimental results show that the present invention outperforms all compared methods on both evaluation metrics, further validating its effectiveness and superior perceptual quality in real-world image editing tasks.

[0057] The electronic device of this invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0058] Multiple components in the device are connected to the I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0059] The processing unit executes the various methods and processes described above, such as methods S1 to S5. For example, in some embodiments, methods S1 to S5 may be implemented as computer software programs tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of methods S1 to S5 described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S5 by any other suitable means (e.g., by means of firmware).

[0060] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload programmable logic devices (CPLDs), and so on.

[0061] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0062] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0063] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A text-guided image editing method based on a diffusion model, characterized in that, include: S1. Obtain the original image, original prompt, and edit prompt; S2. Encode the original image, the original prompts, and the edit prompts to obtain the latent variables of the first image, the text embeddings of the prompts, and the text embeddings of the edit prompts, respectively. S3, Inversion Stage: Based on the text embedding of prompt words, the first image latent variables are diffused and denoised, and the diffused and denoised image latent variables are optimized by combining re-noiseing and structural loss to obtain the second image latent variables; wherein, the stage is a staged structural loss, which is only introduced in the previous denoising time step; S4. Denoising stage: Based on the text embedding of editing prompt words, the second image latent variables are denoised to obtain the third image latent variables; S5. Decode the latent variables of the third image to obtain the edited output image.

2. The text-guided image editing method based on a diffusion model according to claim 1, characterized in that, The noise reduction is performed using fixed-point iteration, and the calculation expression is as follows: , in, , These represent the times within time step t. Second and third The latent variables of the image output at the next iteration; Let be the latent variables of the image output after the (t-1)th time step iteration; To introduce noise; c is for conditional embedding of cue text; The latent variable signal retention coefficient at time step t is used to control the influence of the latent variable from the previous time step on the current latent variable update result. This represents the noise injection coefficient at time step t, used to control the intensity of random noise during the diffusion process; This represents the noise scaling factor at time step t, used to adjust the amplitude of the random noise term.

3. The text-guided image editing method based on a diffusion model according to claim 1, characterized in that, The structured loss is calculated using the following expression: , In the formula: For structured loss; This represents the reconstructed image obtained after denoising and decoding of the latent variables of the image at the current time step; Original image; It is a feature mapping function; This is a structural similarity function.

4. The text-guided image editing method based on a diffusion model according to claim 1, characterized in that, In step S4, based on the text embedding of editing prompt words, the second image latent variable is denoised to obtain the third image latent variable, which specifically includes: S401. Calculate the cross-attention map based on the pixel query Q corresponding to the latent variables of the second image and the text embedding of the editing prompt word K. ,in, The feature dimension of the second image latent variable; S402. Optimize the cross-attention map A based on the edge-aware attention mechanism to obtain the latent variables of the third image.

5. The text-guided image editing method based on a diffusion model according to claim 4, characterized in that, The optimization of the cross-attention graph A based on the edge-aware attention mechanism in S402 is specifically as follows: Cross-attention map Perform gradient calculation to obtain the horizontal gradient. and vertical gradient ; Based on the horizontal gradient corresponding to each pixel and vertical gradient Calculate the gradient magnitude ; Calculate the local difference information for each pixel. ; Based on gradient magnitude and local difference information Attention map Divided into enhanced areas and inhibition region , is represented as: , ,in, The gradient strength threshold, For pixels The corresponding gradient magnitude, For pixels Corresponding local difference information; In the expansion edge mask Under constraints, for cross-attention graphs Modulation is performed, resulting in a modulated cross-attention map. Represented as: , Among them, the first coefficient Used to enhance the area of ​​heightened attention, second coefficient Used to suppress areas of decreased attention.

6. The text-guided image editing method based on a diffusion model according to claim 5, characterized in that, The calculation of local difference information corresponding to each pixel The calculation expression is: , In the formula: For pixels Corresponding local difference information; For pixels The corresponding value; In pixels The average value of the local neighborhood centered on the center.

7. The text-guided image editing method based on a diffusion model according to claim 5, characterized in that, The dilated edge mask The acquisition process is as follows: From the original image Edge information is extracted, and the horizontal and vertical gradients are calculated using the Sobel operator. ,in and These represent Sobel convolution kernels; Based on the horizontal gradient and vertical gradient Calculate edge amplitude ; After thresholding and dilution of the edge amplitude, the resolution is dynamically adjusted according to the number of denoising steps t to match the attention map size, thus obtaining the dilated edge mask. The dilation scale gradually decreases as the number of denoising steps increases.

8. The text-guided image editing method based on a diffusion model according to claim 1, characterized in that, During the noise reduction stage, empty text embeddings are replaced with text embeddings containing editing prompts.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.