Image target removal method and device and electronic equipment

By pre-trained diffusion model combined with self-attention redirection guidance, the problems of inconsistency and instability of image target removal in the prior art are solved, and high-quality and efficient target removal effects are achieved.

CN120298205APending Publication Date: 2025-07-11HANGZHOU HUISHU ZHITONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510355001.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-06
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing image target removal methods do not perform well in maintaining consistency and naturalness of the removal area with the surrounding environment, especially in complex scenarios and large-scale removals, and methods based on generating adversarial networks and diffusion models have problems of slow training speed, instability, or inaccurate results.

Method used

The pre-trained diffusion model is used for encoding and forward diffusion noise addition processing, and the foreground target mask is used for reverse diffusion sampling, and the target removal is achieved through self-attention redirection guidance.

Benefits of technology

It improves the quality and stability of target removal, reduces training overhead, is suitable for various pre-trained diffusion model architectures, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298205A_ABST
    Figure CN120298205A_ABST
Patent Text Reader

Abstract

The invention provides an image target removal method and apparatus, and an electronic device. The method comprises the steps of obtaining a to-be-processed image and a corresponding foreground target mask; inputting the to-be-processed image and the foreground target mask into a pre-training diffusion model; wherein the pre-training diffusion model comprises an auto-encoder module and a noise prediction module; encoding the to-be-processed image through an encoder in the self-encoder module to obtain an initial hidden variable corresponding to the to-be-processed image; performing forward diffusion noise adding processing on the initial hidden variable to obtain a noise-containing hidden variable; based on the foreground target mask, executing a reverse diffusion sampling process on the noise-containing hidden variable through a noise prediction module, and obtaining and correcting a self-attention map of each time step; and guiding a reverse diffusion sampling process based on the corrected self-attention map of each time step and a self-attention redirection guiding mode so as to realize target removal of the to-be-processed image. According to the invention, the target removal quality and stability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to an image target removal method, device, and electronic device. Background Art

[0002] The target removal task, as an important task in the field of computer vision, aims to automatically remove a specified target or object from an image while maintaining the coherence and naturalness of the surrounding environment. This task also has a wide range of practical applications. For example, in image editing software, users can use the target removal function to delete unwanted objects or regions in an image, thereby improving the aesthetics and visual effects of the image, which has important applications in fields such as advertising, graphic design, and photography.

[0003] Traditional target removal methods are block-based or CNN-based methods. Although these target removal methods can achieve certain effects in some cases, they have some limitations, especially in terms of maintaining the consistency and naturalness of the removed area with the surrounding environment. In addition, for complex scenes and large-scale removal cases, the processing capabilities of these methods are limited and often difficult to achieve satisfactory results.

[0004] Recently, with the rapid development of generative models, more and more generative models have been applied to the target removal task. Among them, the more common ones are methods based on generative adversarial networks and methods based on diffusion models. However, these methods also have some deficiencies. The training of generative adversarial networks is slow and unstable, and problems such as mode collapse (the diversity of images generated by the generator decreases) or mode oscillation (the training dynamics between the generator and the discriminator are unstable) are likely to occur. Although the methods based on diffusion models can achieve target removal, due to the randomness of the diffusion model itself and inaccurate guidance operations, these methods may lead to artifacts in the final result or incomplete target removal. Summary of the Invention

[0005] The purpose of the present application is to provide an image target removal method, device, and electronic device. By encoding and forward diffusion denoising the image to be processed through a pre-trained diffusion model, and performing an inverse diffusion sampling process based on a foreground target mask, combined with a self-attention redirection guidance method, an image after target removal is output. The pre-trained diffusion model does not require additional fine-tuning, greatly reducing the training cost, and can improve the quality and stability of target removal.

[0006] In a first aspect, the present application provides an image target removal method, the method comprising: obtaining an image to be processed and a corresponding foreground target mask; inputting the image to be processed and the foreground target mask into a pre-trained diffusion model; wherein the pre-trained diffusion model comprises an autoencoder module and a noise prediction module; encoding the image to be processed through an encoder in the autoencoder module to obtain an initial latent variable corresponding to the image to be processed; performing forward diffusion noise addition processing on the initial latent variable to obtain a noisy latent variable; based on the foreground target mask, performing a reverse diffusion sampling process on the noisy latent variable through the noise prediction module to obtain and correct the self-attention maps at each time step; and guiding the reverse diffusion sampling process based on the corrected self-attention maps at each time step and a self-attention redirection guiding manner to achieve target removal of the image to be processed.

[0007] Further, the step of performing forward diffusion noise addition processing on the initial latent variable to obtain a noisy latent variable comprises: setting a total number of diffusion steps and a variance strategy; calculating a forward diffusion coefficient according to the total number of diffusion steps and the variance strategy; and adding noise to the initial latent variable for the total number of diffusion steps according to the forward diffusion coefficient to obtain a noisy latent variable.

[0008] Further, the noise prediction module is a U-net network architecture; the step of performing a reverse diffusion sampling process on the noisy latent variable through the noise prediction module based on the foreground target mask to obtain and correct the self-attention maps at each time step comprises: for each layer and each time step of the model, performing the following steps: determining a self-attention similarity matrix corresponding to the current layer and the current time step of the model according to the noisy latent variable; correcting the self-attention similarity matrix according to the foreground target mask to obtain a corrected foreground target similarity matrix and a corrected background similarity matrix; determining a corrected foreground attention output result and a corrected background attention output result according to the corrected foreground target similarity matrix, the corrected background similarity matrix, and the corresponding Value matrix; and integrating the corrected foreground attention output result and the corrected background attention output result to obtain a corrected self-attention map corresponding to the current layer and the current time step.

[0009] Further, the step of correcting the self-attention similarity matrix according to the foreground target mask to obtain a corrected foreground target similarity matrix and a corrected background similarity matrix comprises: obtaining the corrected foreground target similarity matrix and the corrected background similarity matrix according to the following first specified formula:

[0010]

[0011] wherein, represents the corrected foreground target similarity matrix; represents the corrected background similarity matrix; inf represents infinity; M l,t represents the flattened foreground object mask; Represents the pixel set of the foreground target area; Represents a set of pixels in the background area; Represents the self-attention similarity matrix corresponding to the t-th time step of the l-th layer of the model;

[0012] The step of determining a modified foreground target attention output result and a modified background attention output result according to the modified foreground target similarity matrix, the modified background similarity matrix, and the corresponding Value matrix comprises: determining the modified foreground target attention output result and the modified background attention output result according to the following second specified formula:

[0013]

[0014] in, Represents the corrected foreground target attention output result; Represents the corrected background attention output result; V l,t Represents the Value matrix corresponding to the tth time step of the lth layer of the model; Represents the corrected foreground target self-attention map; Represents the corrected background self-attention map; softmax() is the normalization function;

[0015] The step of integrating the corrected foreground target attention output result and the corrected background attention output result to obtain a corrected self-attention map corresponding to the current time step of the current layer includes: according to the following third specified formula, the corrected self-attention map corresponding to the current time step of the current layer:

[0016]

[0017] in, Represents the flattened foreground object mask after transposition.

[0018] Further, the step of guiding the reverse diffusion sampling process based on the self-attention maps of each corrected time step and the self-attention redirection guiding method to achieve the target removal of the image to be processed includes: using the first time step as the current time step and the noisy latent variable as the current predicted latent variable, and performing the following prediction correction fusion steps: predicting the current predicted latent variable of the current time step through the noise prediction module to obtain the model predicted noise; correcting the model predicted noise based on the corrected self-attention map of the current time step to obtain the corrected predicted noise; calculating the predicted latent variable of the next time step and the noisy latent variable of the next time step in the reverse diffusion sampling process according to the corrected predicted noise and the forward diffusion coefficient; fusing the predicted latent variable within the foreground target mask and the noisy latent variable outside the mask to retain the background information to obtain the fused latent variable, taking the next time step as the current time step, and re-taking the fused latent variable as the current predicted latent variable, and continuing to perform the prediction correction fusion steps until after all time steps are traversed, taking the fused latent variable of the last time step as the latent variable output result; decoding the latent variable output result through the decoder in the autoencoder module to obtain the output image after target removal.

[0019] Further, the step of correcting the model predicted noise based on the corrected self-attention map of the current time step to obtain the corrected predicted noise includes: correcting the model predicted noise according to the following fourth specified formula to obtain the corrected predicted noise:

[0020]

[0021] where, z t represents the predicted latent variable at the t-th time step in the reverse diffusion sampling process, t ∈ [0, T I ; ∈ θ (z t ) represents the model predicted noise at the t-th time step; s is the target removal guiding coefficient; AAS() represents the entire processing process of performing the reverse diffusion sampling process on the noisy latent variable through the noise prediction module to obtain and correct the self-attention map at the t-th time step; represents the corrected predicted noise at the corrected t-th time step.

[0022] Further, the step of calculating the predicted latent variable of the next time step and the noisy latent variable of the next time step in the reverse diffusion sampling process according to the corrected predicted noise and the forward diffusion coefficient includes: calculating the predicted latent variable of the next time step according to the following fifth specified formula:

[0023]

[0024] where, represents the forward diffusion coefficient at the t-th time step in the reverse diffusion sampling process; βi denotes the variance strategy at the $i$-th moment; denotes the forward diffusion coefficient at the previous time step $t - 1$ in the reverse diffusion sampling process; denotes the corrected predicted noise at time step $t$; $z$ t-1 the predicted latent variable at the previous time step $t - 1$ in the reverse diffusion sampling process;

[0025] Calculate the noisy latent variable at the next time step according to the following sixth specified formula:

[0026]

[0027] where denotes a standard Gaussian distribution with a zero matrix as the mean and an identity matrix $I$ as the covariance, $\in$ is a random Gaussian noise randomly sampled from it; $x_0$ is the latent variable representation of the input image; $x$ t-1 denotes the noisy latent variable at the previous time step $t - 1$.

[0028] In a second aspect, the present application further provides an image target removal device, the device includes: an image mask acquisition module, configured to acquire a to-be-processed image and a corresponding foreground target mask; a model input module, configured to input the to-be-processed image and the foreground target mask into a pre-trained diffusion model; wherein, the pre-trained diffusion model includes an autoencoder module and a noise prediction module; an encoding module, configured to encode the to-be-processed image through an encoder in the autoencoder module to obtain a latent variable corresponding to the to-be-processed image; a noise addition module, configured to perform forward diffusion noise addition processing on the latent variable to obtain a noisy latent variable; a reverse diffusion sampling module, configured to perform a reverse diffusion sampling process on the noisy latent variable through the noise prediction module based on the foreground target mask, and obtain and correct the self-attention maps at each time step; a diffusion guidance module, configured to guide the reverse diffusion sampling process based on the corrected self-attention maps at each time step and the self-attention redirection guidance method to achieve target removal of the to-be-processed image.

[0029] In a third aspect, the present application further provides an electronic device, including a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method described in the first aspect above.

[0030] In a fourth aspect, the present application further provides a computer-readable storage medium, the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method described in the first aspect above.

[0031] In the image target removal method, device, and electronic device provided by this application, first, the image to be processed and the corresponding foreground target mask are obtained; then, the image to be processed and the foreground target mask are input into a pre-trained diffusion model; wherein, the pre-trained diffusion model includes an autoencoder module and a noise prediction module; through the encoder in the autoencoder module, the image to be processed is encoded to obtain the initial latent variable corresponding to the image to be processed; the initial latent variable is subjected to forward diffusion noise addition processing to obtain a noisy latent variable; based on the foreground target mask, the noise prediction module performs a reverse diffusion sampling process on the noisy latent variable to obtain and correct the self-attention maps at each time step; based on the corrected self-attention maps at each time step and the self-attention redirection guidance method, the reverse diffusion sampling process is guided to achieve the removal of the target in the image to be processed. In this method, the image to be processed is encoded and subjected to forward diffusion noise addition processing through the pre-trained diffusion model, and the reverse diffusion sampling process is performed based on the foreground target mask, combined with the self-attention redirection guidance method to output the image after target removal. The pre-trained diffusion model does not require additional fine-tuning, greatly reducing the training cost and improving the quality and stability of target removal. Description of the Drawings

[0032] To more clearly illustrate the specific embodiments of this application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 Flowchart of an image target removal method provided by an embodiment of this application;

[0034] Figure 2 Architecture diagram of a pre-trained model provided by an embodiment of this application;

[0035] Figure 3 Schematic diagram of a forward and reverse diffusion process provided by an embodiment of this application;

[0036] Figure 4 Schematic diagram of a model processing process provided by an embodiment of this application;

[0037] Figure 5 Block diagram of the structure of an image target removal device provided by an embodiment of this application;

[0038] Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of this application. Detailed Description of the Embodiments

[0039] The technical solution of the present application will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0040] In the prior art, although the method based on the diffusion model can achieve target removal, due to the randomness of the diffusion model itself and inaccurate guidance operations, these methods may cause artifacts in the final result or incomplete target removal.

[0041] Based on this, the embodiments of the present application provide an image target removal method, device and electronic device. By using a pre-trained diffusion model to encode and perform forward diffusion noise addition processing on the image to be processed, and performing reverse diffusion sampling based on the foreground target mask, and cooperating with the self-attention redirection guidance method, an image after target removal is output. The pre-trained diffusion model does not require additional fine-tuning, greatly reducing the training cost, and can improve the quality and stability of target removal.

[0042] For the convenience of understanding this embodiment, first, a detailed introduction to an image target removal method disclosed in the embodiments of the present application will be given.

[0043] Figure 1 The following is a flowchart of an image target removal method provided by the embodiments of the present application. The method includes the following steps:

[0044] Step S102, obtain the image to be processed and the corresponding foreground target mask;

[0045] Step S104, input the image to be processed and the foreground target mask into the pre-trained diffusion model; where the pre-trained diffusion model includes an autoencoder module and a noise prediction module;

[0046] Step S106, through the encoder in the autoencoder module, perform encoding processing on the image to be processed to obtain the initial latent variable corresponding to the image to be processed;

[0047] Step S108, perform forward diffusion noise addition processing on the initial latent variable to obtain a noisy latent variable;

[0048] Step S110, based on the foreground target mask, perform reverse diffusion sampling on the noisy latent variable through the noise prediction module, and obtain and correct the self-attention map at each time step;

[0049] Step S112, based on the corrected self-attention maps at each time step and the self-attention redirection guidance method, guide the reverse diffusion sampling process to achieve target removal of the image to be processed.

[0050] In the image target removal method provided by the embodiments of the present application, the input of the pre-trained diffusion model is the image to be processed and the foreground target mask corresponding to the foreground target to be removed, and the output of the model is the image after the specified foreground target is removed. In the above pre-trained diffusion model, the redirected self-attention is used to guide the reverse diffusion sampling process to achieve the guiding method of image target removal, which surpasses the existing state-of-the-art methods in terms of both the quality and stability of target removal. In this embodiment, the model does not require additional fine-tuning, greatly reducing the training cost. Since this embodiment works according to the principle of the diffusion model itself, it can be implemented in various pre-trained diffusion model architectures and checkpoints, with excellent scalability. Different from the training-based methods that can only be applied to specific network architectures, this embodiment can be applied to the newly emerging pre-trained diffusion models generated iteratively, and corresponding updates can be made to further improve the target removal effect by using more powerful models.

[0051] The processing process of the model is elaborated in detail below:

[0052] Step 1, use the encoder in the autoencoder module of the pre-trained diffusion model to encode the image to be processed to obtain the initial latent variable corresponding to the image to be processed;

[0053] In this embodiment, the pre-trained diffusion model adopted is Stable Diffusion (SD). SD is a type of pre-trained latent variable text-to-image diffusion model. As Figure 2 shown, it mainly includes an autoencoder module and a noise prediction network. Given an input image with three RGB channels in the pixel space the encoder ε in the autoencoder module encodes it into a latent variable representation (i.e., the initial latent variable) x0 = ε(X0) in the latent variable space, where

[0054] Step 2, perform forward diffusion and noise addition processing on the initial latent variable to obtain a noisy latent variable;

[0055] As Figure 3 shown, the forward diffusion process aims to add a small amount of Gaussian noise to the input image in T steps, converting the latent variable representation x0 of the input image into random Gaussian noise x T . The forward diffusion process can be defined as follows:

[0056]

[0057] where β t is the variance strategy and I is the identity matrix. It can be deduced that:

[0058]

[0059] where α t= 1 - β t and

[0060] The specific steps of adding noise to the latent variable in the forward diffusion process are as follows:

[0061] Step 2.1, set the total number of diffusion steps T I and the variance strategy

[0062] Step 2.2, according to the total number of diffusion steps T I and the variance strategy β t , calculate the forward diffusion coefficient

[0063] Step 2.3, according to the forward diffusion coefficient, after adding noise to the initial latent variable for the total number of diffusion steps, obtain the noisy latent variable after T I steps where represents the standard Gaussian distribution with a zero matrix as the mean and the identity matrix I as the covariance, and ∈ is the random Gaussian noise randomly sampled from it.

[0064] Step 3, based on the foreground object mask, perform the reverse diffusion sampling process on the noisy latent variable through the noise prediction module to obtain and correct the self-attention maps at each time step;

[0065] In this embodiment, the noise prediction module ∈ of the above pre-trained diffusion model θ mostly uses the U-net network architecture. The U-net network architecture consists of three parts: the downsampling encoder ε U , the bottleneck the upsampling decoder The self-attention module plays an important role in the U-net network. It uses the attention mechanism to aggregate features, thereby controlling the details of image generation more carefully. At the same time, it can also maintain control over the overall structure and content to ensure the consistency and integrity of the generated image. The self-attention map can be obtained from the U-net model where represents the similarity matrix corresponding to the l-th layer and the t-th time step of the model. The specific steps to modify the self-attention maps at each time step are as follows:

[0066] For each layer and each time step of the model, the following steps are performed:

[0067] Step 3.1, according to the noisy latent variable determine the self-attention similarity matrix corresponding to the current layer and the current time step of the model Specifically, given a latent variable where h, w, and c represent the height, width, and channel dimensions respectively, then the corresponding query matrix key matrix and value matrix can be obtained through learnable linear layers l Q , l K and l V respectively, where d is the matrix dimension. The similarity matrix S self can be defined as:

[0068]

[0069] Step 3.2: According to the foreground object mask, correct the self-attention similarity matrix to obtain the corrected foreground object similarity matrix and the corrected background similarity matrix;

[0070] Obtain the corrected foreground object similarity matrix and the corrected background similarity matrix according to the following first specified formula:

[0071]

[0072] where represents the corrected foreground object similarity matrix; represents the corrected background similarity matrix; inf represents infinity; M l,t represents the flattened foreground object mask; set of foreground object region pixels; represents the set of background region pixels; represents the self-attention similarity matrix corresponding to the l-th layer and the t-th time step of the model;

[0073] Step 3.3: According to the corrected foreground object similarity matrix, the corrected background similarity matrix, and the corresponding Value matrix, determine the corrected foreground attention output result and the corrected background attention output result;

[0074] Determine the corrected foreground object attention output result and the corrected background attention output result according to the following second specified formula:

[0075]

[0076] where represents the corrected foreground object attention output result; represents the corrected background attention output result; V l,t represents the Value matrix corresponding to the l-th layer and the t-th time step of the model; represents the corrected foreground object self-attention map; represents the corrected background self-attention map; softmax() is the normalization function;

[0077] Step 3.4, integrate the corrected foreground attention output result and the corrected background attention output result to obtain the corrected self-attention map corresponding to the current layer and the current time step.

[0078] According to the following third specified formula, the corrected self-attention map corresponding to the current layer and the current time step:

[0079]

[0080] Where represents the flattened foreground object mask after transposition. This process can ensure that the above correction operations are performed in the corresponding foreground object region and background region respectively.

[0081] The above method of modifying self-attention is called "Attention Activation and Suppression", abbreviated as AAS in English, and the U-net network processed by the AAS method is denoted as AAS(∈ θ ).

[0082] Step 4, based on the corrected self-attention maps of each time step and the self-attention redirection guidance method, guide the reverse diffusion sampling process to achieve the target removal of the image to be processed;

[0083] Specifically, taking the first time step as the current time step and the noisy latent variable as the current predicted latent variable, perform the following prediction correction fusion steps, as Figure 4 shown:

[0084] Step 4.1, predict the current predicted latent variable of the current time step through the noise prediction module to obtain the model predicted noise;

[0085] Step 4.2, correct the model predicted noise based on the corrected self-attention map of the current time step to obtain the corrected predicted noise;

[0086] According to the following fourth specified formula, correct the model predicted noise to obtain the corrected predicted noise:

[0087]

[0088] Where z t represents the predicted latent variable at the t-th time step in the reverse diffusion sampling process, t ∈ [0, T I ; ∈ θ (z t ) represents the model predicted noise at the t-th time step; s is the target removal guidance coefficient; AAS() represents the entire processing process of performing the reverse diffusion sampling process on the noisy latent variable through the noise prediction module, obtaining and correcting the self-attention map at the t-th time step; Denote the corrected prediction noise at the $t$-th time step after correction.

[0089] Step 4.3, according to the corrected prediction noise and the forward diffusion coefficient, calculate the predicted latent variable at the next time step and the noisy latent variable at the next time step in the reverse diffusion sampling process;

[0090] As Figure 3 shown, starting from $x$ T , the reverse diffusion sampling process aims to obtain real samples by iteratively sampling from $q(x$ t-1 |$x$ t ). However, $q(x$ t-1 |$x$ t ) $\propto q(x$ t-1 ) $q(x$ t |$x$ t-1 ) is difficult to calculate and obtain. Therefore, it is necessary to use the posterior estimation $p$ θ ($x$ t-1 |$x$ t ) of a deep neural network with parameters $\theta$ to fit it. The joint distribution $p$ θ ($x$ 0:T ) is called the reverse diffusion process, which is defined as a Markov chain of learnable Gaussian transition probability distributions starting from :

[0091]

[0092] where the mean $\mu$ θ ($x$ t ) can be obtained by training a deep neural network, and the variance is generally fixed to the same variance $\Sigma$ θ ($x$ t ) = $\beta$ t $I$. To control the randomness of the reverse diffusion sampling process and accelerate sampling, the deterministic DDIM sampling modifies the reverse diffusion process into the following form:

[0093]

[0094] Therefore, calculate the predicted latent variable at the next time step according to the following fifth specified formula:

[0095]

[0096] where, denotes the forward diffusion coefficient at the $t$-th time step in the reverse diffusion sampling process; $\beta$ i denotes the variance strategy at the $i$-th moment; denotes the forward diffusion coefficient at the next time step $t - 1$ in the reverse diffusion sampling process; denotes the corrected prediction noise at the $t$-th time step; $z$ t-1The predicted latent variable at the previous time step t-1 during the reverse diffusion sampling process;

[0097] Calculate the noisy latent variable at the next time step according to the following sixth specified formula:

[0098]

[0099] where represents a standard Gaussian distribution with a mean of the zero matrix and a covariance of the identity matrix I, ∈ is the random Gaussian noise randomly drawn from it; x0 is the latent variable representation of the input image; x t-1 represents the noisy latent variable at the previous time step t-1.

[0100] Step 4.4, fuse the predicted latent variable within the foreground object mask and the noisy latent variable outside the mask to retain background information and obtain the fused latent variable;

[0101] For example, according to the formula z t-1 = z t-1 ⊙M + x t-1 ⊙(1 - M), fuse the predicted latent variable within the foreground object mask and the noisy latent variable outside the mask to obtain the fused latent variable.

[0102] Step 4.5, take the next time step as the current time step, and take the fused latent variable as the current predicted latent variable again, and continue to execute the prediction correction fusion step until after traversing all time steps, take the fused latent variable of the last time step as the latent variable output result;

[0103] When 0 < t < T I Repeat the previous steps 4.1 - 4.4 until the latent variable output result z0 is obtained.

[0104] Step 4.6, perform decoding processing on the latent variable output result through the decoder in the autoencoder module to obtain the output image after target removal

[0105] The embodiment of the present application also provides an image target removal method, which is a method for removing image targets based on a diffusion model without fine-tuning. Generally, it includes the following steps: 1) Use the autoencoder module of the pre-trained diffusion model to encode the input image into a latent variable; 2) Perform forward diffusion noise addition on the latent variable; 3) Input the noise-added latent variable into the denoising module of the pre-trained diffusion model to perform the reverse diffusion sampling process, and obtain and modify the self-attention maps at each time step; 4) Construct a "self-attention redirection guidance" method to guide the reverse diffusion sampling process to achieve target removal. This embodiment can effectively and stably remove the target from a given image using the pre-trained diffusion model without additional fine-tuning, and can be applied to different pre-trained diffusion models.

[0106] Based on the above method embodiment, the embodiment of the present application also provides an image target removal device. Refer to Figure 5 As shown, the device includes: an image mask acquisition module 502, configured to acquire the image to be processed and the corresponding foreground target mask; a model input module 504, configured to input the image to be processed and the foreground target mask into the pre-trained diffusion model; wherein, the pre-trained diffusion model includes an autoencoder module and a noise prediction module; an encoding module 506, configured to encode the image to be processed through the encoder in the autoencoder module to obtain the latent variable corresponding to the image to be processed; a noise addition module 508, configured to perform forward diffusion noise addition processing on the latent variable to obtain a noisy latent variable; a reverse diffusion sampling module 510, configured to perform the reverse diffusion sampling process on the noisy latent variable through the noise prediction module based on the foreground target mask, and obtain and correct the self-attention maps at each time step; a diffusion guidance module 512, configured to guide the reverse diffusion sampling process based on the corrected self-attention maps at each time step and the self-attention redirection guidance method to achieve the target removal of the image to be processed.

[0107] Further, the above noise addition module 508 is configured to set the total number of diffusion steps and the variance strategy; calculate the forward diffusion coefficient according to the total number of diffusion steps and the variance strategy; and after adding noise to the initial latent variable for the total number of diffusion steps according to the forward diffusion coefficient, obtain a noisy latent variable.

[0108] Further, the above noise prediction module is a U-net network architecture; the reverse diffusion sampling module 510 is used to perform the following steps for each layer and each time step of the model: determine the self-attention similarity matrix corresponding to the current layer and the current time step of the model according to the noisy latent variable; correct the self-attention similarity matrix according to the foreground object mask to obtain the corrected foreground object similarity matrix and the corrected background similarity matrix; determine the corrected foreground attention output result and the corrected background attention output result according to the corrected foreground object similarity matrix, the corrected background similarity matrix, and the corresponding Value matrix; integrate the corrected foreground attention output result and the corrected background attention output result to obtain the corrected self-attention map corresponding to the current layer and the current time step.

[0109] Further, the above reverse diffusion sampling module 510 is used to obtain the corrected foreground object similarity matrix and the corrected background similarity matrix according to the following first specified formula:

[0110]

[0111] Where represents the corrected foreground object similarity matrix; represents the corrected background similarity matrix; inf represents infinity; M l,t represents the flattened foreground object mask; Foreground object region pixel set; represents the background region pixel set; represents the self-attention similarity matrix corresponding to the l-th layer and the t-th time step of the model;

[0112] Determine the corrected foreground object attention output result and the corrected background attention output result according to the following second specified formula:

[0113]

[0114] Where represents the corrected foreground object attention output result; represents the corrected background attention output result; V l,t represents the Value matrix corresponding to the l-th layer and the t-th time step of the model; represents the corrected foreground object self-attention map; represents the corrected background self-attention map; softmax() is a normalization function;

[0115] According to the following third specified formula, the corrected self-attention map corresponding to the current layer and the current time step:

[0116]

[0117] Among them, represents the flattened foreground target mask after transposition.

[0118] Furthermore, the above diffusion guidance module 512 is used to take the first time step as the current time step and the noisy latent variable as the current predicted latent variable, and execute the following prediction correction fusion steps: predict the current predicted latent variable at the current time step through the noise prediction module to obtain the model predicted noise; correct the model predicted noise based on the corrected self-attention map at the current time step to obtain the corrected predicted noise; calculate the predicted latent variable at the next time step and the noisy latent variable at the next time step in the reverse diffusion sampling process according to the corrected predicted noise and the forward diffusion coefficient; fuse the predicted latent variable within the foreground target mask and the noisy latent variable outside the mask to retain the background information to obtain the fused latent variable, take the next time step as the current time step, and re-take the fused latent variable as the current predicted latent variable, and continue to execute the prediction correction fusion steps until after traversing all time steps, take the fused latent variable at the last time step as the latent variable output result; decode the latent variable output result through the decoder in the autoencoder module to obtain the output image after target removal.

[0119] Furthermore, the above diffusion guidance module 512 is used to: correct the model predicted noise according to the following fourth specified formula to obtain the corrected predicted noise:

[0120]

[0121] where z t represents the predicted latent variable at the t-th time step in the reverse diffusion sampling process, t ∈ [0, T I ; ∈ θ (z t ) represents the model predicted noise at the t-th time step; s is the target removal guidance coefficient; AAS() represents the entire process of performing the reverse diffusion sampling process on the noisy latent variable through the noise prediction module to obtain and correct the self-attention map at the t-th time step; represents the corrected predicted noise at the corrected t-th time step.

[0122] Furthermore, the above diffusion guidance module 512 is used to calculate the predicted latent variable at the next time step according to the following fifth specified formula:

[0123]

[0124] where represents the forward diffusion coefficient at the t-th time step in the reverse diffusion sampling process; β i represents the variance strategy at the i-th moment; represents the forward diffusion coefficient at the previous time step t - 1 in the reverse diffusion sampling process; represents the corrected predicted noise at time step t; z t-1 the predicted latent variable at the next time step t - 1 in the reverse diffusion sampling process;

[0125] Calculate the noisy latent variable at the next time step according to the following sixth specified formula:

[0126]

[0127] where represents a standard Gaussian distribution with a zero matrix as the mean and the identity matrix I as the covariance, ∈ is a random Gaussian noise randomly sampled from it; x0 is the latent variable representation of the input image; x t-1 represents the noisy latent variable at the next time step t - 1.

[0128] The device provided by the embodiments of the present application has the same implementation principle and the same technical effects as the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the embodiments of the device, reference may be made to the corresponding content in the foregoing method embodiments.

[0129] The embodiments of the present application also provide an electronic device, as Figure 6 shown, which is a schematic structural diagram of the electronic device. Among them, the electronic device includes a processor 61 and a memory 60. The memory 60 stores computer executable instructions that can be executed by the processor 61, and the processor 61 executes the computer executable instructions to implement the above method.

[0130] In Figure 6 the illustrated embodiment, the electronic device further includes a bus 62 and a communication interface 63. Among them, the processor 61, the communication interface 63, and the memory 60 are connected through the bus 62.

[0131] Among them, the memory 60 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 63 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 62 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 62 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a bidirectional arrow is used in Figure 6 , but it does not mean that there is only one bus or one type of bus.

[0132] The processor 61 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 61 or instructions in the form of software. The above-mentioned processor 61 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor 61 reads the information in the memory and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0133] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above method. For specific implementation, reference may be made to the foregoing method embodiments, which will not be elaborated herein.

[0134] The computer program products of the method, device, and electronic device provided by the embodiments of the present application include a computer-readable storage medium storing program codes. The instructions included in the program codes can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference may be made to the method embodiments, which will not be elaborated herein.

[0135] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0136] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0137] In the description of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0138] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: Any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image target removal method, characterized in that, The method includes: Obtaining an image to be processed and a corresponding foreground object mask; Inputting the image to be processed and the foreground object mask into a pre-trained diffusion model; wherein, the pre-trained diffusion model includes an auto-encoder module and a noise prediction module; Encoding the image to be processed through the encoder in the auto-encoder module to obtain an initial latent variable corresponding to the image to be processed; Performing forward diffusion noise addition processing on the initial latent variable to obtain a noisy latent variable; Based on the foreground object mask, performing an inverse diffusion sampling process on the noisy latent variable through the noise prediction module, and obtaining and correcting the self-attention maps at each time step; Guiding the inverse diffusion sampling process based on the corrected self-attention maps at each time step and a self-attention redirection guiding method to achieve the removal of the target in the image to be processed.

2. The method according to claim 1, characterized in that, The step of performing forward diffusion noise addition processing on the initial latent variable to obtain a noisy latent variable includes: Setting a total number of diffusion steps and a variance strategy; Calculating a forward diffusion coefficient according to the total number of diffusion steps and the variance strategy; Adding noise to the initial latent variable for the total number of diffusion steps according to the forward diffusion coefficient to obtain a noisy latent variable.

3. The method according to claim 1, characterized in that, The noise prediction module is a U-net network architecture; the step of performing an inverse diffusion sampling process on the noisy latent variable through the noise prediction module based on the foreground object mask and obtaining and correcting the self-attention maps at each time step includes: For each layer and each time step of the model, the following steps are performed: Determining a self-attention similarity matrix corresponding to the current layer and the current time step of the model according to the noisy latent variable; Correcting the self-attention similarity matrix according to the foreground object mask to obtain a corrected foreground object similarity matrix and a corrected background similarity matrix; Determining a corrected foreground object attention output result and a corrected background attention output result according to the corrected foreground object similarity matrix, the corrected background similarity matrix, and the corresponding Value matrix; Integrating the corrected foreground object attention output result and the corrected background attention output result to obtain a corrected self-attention map corresponding to the current layer and the current time step.

4. The method according to claim 3, characterized in that, The step of correcting the self-attention similarity matrix according to the foreground object mask to obtain a corrected foreground object similarity matrix and a corrected background similarity matrix includes: Obtaining the corrected foreground object similarity matrix and the corrected background similarity matrix according to the following first specified formula: Among them, represents the corrected foreground target similarity matrix; represents the corrected background similarity matrix; inf represents infinity; M l,t represents the flattened foreground target mask; represents the set of foreground target region pixels; represents the set of background region pixels; represents the self-attention similarity matrix corresponding to the l-th layer and the t-th time step of the model; The step of determining a corrected foreground object attention output result and a corrected background attention output result according to the corrected foreground object similarity matrix, the corrected background similarity matrix, and the corresponding Value matrix includes: Determining the corrected foreground object attention output result and the corrected background attention output result according to the following second specified formula: Among them, represents the output result of the corrected foreground target attention; represents the output result of the corrected background attention; V l,t represents the Value matrix corresponding to the l-th layer and the t-th time step of the model; represents the corrected foreground target self-attention map; represents the corrected background self-attention map; softmax() is a normalization function; The step of integrating the corrected foreground target attention output result and the corrected background attention output result to obtain the corrected self-attention map corresponding to the current layer and the current time step includes: According to the following third specified formula, the corrected self-attention map corresponding to the current layer and the current time step: Among them, represents the flattened foreground object mask after transposition.

5. The method according to claim 1, characterized in that The step of guiding the reverse diffusion sampling process based on the corrected self-attention maps of each time step and the self-attention redirection guiding method to achieve the target removal of the image to be processed includes: Taking the first time step as the current time step and the noisy latent variable as the current predicted latent variable, execute the following prediction correction fusion steps: Predict the current predicted latent variable at the current time step through the noise prediction module to obtain the model predicted noise; Based on the corrected self-attention map at the current time step, correct the model predicted noise to obtain the corrected predicted noise; According to the corrected predicted noise and the forward diffusion coefficient, calculate the predicted latent variable and the noisy latent variable at the next time step in the reverse diffusion sampling process; Fuse the predicted latent variable within the foreground target mask and the noisy latent variable outside the mask to retain the background information, obtain the fused latent variable, take the next time step as the current time step, and re-take the fused latent variable as the current predicted latent variable, and continue to execute the prediction correction fusion steps until after traversing all time steps, take the fused latent variable at the last time step as the latent variable output result; Decode the latent variable output result through the decoder in the autoencoder module to obtain the output image after target removal.

6. The method according to claim 5, characterized in that The step of correcting the model predicted noise based on the corrected self-attention map at the current time step to obtain the corrected predicted noise includes: Correct the model predicted noise according to the following fourth specified formula to obtain the corrected predicted noise: where, z t represents the predicted latent variable at the t-th time step in the reverse diffusion sampling process, t ∈ [0, T I ; ∈ θ (z t ) represents the model-predicted noise at the t-th time step; s is the target removal guidance coefficient; AAS() represents the entire process of performing the reverse diffusion sampling process on the noisy latent variable through the noise prediction module to obtain and correct the self-attention map at the t-th time step; represents the corrected predicted noise at the t-th time step after correction.

7. The method according to claim 5, characterized in that, The step of calculating the predicted latent variable and the noisy latent variable at the next time step in the reverse diffusion sampling process according to the corrected predicted noise and the forward diffusion coefficient includes: Calculate the predicted latent variable at the next time step according to the following fifth specified formula: Among them, represents the forward diffusion coefficient at the t-th time step during the reverse diffusion sampling process; β i represents the variance strategy at the i-th moment; represents the forward diffusion coefficient at the next time step t - 1 during the reverse diffusion sampling process; represents the corrected predicted noise at the t-th time step; z t-1 the predicted latent variable at the next time step t - 1 during the reverse diffusion sampling process; Calculate the noisy latent variable at the next time step according to the following sixth specified formula: where represents a standard Gaussian distribution with a mean of the zero matrix and a covariance of the identity matrix I, ∈ is the random Gaussian noise randomly drawn from it; x0 is the latent variable representation of the input image; x t-1 represents the noisy latent variable at the previous time step t - 1.

8. An image target removal device, characterized in that, The device includes: An image mask acquisition module for acquiring the image to be processed and the corresponding foreground target mask; A model input module for inputting the image to be processed and the foreground target mask into a pre-trained diffusion model; wherein, the pre-trained diffusion model includes an autoencoder module and a noise prediction module; An encoding module for encoding the image to be processed through the encoder in the autoencoder module to obtain the latent variable corresponding to the image to be processed; A noise addition module for performing forward diffusion noise addition processing on the latent variable to obtain a noisy latent variable; A reverse diffusion sampling module for performing a reverse diffusion sampling process on the noisy latent variable through the noise prediction module based on the foreground target mask, and obtaining and correcting the self-attention maps of each time step; A diffusion guidance module, configured to guide the reverse diffusion sampling process based on the corrected self-attention maps at each time step and the self-attention redirection guidance method, so as to achieve the target removal of the image to be processed.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the method according to any one of claims 1 to 7.