Generated image background control method capable of embedding diffusion model

By embedding the background control method BgControl into the MS-Diffusion model, combining image and text features, and using cross-attention maps to generate masks, the problem of insufficient background consistency in existing technologies is solved. This achieves consistency between the background and the reference image and the integrity of the target in the generated image, thus improving the adaptability to multiple scenes.

CN121095375APending Publication Date: 2025-12-09INST OF COMPUTING TECH CHINESE ACAD OF SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511197128.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing customized image generation methods are insufficient in terms of background consistency, making it difficult to accurately control the background while maintaining target consistency, especially lacking the ability to flexibly respond to changes in multiple scenes.

Method used

The background control method BgControl is embedded in the MS-Diffusion model. It combines image and text features through a noise fusion module, generates a mask using cross-attention maps to ensure consistency between the background and the reference image, and achieves the integrity and consistency of the target through denoising processing of the Unet network.

Benefits of technology

It effectively maintains the consistency between the background in the generated image and the input background reference image, improving background consistency and target integrity, and enhancing the multi-scene adaptability of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095375A_ABST
    Figure CN121095375A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, and discloses a generated image background control method capable of embedding a diffusion model, which comprises the following steps of: performing noise addition processing of a total reasoning step number T on a background reference image to obtain total Gaussian noise; and inputting and introducing the target reference image and the total Gaussian noise into an MS-Diffusion model of a background control method to obtain a generated image. According to the method, the MS-Diffusion model is improved, so that the background in the generated image can be kept consistent with the input background reference image while the customization capability of the original model is kept.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of customized image generation, and in particular to a method for controlling the background of a generated image that can embed a diffusion model. BACKGROUND

[0002] Customized image generation refers to generating an image containing a specified concept and following a provided text description based on a given set of images with user-specific concepts. Due to its wide application in multiple scenarios, the research on customized image generation methods has always been an important direction in the field of image generation. In customized image generation, the Textual Inversion model learns a new word embedding to represent the given target subject. Through an optimization process, this method can be used for customized subject generation in a plug-and-play manner without affecting the prior knowledge of the T2I model. Similar to the Textual Inversion model, DreamBooth introduces a rare word as a unique identifier to represent the target subject and aligns it with the given subject by fine-tuning the prior parameters. These two methods provide an important theoretical basis for subsequent parameter-based customized generation methods.

[0003] However, although the parameter tuning method is effective, it needs to be trained for each object, and the obtained model parameters can only be used for single-target customized generation, which not only increases the time and storage resources required for image generation, but also reduces the generality of the method. Therefore, researchers have gradually shifted their focus to parameter-free customization methods. These methods train an additional model on open-world images to fuse the features of reference images. Among them, a typical method is IP-Adapter, which uses a pre-trained CLIP image encoder to extract image features and projects these features into a feature space as image embeddings. In order to combine image embeddings with text information to achieve customized generation, IP-Adapter introduces an additional image cross-attention layer in the text cross-attention layer, which is called double cross-attention. By combining the attention mechanisms of text and image, IP-Adapter can effectively integrate reference image information into the generation process.

[0004] Although existing customization methods have proposed many excellent solutions to ensure target consistency, the background consistency problem has not been given enough attention. Among other methods that can control the background, image editing methods modify elements in the background through text prompts, but these methods lack customization capabilities for the target. Image redrawing methods can add targets to specified areas of the background, but the text alignment capabilities of such methods are generally weak, and the generated target action poses are mostly consistent with the target reference image, making it difficult to flexibly respond to changes in different scenarios. Therefore, existing image generation methods cannot well meet the content consistency requirements in story creation. SUMMARY

[0005] The purpose of the present application is to enable the background of the generated image to be consistent with the input background reference image while maintaining the original model customization capability, and to provide a diffusion model-embeddable generated image background control method.

[0006] To achieve the above-mentioned purpose of the application, the embodiments of the present application provide the following technical solutions:

[0007] A diffusion model-embeddable generated image background control method, comprising the following steps:

[0008] Step 1: performing total noise addition processing on the background reference image for a total number of reasoning steps T to obtain total Gaussian noise Z T , t = 1, 2,..., T;

[0009] Step 2: inputting the target reference image and the total Gaussian noise into the MS-Diffusion model to obtain a generated image;

[0010] The step 2 specifically comprises the following steps:

[0011] Step 2-1: setting a start time step t1 and an end time step t2, and 0 < t1 < t2 < T;

[0012] Step 2-2: when t = 1, performing denoising processing on the total Gaussian noise Z T as the input of the Unet network to obtain Gaussian noise z T-1 , and inputting the Gaussian noise z T-1 as the input of the Unet network at the next moment; repeating step 2-2 until t = t1, and then entering step 2-3;

[0013] Step 2-3: when t = t1, performing t times of noise addition processing on the background reference image to obtain Gaussian noise c t , inputting the Gaussian noise z T-t and the Gaussian noise c t into a noise fusion module to obtain fusion noise X t; the fusion noise X t ; the Gaussian noise z T-(t+1) ; repeat step 2-3 until t > t2, and then go to step 2-4;

[0014] Step 2-4: when t > t2, the Gaussian noise z T-t ; the Gaussian noise z T-(t+1) ; repeat step 2-4 until t = T, and then obtain the generated image z0.

[0015] Further, when 1 ≤ t ≤ T, at each time step t, the input of the Unet network includes image features and text features in addition to the Gaussian noise or the fusion noise of the previous time step. The specific steps are as follows:

[0016] Step p1: input the target reference image into the image encoder to extract target features;

[0017] Step p2: set the initial position box of the target in the background reference image, and record the coordinate information of the initial position box, which includes the top-left corner coordinates (x0, y0), the bottom-right corner coordinates (x1, y1), and the center coordinates (x c , y c );

[0018] Step p3: input the embedded text description about the target in the target reference image into the text encoder to obtain text features C;

[0019] Step p4: input the extracted target features, the set position box, and the embedded text description into the resampler for interactive processing to generate refined image features F;

[0020] Step p5: input the text features C and the refined image features F into the Unet network.

[0021] Further, after step p5, it further includes:

[0022] Step p6: the Unet network outputs a cross-attention map M attn , which is transmitted to the noise fusion module together with the Gaussian noise when t1 ≤ t ≤ t2, or returned to the Unet network together with the Gaussian noise when 1 ≤ t < t1 and t2 < t ≤ T.

[0023] Further, when t1 ≤ t ≤ t2, the noise fusion module generates a mask for the target at each time step t, and according to the mask, the Gaussian noise z T-t with target information and the Gaussian noise c tThe fusion is performed:

[0024]

[0025] wherein M t represents the mask.

[0026] Further, when t1≤t≤t2, the noise fusion module uses the initial position box M box and the cross-attention map M attn to generate the mask simultaneously, and the calculation formula is represented as:

[0027] wherein (x, y) represents any position in the background reference map; P(x, y, t) represents the probability that the noise vector at (x, y) in the background reference map belongs to the target at the time step t; M attn (x, y) represents the cross-attention map information at (x, y) in the background reference map; G(M attn (x, y)) represents a filter for smoothing M attn (x, y); represents the attention map weight varying with the time step t; the upper left corner coordinates of the initial position box are (x0, y0), the lower right corner coordinates are (x1, y1), and the center coordinates are (x c , y c ).

[0028] Further, the calculation formula of the attention map weight is:

[0029] wherein t is the current time step; S start is the starting time step; S end is the ending time step; C1 and C2 are hyperparameters.

[0030] Further, the calculation formula of the mask M t is:

[0031] wherein Y t is a threshold value; C3 is a hyperparameter for controlling the threshold value.

[0032] Compared with the prior art, the present application has the beneficial effects:

[0033] The application embeds a background control method BgControl (noise fusion module) of customizable image generation on the basis of the MS-Diffusion model, which can be embedded into an existing customization method, while keeping the original customization ability, so that the background in the generated image can be consistent with the input background reference image. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows, and it should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0035] Figure 1 It is a schematic diagram of the prior art MS-Diffusion model;

[0036] Figure 2 It is a schematic diagram of the improved MS-Diffusion model of the present application;

[0037] Figure 3 It is an example of experimental results of the model in embodiment 2;

[0038] Figure 4 It is an example of experimental results of the comparison of the capabilities of the MS-Diffusion model before and after adding the BgControl method in embodiment 2;

[0039] Figure 5 It is an example of experimental results of the comparison of the present method with the traditional Omnigen model and DreamEngine model in embodiment 2. DETAILED DESCRIPTION

[0040] The technical solutions of the embodiments of the present application will be described clearly and completely in the embodiments of the present application combined with the drawings, and obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present application.

[0041] It should be noted that similar reference numerals and letters refer to like items in the accompanying drawings, and once an item is defined in one drawing, it should not require further definition and explanation in the subsequent drawings. Also, in the description of the present application, the terms "first", "second", and the like are used only to distinguish different descriptions, and cannot be understood as indicating or implying relative importance, or implying any such actual relationship or order between these entities or operations. In addition, the terms "connected", "connected" and the like can be direct connection between elements, or indirect connection via other elements.

[0042] Embodiment 1:

[0043] The method is improved based on the customized image method MS-Diffusion model. The method framework of MS-Diffusion is shown in Figure 1 Two key modules, Grounding Resampler and Multi-subject Cross-attention, are added to the LDM model. Grounding Resampler queries and extracts relevant information from image features, and then Multi-subject Cross-attention uses attention masks to enable the model to introduce the image features of the corresponding target as a reference when generating each position box region, thereby generating an image consistent with the given reference image.

[0044] The present application introduces a background control method BgControl (noise fusion module) based on MS-Diffusion, which is implemented by the following technical scheme, as shown in Figure 2 A generated image background control method that can be embedded in a diffusion model includes the following steps:

[0045] Step 1, the background reference image is subjected to total reasoning step noise processing to obtain total Gaussian noise.

[0046] The background reference image is subjected to T-step (T is the total reasoning step, t=1,2,...,T) noise processing using a variational encoder to obtain total Gaussian noise Z T .

[0047] Step 2, input the target reference image and the total Gaussian noise into the MS-Diffusion model to obtain the generated image.

[0048] The step 2 specifically includes the following steps:

[0049] Step 2-1, set the start time step t1, the end time step t2, and 0<t1<t2<T.

[0050] Step 2-2, at the initial time (t=1), total Gaussian noise Z T is input into the Unet network for denoising to obtain Gaussian noise z T-1 , and Gaussian noise z T-1 is input into the Unet network as the input of the next time; Step 2-2 is repeated until t=t1, and Step 2-3 is entered.

[0051] Step 2-3: When t=t1, t-time noise is added to the background reference image to obtain Gaussian noise c t , Gaussian noise z T-t and Gaussian noise c t are input into the noise fusion module for noise fusion to obtain the fusion noise X t of the current time; Gaussian noise z t of the current time is input into the Unet network as the input of the next time for denoising to obtain Gaussian noise z T-(t+1) ; Step 2-3 is repeated until t>t2, and Step 2-4 is entered.

[0052] Step 2-4: When t>t2, Gaussian noise z T-t output by the Unet network at time t is input into the Unet network as the input of the next time for denoising to obtain Gaussian noise z T-(t+1) ; Step 2-4 is repeated until t=T to obtain the generated image z0.

[0053] As an example, assuming that the total reasoning step number T=50 (t=1, 2,..., T), setting the start time step t1=5 and the end time step t2=35 (satisfying the condition 0<t1<t2<T). The background reference image is added with noise for 50 steps using the variational encoder to obtain total Gaussian noise Z 50 , at the initial time t=1, total Gaussian noise Z 50 is input into the Unet network for denoising to obtain Gaussian noise z 49 ; Gaussian noise z 49 is input into the Unet network as the input of the next time to obtain Gaussian noise z 48 ; and so on, until t=5, Gaussian noise z 45 output by the Unet network is obtained. When t=5, Gaussian noise c5 is obtained by adding noise to the background reference image for 5 steps using the variational encoder; Gaussian noise c5 and Gaussian noise z 45 are input into the noise fusion module for noise fusion to obtain fusion noise X5; fusion noise X5 is input into the Unet network as the input of the next time for denoising to obtain Gaussian noise z 44At the same time, the background reference image is subjected to 6-step noise addition using a variational encoder to obtain Gaussian noise c6, and then the Gaussian noise c6 and the Gaussian noise z 44 are input into a noise fusion module for noise fusion to obtain fusion noise X6; this cycle is repeated until t=35 to obtain fusion noise X 35 . The fusion noise X 35 is input into the Unet network as input for denoising at the next moment to obtain Gaussian noise z 14 ; the Gaussian noise z 14 is input into the Unet network as input for denoising at the next moment to obtain Gaussian noise z 13 ; this cycle is repeated until t=50 to obtain the generated image z0.

[0054] Further, when 1≤t≤T, at each time step t, the input of the Unet network includes the image feature and the text feature in addition to the Gaussian noise (or the fusion noise) at the previous moment, please refer to Figure 2 , at each time step t, the following steps are further included:

[0055] Step p1, inputting the target reference image into an image encoder to extract the target feature;

[0056] Step p2, setting an initial position box of the target in the background reference image, recording the coordinate information of the initial position box, the coordinate information of the initial position box including the upper left corner coordinate (x0, y0), the lower right corner coordinate (x1, y1) and the center coordinate (x c ,y c );

[0057] Step p3, inputting the embedded text description about the target in the target reference image into a text encoder to obtain the text feature C;

[0058] Step p4, inputting the extracted target feature, the set position box and the embedded text description into a grounding resampler for interactive processing to generate a refined image feature F;

[0059] Step p5, inputting the text feature C and the refined image feature F into the Unet network;

[0060] In this scheme, the Unet network in the MS-Diffusion injects a Multi-subject Cross-Attention mechanism compared with the traditional Unet, which is the core of text-controlled image generation, and is not simply connected or added, but dynamically and adaptively fused through the attention mechanism to obtain the cross-attention map M attnThrough the cross-attention mechanism, the image features can query the part of the text description that is most relevant to its visual content, thereby accurately injecting text information into the visual features.

[0061] Therefore, after the step p5, a step p6 of outputting a cross-attention map M by the Unet network can be further included. attn , is transmitted to the noise fusion module together with the Gaussian noise (when t1≤t≤t2) or is returned to the Unet network together with the Gaussian noise (when 1≤t

[0062] Further, when t1≤t≤t2, the noise fusion module generates a mask M for the target at each time step t t , according to the mask M t , the Gaussian noise z T-t with target information, and the Gaussian noise c t with background information are fused. The fusion process is represented by the following formula:

[0063] Next, the algorithm of the mask M t is introduced. The initial position box set in the background reference image in step 1 is used to locate the target, but the initial position box only plays a role in preliminary positioning. In the actual generated image, the target may have part of the range exceeding the initial position box. Even if the target is completely located within the initial position box, the initial position box often contains a large amount of background area. Therefore, if the initial position box is directly used as the mask, two problems may be caused: first, part of the noise that should belong to the target may be replaced by the background noise, resulting in incomplete target in the generated image; second, part of the noise that should belong to the background may be replaced by the target noise, resulting in serious segmentation between the background and the target in the generated image.

[0064] To solve the above problems, the noise fusion module of the present scheme introduces the cross-attention map M attn in the Unet network to generate the mask. The cross-attention map M attn can reflect the approximate area corresponding to each word in the noise space. Based on this, the noise fusion module determines the area corresponding to the target in the noise space by calculating the cross-attention between the Gaussian noise z T-t with target information and the Gaussian noise c t with background information. This method can generate a more accurate mask, effectively avoiding the segmentation problem between the target and the background, while ensuring the integrity and consistency of the target in the generated image.

[0065] It is worth noting that the cross-attention map M attnThe correspondence between the target and the background is not always completely accurate, especially in the early stage of the generation process. In the later stage of the generation process, the Gaussian noise vector distribution after the Unet network denoising processing is close to the image domain, and at this time, noise fusion may still cause the problem of segmentation between the target and the background. Therefore, the present scheme introduces two parameters of the starting time step t1 and the ending time step t2. Before the t1 moment, the noise fusion module uses the position box as a mask, which helps to control the target as much as possible within the position box. After the t1 moment, the fusion of noise is stopped to avoid the segmentation problem between the target and the background, so as to realize a more natural fusion. The fusion Gaussian noise at the ending time step t2 will undoubtedly affect the final generated image. Experiments show that when the ending time step t2 is closer to T, the background in the generated image will be more restored to the background reference image, but at the same time, it may also lead to a more serious phenomenon of target edge segmentation; on the contrary, when the ending time step t2 is closer to the starting time step t1, the fusion between the target and the background will be more natural, but the restoration degree of the background will be reduced. Therefore, through experimental testing, when the ending time step t2 is set between 40% and 70% of the total reasoning step T, the quality of the generated image can be guaranteed while maintaining good background consistency. Within this interval, users can adjust the ending time step to control the similarity between the background in the generated image and the background reference image.

[0066] Between the starting time step t1 and the ending time step t2, the noise fusion module uses the initial position box M box and the cross-attention map M attn to generate a mask, and the calculation formula can be represented as:

[0067] Where (x, y) represents any position in the background reference image; P(x, y, t) represents the probability that the noise vector at (x, y) in the background reference image belongs to the target at time step t; M attn (x, y) represents the cross-attention map information at (x, y) in the background reference image; G(M attn (x, y)) represents a filter for smoothing M attn (x, y), and in the present embodiment, a Gaussian filter is used for smoothing; represents the attention map weight changing with time step t, and the calculation formula is:

[0068] Where t is the current time step; S start is the starting time step; S endis the end time step; C1 and C2 are hyperparameters, both of which are set to 0.5 by default. Specifically, C2 represents the proportion of the initial time step at which the cross-attention map M attn occupies, and C1 affects the rate at which the proportion of the cross-attention map M attn occupies increases linearly. That is, C2 controls the initial importance of the cross-attention map M attn at the beginning of the generation process, while C1 determines the growth rate of the attention map weight over time steps.

[0069] By using the probability P(x, y, t) obtained above, a threshold Y t that changes over time steps can be set, and a mask M t is calculated according to the threshold Y t . The calculation formula is as follows:

[0070] where C3 is a hyperparameter that controls the threshold, and is set to 0.5 by default. By setting an appropriate threshold Y t , the specific boundary of the target region can be determined, and C3 controls the sensitivity and adjustment range of the threshold, thereby fine-tuning the generation effect of the target region.

[0071] Example 2:

[0072] This example is based on Example 1 and performs experimental verification. The MS-Diffusion model after introducing the BgControl method is subjected to ablation experiments and comparative experiments, and the experimental results are quantitatively and qualitatively analyzed. Figure 3 As shown in FIG. 6, it can be seen that the cross-attention map corresponding to the embedded text information can reflect the approximate region of the target in the image, and according to Figure 3 the corresponding mask of the target in the noise can be approximately calculated, and then fused with the background information.

[0073] Dataset:

[0074] Since there are differences in the effects of customized methods on human and non-human objects, this embodiment constructs different data sets for humans and animals. The animal data set contains 9 animal reference images and 11 background reference images, each background corresponding to an embedded text description and an initial position box. The human data set contains 15 human reference images and 15 background reference images, each background also with an embedded text description and an initial position box. The actions in the data set include running, walking, lying down, standing, etc., to test whether the model can learn the appearance characteristics of the target, rather than just generating the same pose as the target reference image. Through the arrangement and combination of targets and backgrounds, this experiment finally constructs 324 test samples in the format [target reference image, background reference image, embedded text description, initial position box].

[0075] Evaluation criteria:

[0076] This experiment uses the CLIP model to evaluate the model in three aspects: the ability to maintain target consistency, the ability to maintain background consistency, and the ability to align with text information. CLIP (Contrastive Language-Image Pre-training) is a multi-modal pre-training model proposed by OpenAI in 2021, which achieves cross-modal semantic alignment between images and text through contrastive learning. In this experiment, CLIP-I is used to measure the similarity of the generated image to the background reference image and the target reference image, and the target similarity, and CLIP-T is used to measure the similarity between the content of the generated image and the embedded text description. To avoid the influence of background information in the target reference image, when calculating the target similarity, the experiment uses the target reference image without the background.

[0077] Experimental results:

[0078] The experimental results show that the addition of the BgControl method significantly enhances the ability of the MS-Diffusion model to maintain background consistency, and also has a great advantage in controlling background consistency compared to other customized methods. The comparison of the abilities of the MS-Diffusion model before and after the addition of the BgControl method is shown in Table 1, and examples of the comparison of the abilities of the MS-Diffusion model before and after the addition of the BgControl method are shown in Figure 1. Figure 4 The comparison of the abilities of the MS-Diffusion model before and after the addition of the BgControl method is shown in Table 1, and examples of the comparison of the abilities of the MS-Diffusion model before and after the addition of the BgControl method are shown in Figure 1. Figure 5 The comparison of the abilities of the MS-Diffusion model before and after the addition of the BgControl method is shown in Table 1, and examples of the comparison of the abilities of the MS-Diffusion model before and after the addition of the BgControl method are shown in Figure 1.

[0079] Table 1

[0080] Table 2

[0081] The above description is merely that of a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, and all of them should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for controlling a generated image background which can embed a diffusion model, characterized by, The method comprises the following steps: Step 1, total noise processing of background reference map T, get total Gaussian noise Z T , t = 1, 2,..., T; Step 2, inputting the target reference image and the total Gaussian noise into the MS-Diffusion model to obtain a generated image; The step 2 specifically comprises the following steps: Step 2-1, setting a start time step t1, an end time step t2, and 0 < t1 < t2 < T; Step 2-2, when t = 1, total Gaussian noise Z T As the input of the Unet network, the Gaussian noise z T-1 is obtained by denoising processing, and the Gaussian noise z T-1 is taken as the input of the Unet network at the next moment; loop step 2-2 until t = t1, enter step 2-3; Steps 2-3: When t=t1, perform t noise addition processing on the background reference image to obtain Gaussian noise c. t Gaussian noise z T-t and Gaussian noise c t The noise fusion module performs noise fusion to obtain the fused noise X at the current time. t ; The fusion noise X at the current moment t The noise is then processed as the input to the Unet network at the next time step to obtain Gaussian noise z. T-(t+1) Repeat steps 2-3 until t>t2, then proceed to step 2-4. Step 2-4, when t > t2, the Gaussian noise z output by the Unet network at time t T-t As the input of the Unet network at the next time, the Gaussian noise z is obtained by denoising processing T-(t+1) ; Loop step 2-4 until t = T, and generate image z0.

2. The method of claim 1, wherein, When 1 <= t <= T, at each time step t, the input of the Unet network includes image features and text features in addition to the Gaussian noise or the fused noise at the last time, and the specific steps are as follows: Step p1, inputting the target reference image into an image encoder to extract target features; Step p2, set the target in the initial position box of the background reference map, record the coordinate information of the initial position box, the coordinate information of the initial position box includes the left upper corner coordinate (x0, y0), the right lower corner coordinate (x1, y1) and the center coordinate (x c ,y c ) of the initial position box; Step p3, inputting the embedded text description about the target in the target reference image into a text encoder to obtain text features C; Step p4, inputting the extracted target features, the set position frame and the embedded text description into a resampler for interactive processing to generate refined image features F; Step p5, inputting the text features C and the refined image features F into the Unet network.

3. The method of claim 2, wherein, After the step p5, the method further comprises the following steps: Step p6: Unet network outputs cross-attention map M attn When t1≤t≤t2, it is transmitted to the noise fusion module together with the Gaussian noise; or when 1≤t<t1 and t2<t≤T, it is returned to the Unet network together with the Gaussian noise.

4. The method of claim 1, wherein, When t1≤t≤t2, the noise fusion module generates a mask for the target at each time step t, and according to the mask, the Gaussian noise z T-t with target information and the Gaussian noise c t is fused: wherein M t represents a mask.

5. The method of claim 4, wherein, When t1≤t≤t2, the noise fusion module uses the initial position box M box and the cross attention map M attn to generate a mask, and the calculation formula is represented as: wherein (x, y) represents any position in the background reference map; P(x, y, t) represents the probability that the noise vector at position (x, y) in the background reference map belongs to the target at time step t; M attn (x, y) represents the cross-attention map information at position (x, y) in the background reference map; G(M attn (x, y)) represents a filter for smoothing M attn (x, y); represents the attention map weight varying with time step t; the top-left corner coordinate of the initial position box is (x0, y0), the bottom-right corner coordinate is (x1, y1), and the center coordinate is (x c , y c ).

6. The method of claim 5, wherein, The attention map weight The calculation formula is: where t is the current time step; S start is the start time step; S end is the end time step; C1 and C2 are hyperparameters.

7. The method of claim 6, wherein the method is characterized by: Mask M t The calculation formula is: where Y t is a threshold value; C3is a hyperparameter controlling the threshold value.