Flexible and consistent appearance migration method based on training-free diffusion model
By adopting the training-free diffusion model, dual-guided branching and an improved attention mechanism MAA in the appearance migration technology, the problems of insufficient structural and background consistency and lack of flexible control in the existing technology are solved, and an efficient, flexible and consistent appearance migration effect is achieved.
Patent Information
- Application Number
- CN202510161968.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The existing appearance migration technology has shortcomings in ensuring the consistency of source image structure and background, and lacks flexible regional control, which cannot meet the increasingly diverse user needs.
Using a flexible and consistent appearance migration method based on the training-free diffusion model, the structure of the source image and the appearance characteristics of the reference image are extracted through the dual-guided branch and the improved attention mechanism Mask-Appearance-Attention (MAA), and fuse it in the generated branch to limit the calculation area of attention to ensure the consistency of structure and background.
It realizes high-quality appearance migration, significantly improves generation efficiency, ensures the consistency of source image structure and background, and provides flexible control to meet diverse user needs.
Smart Images

Figure CN120107414A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular to image editing and personalized generation technology. Specifically, the present invention relates to a flexible and consistent appearance migration method based on a training-free diffusion model, which aims to achieve high-quality appearance migration while ensuring the consistency of the source image structure and background, and at the same time provide users with more flexible control to achieve migration of specified areas and multiple appearances. Background Art
[0002] In recent years, generative artificial intelligence has made breakthrough progress and has also spawned many more interesting and personalized tasks. Among them, appearance transfer is a common problem, which aims to seamlessly blend the appearance of a reference image into the source image.
[0003] Early methods for achieving appearance transfer mainly used generative adversarial networks and autoencoders. These methods can achieve simple appearance transfer, but require a lot of computing resources for training, and often show inconsistency and instability when dealing with complex image structures and backgrounds.
[0004] In recent years, with the emergence of diffusion models, image editing methods based on diffusion models have shown significant advantages in various image generation and editing tasks. Some studies have also attempted to apply diffusion models to appearance transfer tasks.
[0005] For example, DiffuseIT uses the pre-trained DINO-ViT to extract the structural and appearance features of the image, and then guides the denoising process to achieve appearance transfer. Due to the powerful generation ability of the diffusion model, DiffuseIT has achieved good results, but it still faces the problems of poor structural consistency and inaccurate background transfer. DiffEditor proposes another image editing method based on the diffusion model, which integrates image and text embedding to guide the denoising process, thereby improving the accuracy and flexibility of the transfer. DiffEditor can ensure the consistency of the background to a certain extent, but it still performs poorly in terms of structural consistency, especially when dealing with scenes with complex structures or backgrounds, the effect is relatively limited. In addition, although this method enhances the flexibility of appearance transfer, its performance is still unstable for complex image content, and it is prone to background interference or loss of details. Cross-image aligns the semantics of the reference image and the target image through the self-attention mechanism of the diffusion model, so that the generated image can query the features of the reference image, thereby achieving appearance transfer. Although this method can achieve high-quality appearance transfer, it has obvious shortcomings in ensuring structural consistency. Especially in the processing of complex backgrounds, the background of the reference image may be unnecessarily transferred to the target image, resulting in confusion in the generated image. In addition, most existing methods lack flexible region control and cannot perform appearance transfer for specific regions.
[0006] In summary, although these methods have made some progress in the appearance transfer task, they still face many challenges. In particular, existing methods generally have limitations in how to ensure the consistency of the source image structure and background and achieve flexible and diverse appearance transfer. Most methods fail to distinguish the boundaries of the appearance transfer area, resulting in the leakage of background information and the loss of structural features. At the same time, existing methods are difficult to provide flexible transfer control and cannot meet the increasingly diverse user needs. Therefore, how to ensure the consistency of the source image structure and background during the transfer process and provide users with flexible control is still a key problem. Summary of the invention
[0007] 1. In view of the above problems, the present invention proposes a flexible and consistent appearance transfer method based on a training-free diffusion model, and includes the following steps:
[0008] S1, based on the source image and the reference image, uses a pre-trained diffusion model to transfer the appearance of the reference image to the source image, while keeping the structure and background of the source image unchanged.
[0009] S2, segment the source image and the reference image, and add noise to the source image and the reference image through DDIM inversion to prepare for subsequent input.
[0010] S3, builds dual-guide branches, each of which is a separate diffusion model. By analyzing the impact of the self-attention mechanism on structure and appearance, the structure of the source image and the appearance features of the reference image are extracted respectively to prepare for subsequent fusion.
[0011] S4, constructs a generation branch, which is used to generate the final image and also provides structural features of the source image.
[0012] S5, designs an improved attention mechanism Mask-Appearance-Attention (MAA) to fuse the structure of the source image and the appearance features of the reference image. And selects the appropriate time step and U-Net layer to replace the self-attention in the generation branch with MAA.
[0013] S6, for flexible appearance migration, users can customize masks to achieve appearance migration of specified areas and migration of multiple appearances.
[0014] 2. The flexible and consistent appearance transfer method based on a training-free diffusion model according to claim 1, wherein step S1 specifically comprises the following steps:
[0015] S11, prepare source image I c and reference image I r,The source image and the reference image have at least one subject ,as the transfer target and the appearance source respectively. The categories can be animals, buildings, cars, food, etc.
[0016] 3. The flexible and consistent appearance transfer method based on a training-free diffusion model according to claim 1, wherein step S2 specifically comprises the following steps:
[0017] S21, use Segment Anything Model (SAM) to segment the source image and the reference image to obtain the segmentation map M of the corresponding subject c ,M r .
[0018] S22, invert the source image and the reference image through DDIM inversion to obtain the corresponding noise map As the input of the subsequent model, this part can be expressed as:
[0019]
[0020] in represents the latent space representation z of the input image predicted at time step t 0 ,∈ θ (z t ,t) is the noise predicted at time step t, α t+1 is a noise scaling factor that is time-step dependent.
[0021] 4. The flexible and consistent appearance transfer method based on a training-free diffusion model according to claim 1 is characterized in that the step S3 constructs a dual-guide branch to extract the structure of the source image and the appearance features of the reference image, specifically comprising the following steps:
[0022] S31, the input of the guiding branch is the noise map of the source image and the reference image It includes two diffusion models that are performed independently at the same time to reconstruct the source image and the reference image respectively. The method of the present invention mainly focuses on the self-attention layer of the diffusion model U-Net. During the reconstruction process, conventional self-attention calculation is performed, which can be expressed as:
[0023] Q c ,K c ,V c =W Q f i c ,W K f i c ,W V f i c (2)
[0024] Q r ,K r ,V r =W Q f i r ,W K f i r ,W V f i r (3)
[0025]
[0026] Where W Q ,W K ,W V Represents different parameter matrices, f i Represents the feature map of the latent space of the input image in different U-Net layers, and d represents the dimension of K, V features.
[0027] S32, extracts the structure of the source image and the appearance features of the reference image from the conventional self-attention layer. This method analyzes the impact of Q, K, and V of the self-attention mechanism on the structure and appearance during the reconstruction process, and believes that Q affects the structure of the image more, while K and V affect the appearance of the image. Therefore, Q of the source image reconstruction branch is retained. c and the K of the reference image reconstruction branch r ,V r Used for subsequent feature fusion.
[0028] 5. The flexible and consistent appearance transfer method based on a training-free diffusion model according to claim 1, wherein the step S4 of constructing a generation branch specifically comprises the following steps:
[0029] S41, generate a noise map of the source image as the input of the branch, denoted as The generation branch is also an independent diffusion model, and its self-attention mechanism calculation can be expressed as:
[0030] Q o ,K o ,V o =W Q f i o ,W K f i o ,W V f i o (5)
[0031] S42, extract features from the attention of the generation branch. The input of the generation branch is the noise map of the source image, retaining Qo To obtain richer structural features.
[0032] 6. A flexible and consistent appearance transfer method based on a training-free diffusion model according to claim 1, characterized in that the step S5 designs an improved attention mechanism Mask-Appearance-Attention (MAA) specifically comprising the following steps:
[0033] S51, the structure and appearance features obtained by the fusion guidance branch. By combining the attention of different branches, the fusion guidance is made so that the Q c and Q o The K of the reference image r ,V r Calculating attention, this part can be expressed as:
[0034] A o =Attn(Q o ,K r ,V r ) (6)
[0035] A c =Attn(Q c ,K r ,V r ) (7)
[0036] Among them A o and A c Represents the calculation results of two attentions.
[0037] S52, limit the area of appearance migration to ensure the consistency of the source image structure and background. The mask is used to limit the operation area of attention, so that the combined attention is only performed in the specified area of appearance migration to prevent background information leakage. This part can be expressed as:
[0038]
[0039] in and Represents the attention calculation result of two combined masks.
[0040] S53, integrate the two attention results. By Q c Calculated, Q c The source image reconstruction process from the guidance branch, Contains more accurate structural information, Q generated by the branch o It contains richer appearance information. Further linear integration of the two can better integrate the structure and appearance features. This part can be expressed as:
[0041]
[0042] Where ω represents the appearance coefficient, which can be used to control the strength of appearance migration.
[0043] S54, calculate the attention of the background area to ensure the consistency of the background. The background area performs conventional self-attention and is restricted by the background mask. The final attention result consists of the migration area and the background area, which can be expressed as:
[0044]
[0045] Among them A o is the final attention calculation result, M -1 Indicates 1-M c . A o The generated branch self-attention result will be replaced for subsequent calculations.
[0046] S55, select the appropriate time step and U-Net layer to replace MAA. Set the time step and layer thresholds S, L. If and only if the time step t is greater than S and the number of layers l is greater than L, replace the self-attention of the generated branch with MAA. This part can be expressed as:
[0047]
[0048] The generation branch generates appearance transfer results that meet the requirements through denoising at multiple time steps.
[0049] 7. The flexible and consistent appearance transfer method based on a training-free diffusion model according to claim 1, wherein the step S6 of implementing the flexible appearance transfer specifically comprises the following steps:
[0050] S61, realize the appearance migration of the specified area. The user can customize the mask to specify the migration area of interest. This mask can be a part of the main body of the source image. This method can well migrate the appearance only to the specified area and ensure that the rest of the area remains unchanged.
[0051] S62, achieving migration of multiple appearances. Given multiple reference images and corresponding migration regions, which may be different parts of the source image subject, the method can gradually migrate the appearance of each image to the corresponding region to generate a more interesting image.
[0052] Beneficial effects: Compared with the prior art, the present invention provides a flexible and consistent appearance transfer method based on a training-free diffusion model, which can achieve more consistent and flexible appearance transfer and produce the following beneficial effects:
[0053] 1. Efficient appearance transfer: Most existing methods require a lot of computational training, while the present invention guides the pre-trained diffusion model by designing attention. It does not require additional training or fine-tuning and only takes 20 seconds to obtain high-quality transfer results, which significantly improves the generation efficiency.
[0054] 2. More consistent appearance transfer: Existing methods mostly operate on the entire image as a unit, and cannot distinguish the boundary between the transfer area and the background, resulting in problems such as loss of source image structure and mis-transfer of reference image background. This paper uses dual-guide branches to extract structural and appearance features, and designs an improved attention mechanism for fusion, operating on the transfer area and background area separately to ensure the consistency of structure and background.
[0055] 3. More flexible control: The present invention provides more flexible control. Users can customize masks to specify areas of interest and only transfer the appearance of the reference image to the corresponding area. At the same time, given multiple reference images and corresponding areas, the present invention can transfer the appearance of each image to the source image to obtain interesting generation results. More flexible control can also meet the increasingly diverse user needs.
[0056] 4. Broad application prospects: The present invention is applicable to various image generation and editing tasks, such as art creation, advertising production, and virtual try-on. It can provide users with a more realistic and convenient experience, reduce costs for creators, and has high commercial value. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A flowchart of a flexible and consistent appearance transfer method based on a training-free diffusion model, describing in detail the various steps from source images and reference images to the final generated image.
[0058] Figure 2 It is a comprehensive display result of the present invention, including global (main) appearance migration, specified area and multi-appearance migration.
[0059] Figure 3 It is the visual effect of the method of the present invention on the appearance migration of the specified area.
[0060] Figure 4 The method of the present invention achieves multiple appearance migration results.
[0061] Figure 5 This is a visual effect comparison diagram of the method of the present invention and other methods.
[0062] Figure 6 , Figure 7 and Figure 8 It is a comparison diagram of the present invention in terms of structural consistency, appearance similarity and background consistency. DETAILED DESCRIPTION
[0063] The flexible and consistent appearance transfer method based on the training-free diffusion model of the present invention achieves high-quality appearance transfer by adopting dual-guide branches and an improved attention mechanism MAA, which can ensure the consistency of the source image structure and background, and provides flexible control, including the appearance transfer of a specified area and the transfer of multiple appearances. The specific implementation process is as follows.
[0064] The invention will be further described below in conjunction with specific embodiments.
[0065] Embodiment 1:
[0066] like Figure 1 As shown in the figure, given a source image and a reference image, the mask image is first obtained through SAM, and the noise map is obtained through DDIMInversion and input into the model. The model includes a guiding branch and a generating branch, both of which are performed simultaneously. The guiding branch reconstructs the source image and the reference image, retains the structure and appearance related parameters, and generates the structure related parameters. The retained parameters are input into the improved attention mechanism MAA for fusion, and the mask is used to limit the calculation area of attention to ensure the consistency of structure and background. The attention calculation results obtained after fusion are returned to the generating branch for subsequent operations. After several time steps, the generating branch obtains the final result.
[0067] Figure 2 It is a comprehensive display result of the present invention, including global (main) appearance migration, specified area and multi-appearance migration.
[0068] Figure 3 It is the visual effect of the method of the present invention on the appearance migration of the specified area.
[0069] Figure 4 The method of the present invention achieves multiple appearance migration results.
[0070] Figure 5 This is a comparison chart of the visual effects of the method of the present invention and other methods. It can be seen that the present method can achieve high-quality appearance migration while ensuring the consistency of structure and background.
[0071] Figure 6 , Figure 7 and Figure 8 The figure is a comparison chart of the present invention in terms of structural consistency, appearance similarity and background consistency. It can be seen that the method in this paper is superior to the existing methods in terms of structural and background consistency, and is also very competitive in terms of appearance similarity.
[0072] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0073] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A flexible and consistent appearance transfer method based on a training-free diffusion model, characterized in that: The following steps are involved: S1, based on the source image and the reference image, uses a pre-trained diffusion model to transfer the appearance of the reference image to the source image, while keeping the structure and background of the source image unchanged. S2, segment the source image and the reference image, and add noise to the source image and the reference image through DDIM inversion to prepare for subsequent input. S3, builds dual-guide branches, each of which is a separate diffusion model. By analyzing the impact of the self-attention mechanism on structure and appearance, the structure of the source image and the appearance features of the reference image are extracted respectively to prepare for subsequent fusion. S4, constructs a generation branch, which is used to generate the final image and also provides structural features of the source image. S5, designs an improved attention mechanism Mask-Appearance-Attention (MAA) to fuse the structure of the source image and the appearance features of the reference image. And selects the appropriate time step and U-Net layer to replace the self-attention in the generation branch with MAA. S6, for flexible appearance migration, users can customize masks to achieve appearance migration of specified areas and migration of multiple appearances.
2. A flexible and consistent appearance transfer method based on a training-free diffusion model as claimed in claim 1, characterized in that: The step S1 specifically includes the following steps: S11, prepare source image I c and reference image I r ,The source image and the reference image have at least one subject ,as the transfer target and the appearance source respectively. The categories can be animals, buildings, cars, food, etc.
3. The flexible and consistent appearance transfer method based on a training-free diffusion model as claimed in claim 1, characterized in that: The step S2 specifically includes the following steps: S21, use Segment Anything Model (SAM) to segment the source image and the reference image to obtain the segmentation map M of the corresponding subject c ,M r . S22, invert the source image and the reference image through DDIM inversion to obtain the corresponding noise map As the input of the subsequent model, this part can be expressed as: in represents the latent space representation z0 of the input image predicted at time step t, ∈ θ (z t ,t) is the noise predicted at time step t, α t+1 is a noise scaling factor that is time-step dependent.
4. The flexible and consistent appearance transfer method based on a training-free diffusion model as claimed in claim 1, characterized in that: The step S3 constructs a dual-guide branch to extract the structure of the source image and the appearance features of the reference image, specifically comprising the following steps: S31, the input of the guiding branch is the noise map of the source image and the reference image It includes two diffusion models that are performed independently at the same time to reconstruct the source image and the reference image respectively. The method of the present invention mainly focuses on the self-attention layer of the diffusion model U-Net. During the reconstruction process, conventional self-attention calculation is performed, which can be expressed as: Q c ,K c ,V c =W Q f i c ,W K f i c ,W V f i c (2) Q r ,K r ,V r =W Q f i r ,W K f i r ,W V f i r (3) Where W Q ,W K ,W V Represents different parameter matrices, f i Represents the feature map of the latent space of the input image in different U-Net layers, and d represents the dimension of K, V features. S32, extracts the structure of the source image and the appearance features of the reference image from the conventional self-attention layer. This method analyzes the impact of Q, K, and V of the self-attention mechanism on the structure and appearance during the reconstruction process, and believes that Q affects the structure of the image more, while K and V affect the appearance of the image. Therefore, Q of the source image reconstruction branch is retained. c and the K of the reference image reconstruction branch r ,V r Used for subsequent feature fusion.
5. The flexible and consistent appearance transfer method based on a training-free diffusion model as claimed in claim 1, characterized in that: The step S4 of constructing the generation branch specifically includes the following steps: S41, generate a noise map of the source image as the input of the branch, denoted as The generation branch is also an independent diffusion model, and its self-attention mechanism calculation can be expressed as: Q o ,K o ,V o =W Q f i o ,W K f i o ,W V f i o (5) S42, extract features from the attention of the generation branch. The input of the generation branch is the noise map of the source image, retaining Q o To obtain richer structural features.
6. A flexible and consistent appearance transfer method based on a training-free diffusion model as claimed in claim 1, characterized in that: The step S5 designs an improved attention mechanism Mask-Appearance-Attention (MAA) which specifically includes the following steps: S51, the structure and appearance features obtained by the fusion guidance branch. By combining the attention of different branches, the fusion guidance is made so that the Q c and Q o The K of the reference image r ,V r Calculating attention, this part can be expressed as: A o =Attn(Q o ,K r ,V r ) (6) A c =Attn(Q c ,K r ,V r ) (7) Among them A o and A c Represents the calculation results of two attentions. S52, limit the area of appearance migration to ensure the consistency of the source image structure and background. The mask is used to limit the operation area of attention, so that the combined attention is only performed in the specified area of appearance migration to prevent background information leakage. This part can be expressed as: in and Represents the attention calculation result of two combined masks. S53, integrate the two attention results. By Q c Calculated, Q c The source image reconstruction process from the guidance branch, Contains more accurate structural information, Q generated by the branch o It contains richer appearance information. Further linear integration of the two can better integrate the structure and appearance features. This part can be expressed as: Where ω represents the appearance coefficient, which can be used to control the strength of appearance migration. S54, calculate the attention of the background area to ensure the consistency of the background. The background area performs conventional self-attention and is restricted by the background mask. The final attention result consists of the migration area and the background area, which can be expressed as: Among them A o is the final attention calculation result, M -1 Indicates 1-M c . A o The generated branch self-attention result will be replaced for subsequent calculations. S55, select the appropriate time step and U-Net layer to replace MAA. Set the time step and layer thresholds S, L. If and only if the time step t is greater than S and the number of layers l is greater than L, replace the self-attention of the generated branch with MAA. This part can be expressed as: The generation branch generates appearance transfer results that meet the requirements through denoising at multiple time steps.
7. The flexible and consistent appearance transfer method based on a training-free diffusion model as claimed in claim 1, characterized in that: The step S6 for implementing flexible appearance migration specifically includes the following steps: S61, realize the appearance migration of the specified area. The user can customize the mask to specify the migration area of interest. This mask can be a part of the main body of the source image. This method can well migrate the appearance only to the specified area and ensure that the rest of the area remains unchanged. S62, achieving migration of multiple appearances. Given multiple reference images and corresponding migration regions, which may be different parts of the source image subject, the method can gradually migrate the appearance of each image to the corresponding region to generate a more interesting image.
Citation Information
Patent Citations
Arbitrary style migration method based on multi-attention network
CN114170066A
Image style migration method based on diffusion model
CN117689532A
Training-free image style migration method and device based on diffusion model and medium
CN119273535A
Utilizing cross-attention guidance to preserve content in diffusion-based image modifications
US20240331236A1
METHOD AND SYSTEM FOR SEMANTIC APPEARANCE TRANSFER USING SPLICING ViT FEATURES
US20240419382A1