A general style character image customization generation method based on a diffusion model
By employing a dual-path guided generation process that combines referencing a sensory self-attention mechanism and a region grouping hybrid attention mechanism, the problems of cumbersome training and insufficient multi-concept generation in existing technologies are solved, achieving efficient and high-fidelity customized generation of human images.
Patent Information
- Application Number
- CN202411156542.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-08-22
AI Technical Summary
Existing methods for personalized generation of portrait images require tedious training or fine-tuning, are prone to overfitting, and cannot effectively handle customized generation of portrait images with multiple concepts, especially in terms of detail and facial identity fidelity.
A coarse-to-fine forward control generation process is adopted, which combines Reference Aware Self-Attention (RSA) and Region Grouping Hybrid Attention (RBA) mechanisms. Through dual-path guidance of reference path and customized generation path, visual features are gradually extracted to generate high-fidelity human images.
It achieves high-fidelity single-concept or multi-concept character image customization generation without additional training, ensuring that the generated concept is highly similar to the reference image, thus improving generation efficiency and detail fidelity.
Smart Images

Figure CN119295574B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal image customized generation, and particularly relates to a universal style character image customized generation method based on a diffusion model. BACKGROUND
[0002] The goal of personalized generation of character images is to generate images of a specific individual in a specified scene, style, and action. Current methods mainly fall into two categories: tuning-based personalized customization and zero-shot personalized customization.
[0003] Tuning-based personalized customization. Tuning-based methods require additional fine-tuning of the model based on the reference image of the specific subject during testing. The pioneering work Dreambooth [1] fine-tunes a pre-trained diffusion model using a large number of reference images, binding an unique identifier with the identity features of a given subject. Concurrent work Textual Inversion [2] converts subject images into simple learnable text embeddings to encode the identity of the subject. Subsequent work such as NeTI [3] and XTI [4] respectively introduce implicit time-aware representations and layer-wise learnable embeddings to achieve better performance. In addition, Tuning-Encoder [5] generates a set of initial latent codes by using a pre-trained encoder, and optimizes these codes through fewer fine-tuning iterations to better preserve the subject identity. Although these methods are effective, they are inefficient due to the need for a large amount of time and computational resources for fine-tuning during testing.
[0004] Zero-shot personalized customization. Zero-shot methods attempt to use a single image and generate customized images through a one-time forward denoising process, thereby significantly speeding up the personalization process. For example, ELITE [6] and InstantBooth [7] achieve this goal by utilizing a global mapping network to encode the reference image into a word embedding, and by injecting local features of the reference image into a cross-attention layer through a local mapping network. Fastcomposer [8] and PhotoMaker [9] obtain identity-centered embedding features by fine-tuning image encoders and fusing class words and image embeddings. Face-Diffuser
[10] reveals the problems of training imbalance and quality compromise in these methods, and solves these problems by proposing a novel collaborative generation method. InstantID
[11] proposes an IdentityNet that guides generation by combining facial features, pose conditions, and text prompts. FlashFace
[12] encodes the reference image into a series of feature maps to preserve more details. Although these methods have achieved impressive results, they only focus on single-concept portrait personalization and cannot handle complex scenes involving multiple concepts.
[0005] In addition, recent research has focused on multi-concept personalized generation. Custom Diffusion
[13] achieves multi-concept personalized generation through closed-form constraint optimization. Perfusion
[14] proposes a dynamic rank-1 update strategy to ensure visual fidelity of concepts. FreeCustom
[15] proposes a multi-reference self-attention mechanism that allows potential images to interact with input concepts. ClassDiffusion
[16] explicitly adjusts the concept space by using semantic preservation loss to maintain semantic consistency when learning new concepts. Although these methods are effective in handling general objects with coarse-grained textures, they still have deficiencies in handling portrait personalization due to the subtle semantics of portraits and the higher detail and fidelity requirements for facial identity.
[0006] In summary, existing methods for personalized generation of portrait images usually require tedious training: either fine-tuning a pre-trained model with a small number of reference images or retraining a pre-trained model on a large-scale dataset. However, these training-based methods are prone to overfitting, making them unable to generate personalized portraits for different styles. In addition, existing methods for personalized generation of portrait images cannot achieve multi-concept customization of portraits, while existing methods for multi-concept personalized customization are only applicable to general objects with coarse-grained textures and still have deficiencies in portrait customization. SUMMARY
[0007] To solve the above problems of the prior art, the present application proposes a method for personalized generation of images of general style characters, which can generate high-fidelity personalized portraits of any style characters with single or multiple concepts without additional training.
[0008] The present application introduces a coarse-to-fine forward control generation process, including two consecutive stages: semantic layout construction and concept feature injection. During the generation process, the technical solution gradually extracts visual features from the input reference concept image to guide the generation, thereby achieving high-fidelity personalized generation of portrait images. The core of the present application is the reference-aware self-attention (RSA) and region-grouped blend attention (RBA) mechanisms proposed by the present application. The technical solution of the present application is described as follows.
[0009] This invention provides a general-style character image customization generation method based on a diffusion model. This method achieves single-concept or multi-concept customized image generation for characters of any style through dual-path guidance of a reference path and a customized generation path. In the customized generation path, a coarse-to-fine forward control generation process is achieved through a reference-aware self-attention (RSA) and region-grouping hybrid attention (RBA) mechanism. The generation process is divided into the following two stages:
[0010] Phase 1 Semantic Scene Construction
[0011] RSA enables the latent image to query features from reference images of all concepts simultaneously, adaptively integrating coarse-grained semantic information to extract a comprehensive semantic understanding from the reference images in order to establish an initial semantic layout.
[0012] Phase Two Concept Feature Injection
[0013] First, a latent semantic map for each concept is calculated based on the attention map to accurately locate the generation position of all concepts in the latent image. Then, RBA divides the latent image into multiple semantic groups and allows each group to query fine-grained features from its corresponding reference concept image to ensure accurate attribute alignment and feature injection. Finally, the features of each reference concept image are accurately injected into the corresponding position to ensure that the generated concept is highly similar to the reference image.
[0014] In this invention, the reference path uses the pre-trained diffusion model Stable Diffusion as the base model, and utilizes its U-Net ∈ θ As a denoiser, visual features are progressively extracted from the input reference concept image to guide the generation; in the customized generation path, the self-attention layer in the first stage of U-Net is replaced by the reference-aware self-attention RSA module, and in the second stage it is replaced by the region grouping hybrid attention RBA module.
[0015] In this invention, in the reference path, for each reference concept image, a forward diffusion process is first applied to calculate the noise latent variable z of the reference image. T '; In each time step t, the noise latent variable z T The corresponding text prompt is input into the denoiser; in the self-attention layer l and time step t, the key features K of the i-th concept are extracted. i,l,t Sum characteristic V i,l,t This guides the generation of images using customized paths.
[0016] In this invention, in the first stage, in each self-attention layer l and time step t, the latent variable z t The query, key, and value characteristics are Q, respectively. l,t K l,t and Vl,t For simplicity, the l, t are omitted later, the key and value features are concatenated with the N key features and value features obtained from the reference path respectively, and and The segmentation masks of the reference concept are introduced, and they are concatenated with the all-1 matrix to obtain M = [1, M1, M2,..., M N ], so as to correct the attention of the model; finally, the RSA is represented as:
[0017]
[0018] Where, ⊙ represents Hadamard product, Q is the query feature obtained from the latent image feature through different linear mappings, and d represents the dimension of Q.
[0019] In the second stage of the application, the steps of calculating the latent semantic graph of each concept based on the attention graph are as follows: 1) generating an attention graph
[0020] In the attention layer l and the time step t, the self-attention graph and the cross-attention graph are calculated by linearly projecting the spatial feature or the text embedding e of the latent image.
[0021]
[0022] Where, Q * (·) and K * (·) are linear projections with dimension d.
[0023] 2) Cross-attention-based semantic segmentation
[0024] The latent variable z t is segmented into a set of mask regions [m1, m2,..., m K ], where K represents the number of text tokens, and m i ∈{0,1} represents the latent semantic region of the text token P i ; specifically as follows:
[0025] First, all cross-attention graphs are upsampled to the same size, then they are averaged and renormalized in the spatial dimension to obtain the final attention graph C t , which estimates the probability of assigning the image block s to the text token P i , finally, the argmax operation is applied in the text token dimension to determine the activation of each image block:
[0026]
[0027] By setting the elements in the image block set {s:i s = 1 and other elements to 0, m i is calculated
[0028] The number of semantic masks of the reference concept for feature injection is N, so the concept-specific mask is reserved, and the remaining masks are merged into a new mask m0 to represent the background area, and finally the latent semantic mask M = [m0, m1, m2,..., m N ] is obtained
[0029] 3) Self-attention-based segmentation completion
[0030] Refine the cross-attention map by multiplying the self-attention map with the corresponding cross-attention map:
[0031]
[0032] Further average and renormalize all the cross-attention maps obtained in the spatial dimension , and input the result into formula (1) to calculate a more refined latent semantic mask.
[0033] In the second stage of the present application, RBA divides the latent image into multiple concept semantic groups, and makes each group query fine-grained features and the specific method of feature injection from its corresponding reference image as follows:
[0034] At each self-attention layer l and time step t, the query, key and value features from the latent variable z t sampled from the Gaussian distribution are Q, K and V, respectively, and l and t are omitted here for brevity, next, RBA divides Q into multiple groups [q0, q1, q2,..., q N ] according to the latent semantic mask M.
[0035] For each group q i , i>0, K and V are concatenated with the key and value features in the corresponding reference path to obtain and For the feature group q0 corresponding to the background, there are and
[0036] Further, the attention output x i of each region is calculated as follows:
[0037]
[0038] where
[0039] Finally, each output pixel group is placed to its corresponding position based on the latent semantic mask M for blending, resulting in the final output.
[0040] In the first and second stages, the RSA and RBA are used to segment the mask M of each reference concept image i A scaling factor w is introduced i to enhance the model's focus on the features of each reference concept image.
[0041] Compared with the prior art, the present application has the beneficial effects that:
[0042] The present application discloses a first high-fidelity general style character image personalization generation technical solution, which does not require additional training and supports single / multiple concept customization. The two-stage control generation process from coarse to fine, combined with our reference perception self-attention and region grouping hybrid attention mechanism, ensures accurate attribute alignment and feature injection of each generated concept. The reference perception self-attention mechanism is proposed to query coarse-grained features from all concept reference images simultaneously, extract comprehensive semantic understanding, and promote initial semantic layout construction of the generated image. The attention-based semantic segmentation method in the diffusion model is proposed to accurately locate the potential generation position of each generated concept. The region grouping hybrid attention mechanism accurately injects fine-grained features of the reference concept image into the corresponding potential generation area on the generated image, ensuring that the generated concept is highly similar to the reference image. A large number of experiments show that the method of the present application has significant advantages in human-centered image generation and multi-concept portrait customization. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 The model architecture of the present application.
[0044] Figure 2 Segmentation example of the latent semantic mask. DETAILED DESCRIPTION
[0045] The technical solutions of the present application will be described in detail below in conjunction with the drawings and examples.
[0046] The present application aims to propose a general style character image customization generation method, which is as follows.
[0047] I. Method Overview
[0048] The present application aims to realize high-fidelity personalized generation of characters, and realize single-concept or multi-concept customized image generation of multi-style characters without training. To this end, the present application proposes a coarse-to-fine control generation process combined with the reference perception self-attention (RSA) and region grouping hybrid attention (RBA) mechanisms proposed by the present application. Specifically, the sampling process of the present technical solution is divided into two consecutive stages, with a total of T denoising steps.
[0049] In the first stage, the present application uses RSA to gradually extract overall semantic understanding from the reference images of all concepts, promoting the initial semantic scene construction. This process gradually extracts comprehensive semantic understanding from the reference images, and once the overall scene is preliminarily established, it enters the second stage to refine the features of the generated concepts:
[0050] In the second stage, the present application first generates a latent semantic map at each step to accurately locate the generation positions of all concepts. Then, RBA accurately injects the features of each reference concept image into the corresponding positions, ensuring that the generated concepts are highly similar to the reference images. In each step, the present application proposes a training-free semantic segmentation method to identify the potential generation areas of all concepts. Specifically, the present application first normalizes the cross-attention map obtained during the generation process, assigning each image block to its corresponding concept semantics, while using a self-attention map to refine and supplement the semantic regions. Subsequently, according to the obtained latent semantic map, RBA divides the latent image into multiple semantic groups, and makes each group query fine-grained features from its respective reference concept image. This process effectively ensures accurate alignment and detailed feature injection of each concept's attributes. To guide the model to pay more attention to the given concept features and eliminate irrelevant information in the reference images, the present application implements a weighted mask strategy during the generation process.
[0051] 1. Self-attention layer in the original diffusion model
[0052] The present application adopts the pre-trained diffusion model Stable Diffusion (SD) v2.1
[17] as the basis model of the present application, and uses the U-Net ∈ θ in it as the denoiser. The original SD U-Net consists of 16 layers, each including a residual block, a self-attention module, and a cross-attention module. The self-attention (SA) layer in the U-Net is crucial for image layout establishment and feature refinement, and its mathematical expression is:
[0053]
[0054] where Q, K, and V are query, key, and value features obtained from latent image features through different linear mappings, with dimension d.
[0055] 2. Customized generation of dual-path bootloader
[0056] like Figure 1 As shown, the overall generation process includes two paths: the reference path and the custom generation path. These two paths will be explained in detail below.
[0057] Reference path: For each reference concept image in the reference path, the present invention first applies a forward diffusion process
[18] to calculate the noise latent variable z of the reference image. T In each time step t, the present invention will z T 'and the corresponding text prompts are entered into U-Net∈ θ In the self-attention layer l and time step t, this invention extracts the key and value features K of the i-th concept. i,l,t and V i,l,t This guides the generation of images using customized paths.
[0058] Customized generation path: This path first samples the latent variable z from the Gaussian distribution N(0,I). T Next, this invention modifies the original self-attention module of SDU-Net, extending it to the RSA / RBA of this invention. In each step t, this invention... t Input the target text prompt P into the modified U-Net In the middle, K is obtained from the reference path i,l,t and V i,l,t Features are extracted from each reference concept image to ensure that the generated concept is highly similar to its reference image. The final denoising result of this process is the personalized image generated.
[0059] II. Coarse-to-fine control generation process
[0060] In customizing the generation path, this invention proposes a coarse-to-fine controlled generation process to generate images: coarse-grained semantic scene construction is performed in the first αT steps, and fine-grained conceptual feature injection is performed in the subsequent T(1-α) steps. In the latter stage, this invention proposes an attention-based semantic segmentation method, laying the foundation for accurate feature injection. The specific details of each stage will be described in detail below.
[0061] 1. Semantic Scene Construction
[0062] Reference-Aware Self-Attention (RSA): RSA enables a latent image to simultaneously query the reference image features of all concepts to integrate coarse-grained semantic information for the construction of an initial semantic scene. Specifically, in each self-attention layer l and time step t, from z... t The retrieved query, key, and value features are Q. l,t, K l,t and V l,t . The present application will K l,t and V l,t
[0063] and concatenate with N key-value features obtained from the reference paths respectively (N is the number of reference concept images,) to get and For brevity, l, t are omitted here. However, the goal of the present application is to query features from concept regions in reference images only, because irrelevant background information in reference images can distract the attention of the model and reduce its effectiveness. To solve this problem, the present application introduces segmentation masks of reference concepts and concatenates them with all-1 matrices to get M = [1, M1, M2,..., M N ], so as to correct the attention of the model. Finally, RSA can be represented as:
[0064]
[0065] where ⊙ denotes Hadamard product. This process ensures that the generated image can effectively interact with the overall semantic information in the reference image while filtering out irrelevant noise.
[0066] 2. Latent semantic mask generation
[0067] To ensure accurate feature injection for each generated concept at the pixel level, the present application needs to obtain the latent generated region of each concept at each step. The attention layer in U-Net contains rich semantic information and can be used to effectively identify semantic units. Therefore, the present application proposes to calculate the latent semantic map of each concept based on attention map, which includes two consecutive steps: semantic segmentation based on cross-attention and segmentation completion based on self-attention.
[0068] Generating attention map:
[0069] In attention layer l and time step t, the present application calculates the self-attention map and cross-attention map by linearly projecting the spatial features or text embedding e of the latent image.
[0070]
[0071] where Q * (·) and K * (·) are linear projections with dimension d.
[0072] Semantic segmentation based on cross-attention:
[0073] Cross-attention maps contain image patches s and text tags P. i The correlation value between them. Therefore, in In each row, a higher probability Indicates s and P i The relationship between them is closer. Based on this, the present invention will... t Divide into [m1, m2, ..., m K The set of regions in the mask, where K represents the number of text tags, m i ∈{0,1} represents the label P i The potential semantic region.
[0074] Specifically, this invention first combines all cross-attention maps Upsampled to the same size, then averaged and renormalized along the spatial dimension to obtain the final attention map C. t The figure estimates the allocation of image patches s to text tags P. i The probability of activation is determined by applying the argmax operation along the text tag dimension.
[0075]
[0076] Following this line of thought, the present invention uses an image patch set {s:i s The elements in =i} are set to 1, and all other elements are set to 0 to calculate m. i It is important to note that this invention only requires semantic masks of N reference concepts for feature injection. Therefore, this invention retains the masks of concept-specific markers and merges the remaining masks into a new mask m0 to represent the background region, ultimately obtaining M = [m0, m1, m2, ..., m...]. N ].
[0077] Figure 2 The first line shows an example of the semantic segmentation results described above. These semantic maps are effective at identifying the approximate location of generated concepts. However, they often exhibit unclear boundaries and may contain internal gaps, leading to unsatisfactory results. To address this issue, this invention uses self-attention maps to further refine and complete the latent semantic maps, a process that will be described in detail in the following sections.
[0078] Segmentation completion based on self-attention:
[0079] Self-attention map The correlation between image patches is estimated, so the present application refines the incomplete activation regions in cross-attention maps by passing semantic information between image patches. This method is similar to feature passing in spectral domain graph convolution, because the self-attention map can be regarded as a transfer matrix, where each element is non-negative, and the sum of each row is 1. Therefore, we refine the cross-attention map by multiplying the self-attention map with the corresponding cross-attention map: Based on this, the present application further averages and renormalizes all the cross-attention maps obtained in the spatial dimension , and inputs the results into (1) to calculate a more refined latent semantic mask.
[0080] Figure 2 The last row of the figure shows the refinement effect, where the generated segmentation map clearly indicates the boundaries of the objects, and the internal holes are significantly reduced.
[0081] 3. Concept feature injection
[0082] Region grouping attention (RBA): According to the latent semantic mask M, RBA divides the latent image into multiple concept semantic groups, and makes each generated concept query features from its corresponding reference image. Specifically, at each self-attention layer l and time step t, the query, key and value features from z t are Q, K and V, respectively. For brevity, l and t are omitted here. Next, RBA divides Q into multiple groups [q0, q1, q2,..., q N ] according to M.
[0083] For each group q i (i>0), we concatenate K and V with the key and value features in the corresponding reference path to obtain and For the feature group q0 corresponding to the background, we have and
[0084] Based on this, we calculate the attention output x i of each region as follows:
[0085]
[0086] where
[0087] Finally, we place each output pixel group to its corresponding position based on M for mixing, thus obtaining the final output. This process effectively ensures the accurate attribute alignment and feature injection of each generated concept.
[0088] 4. Weighted mask strategy
[0089] Although the semantic segmentation mask of each introduced concept helps to avoid the influence of irrelevant features in the reference images, the existing model is still difficult to accurately capture the unique attributes of the target concept, especially the information in preserving the identity of the person. To solve this problem, the present invention introduces a scaling factor W i A scaling factor W is introduced i , which aims to enhance the model's focus on the required concept features.
[0090] Embodiment 1
[0091] The present invention compares with the existing latest public technical solutions PhotoMaker[9], FlashFace
[12] , Fastcomposer[8] and InstantID
[11] for single-concept person image customization generation, and compares with ClassDiffusion
[13] , FreeCustom
[15] , CustomDiffusion
[16] and Prefusion
[14] for multi-concept person image customization generation.
[0092] To further verify the effectiveness of the proposed method, the present invention first visualizes the correlation map between the latent image and all reference concept images according to the attention map generated by RSA in the first stage of the customization generation path, where the areas with the highest correlation on the latent image and the reference concept image are represented using the same color. The present invention also visualizes the attention map between certain image blocks on the generated image in the second stage of the customization generation path and their corresponding reference concept images in RBA. According to the visualization results, it can be seen that the person image customization generation process proposed by the present invention accurately extracts features from the corresponding reference image in the pixel of the latent generation area of each concept, thereby ensuring that the generated concept is similar to the reference concept.
[0093] References
[0094] [1] Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine-tuning text-to-image diffusion models for subject-driven generation. In CVPR, pages 22500-22510, 2023.
[0095] [2] Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion. In ICLR, 2023.
[0096] [3] Yuval Alaluf, Elad Richardson, Gal Metzer, and Daniel Cohen-Or. A neural space-time representation for text-to-image personalization. ACM Transactions on Graphics, 42(6):1–10, 2023.
[0097] [4] Andrey Voynov, Qinghao Chu, Daniel Cohen-Or, and Kfir Aberman. p+: Extended textual conditioning in text-to-image generation. arXiv preprint arXiv:2303.09522, 2023.
[0098] [5] Moab Arar, Rinon Gal, Yuval Atzmon, Gal Chechik, Daniel Cohen-Or, Ariel Shamir, and Amit H. Bermano. Domain-agnostic tuning-encoder for fast personalization of text-to-image models. In SIGGRAPH Asia, pages 1–10, 2023.
[0099] [6] Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation. arXiv preprint arXiv:2302.13848, 2023.
[0100] [7] Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung. Instantbooth: Personalized text-to-image generation without test-time finetuning. arXiv preprint arXiv:2304.03411, 2023.
[0101] [8] Guangxuan Xiao, Tianwei Yin, William T Freeman, Fr'edo Durand, and Song Han. Fastcomposer: Tuning-free multi-subject image generation with localized attention. arXiv preprint arXiv:2305.10431, 2023.
[0102] [9] Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming-Ming Cheng, and Ying Shan. Photomaker: Customizing realistic human photos via stacked id embedding. In CVPR, pages 8640-8650, 2024.
[0103]
[10] Yibin Wang, Weizhong Zhang, Jianwei Zheng, and Cheng Jin. High-fidelity person-centric subject-to-image synthesis. In CVPR, pages 7675-7684, 2024.
[0104]
[11] Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, and Anthony Chen. Instantid: Zero-shot identity-preserving generation in seconds. arXiv preprint arXiv:2401.07519, 2024.
[0105]
[12] Shilong Zhang, Lianghua Huang, Xi Chen, Yifei Zhang, Zhi-Fan Wu, Yutong Feng, Wei Wang, Yujun Shen, Yu Liu, and Ping Luo. Flashface: Human image personalization with high-fidelity identity preservation. arXiv preprint arXiv:2403.17008, 2024.
[0106]
[13] Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In CVPR, pages 1931-1941, 2023.
[0107]
[14] Yoad Tewel, Rinon Gal, Gal Chechik, and Yuval Atzmon. Key-locked rank one editing for text-to-image personalization. In ACM SIGGRAPH, pages 1-11, 2023.
[0108]
[15] Ganggui Ding, Canyu Zhao, Wen Wang, Zhen Yang, Zide Liu, Hao Chen, and Chunhua Shen. Freecustom: Tuning-free customized image generation for multi-concept composition. In CVPR, pages 9089-9098, 2024.
[0109]
[16] Jiannan Huang, Jun Hao Liew, Hanshu Yan, Yuyang Yin, Yao Zhao, and Yunchao Wei. Classdiffusion: More aligned personalization tuning with explicit class guidance. arXiv preprint arXiv:2405.17532, 2024.
[0110]
[17] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10684-10695, 2022.
[0111]
[18] Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. NeurIPS 33, 6840-6851, 2020.
Claims
1. A universal style character image customization generation method based on a diffusion model, characterized in that, The method realizes single-concept or multi-concept customized image generation of any style character based on dual-path guidance of reference path and customized generation path; in the customized generation path, forward control generation from coarse to fine is realized through reference perception self-attention (RSA) and region grouping hybrid attention (RBA) mechanisms, and the generation process is divided into the following two stages: First stage: semantic scene construction RSA enables the latent image to query features from all reference images of concepts at the same time, and adaptively integrates coarse-grained semantic information, so as to extract comprehensive semantic understanding from the reference images to establish an initial semantic layout; Second stage: concept feature injection Firstly, the latent semantic map of each concept is calculated based on the attention map to accurately locate the generation position of all concepts in the latent image; then, RBA divides the latent image into multiple concept semantic groups, and makes each group query fine-grained features from the corresponding reference concept image to ensure accurate attribute alignment and feature injection; finally, the features of each reference concept image are accurately injected into the corresponding position to ensure that the generated concept is highly similar to the reference image.
2. The diffusion model based general style character image customization generation method according to claim 1, characterized in that, In the reference path, the pre-trained diffusion model Stable Diffusion is used as the base model, and the U-Net in it is used as the reference perception module θ As a denoiser, visual features are gradually extracted from the input reference concept image to guide generation; in the customized generation path, the self-attention layer in the first stage U-Net is replaced with the reference perception self-attention RSA module, and in the second stage it is replaced with the region grouping hybrid attention RBA module.
3. The diffusion model based general style character image customization generation method according to claim 1, characterized in that, In the reference path, for each reference concept image, a forward diffusion process is first applied to compute the noise latent variable z T At each time step t, the noise latent variable z T and the corresponding text prompt are input into the denoiser; at self-attention layer l and time step t, the key features K i,l,t and the value features V i,l,t are extracted for the i-th concept to guide the image generation of the custom generation path.
4. The diffusion model based general style character image customization generation method according to claim 1, characterized in that, In the first stage, in each self-attention layer l and time step t, the latent variable z t The query, key and value features, respectively, Q l,t , K l,t and V l,t , are omitted for brevity, l, t, and are concatenated with the N key and value features obtained from the reference path, respectively, to obtain and Split masks of the reference concept are introduced and concatenated with an all-1 matrix to obtain M = [1, M1, M2,..., M N ], thereby correcting the attention of the model; finally, RSA is represented as: Wherein, ⊙ represents Hadamard product, and d represents the dimension of Q.
5. The diffusion model based general style character image customization generation method according to claim 1, characterized in that, In the second stage, the steps of calculating the latent semantic map of each concept based on the attention map are as follows: 1) generating an attention map In attention layer l and time step t, the self-attention map and the cross-attention map are computed by linearly projecting the spatial features where Q * (·) and K * (·) are linear projections of dimension d; 2) semantic segmentation based on cross-attention The latent variable z t is partitioned into a set of regions [m1, m2,..., m K K-1] masks, where K represents the number of text tokens, m i ∈ {0, 1} represents the latent semantic region of a text token P i ; in particular as follows: First, all cross-attention maps are up-sampled to the same size and then averaged and re-normalized over the spatial dimensions to obtain the final attention map C t which estimates the likelihood of assigning an image patch s to a text token P i Finally, an argmax operation is applied over the text token dimension to determine the activation of each image patch: By setting the elements in the set of image blocks {s:i s = i} to 1 and the other elements to 0, m i is calculated. The number of semantic masks of reference concepts for feature injection is N, so the mask of concept-specific markers is reserved and the remaining masks are merged into one new mask m0 to represent the background region, resulting in the final latent semantic mask M = [m0, m1, m2,..., m N ]. 3) segmentation completion based on self-attention The cross-attention map is refined by multiplying the self-attention map and the corresponding cross-attention map: Further, all cross-attention maps obtained are averaged and renormalized in the spatial dimension and the result is input into equation (1) to compute a more refined latent semantic mask M. Further, all cross-attention maps obtained are averaged and renormalized in the spatial dimension and the result is input into equation (1) to compute a more refined latent semantic mask M.
6. The diffusion model based general style character image customization generation method according to claim 5, characterized in that, In the second stage, RBA divides the latent image into multiple concept semantic groups, and makes each group query fine-grained features from the corresponding reference image, and the specific method of feature injection is as follows: At each self-attention layer l and time step t, the latent variable z t The query feature, key feature, value feature are Q l,t , K l,t and V l,t , respectively, l, t are omitted for brevity, RBA divides Q into multiple groups [q0, q1, q2,..., q N ] according to the latent semantic mask M; For each group q i where i > 0, K and V are concatenated with their corresponding key and value features in the reference path, resulting in and For the feature group q0 corresponding to the context, there are and Further, the attention output x of each region is calculated respectively i As follows: wherein Finally, based on the latent semantic mask M, each output pixel group is placed in the corresponding concept generated position for mixing, so as to obtain the final output.
7. The diffusion model based general style character image customization generation method according to claim 1, characterized in that, In the first and second stages, the RSA and RBA are trained to generate segmentation masks M for each reference concept image i A scaling factor w is introduced i to enhance the model's focus on the features of each reference concept image.
Citation Information
Patent Citations
Face image generation method and system and model training method
CN117522697A
Virtual fitting method and system for reversely generating portrait fitting effect according to image
CN117974950A