Identity preserving image generation method and terminal

By layering the extraction of facial and body characteristics and adopting selective injection strategies, the problems of facial and body dissonance and limited text control in the prior art are solved, and efficient generation of identity-keeping images is achieved, reducing the generation time and cost.

CN120260103AActive Publication Date: 2025-07-04HANGZHOU DISHI CHUANGXIANG TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510741913.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing identity-keeping image generation methods require multiple reference characters, long training time, incongruence between the face and the body, and limited text control ability.

Method used

By extracting facial and body features in layers and adopting selective injection strategies, the core identity features are limited to the key layers, and the identity-keeping image is generated by combining the pre-trained diffusion model. Only a single reference image is required, and the traditional fine-tuning mode is abandoned.

Benefits of technology

It realizes synchronous maintenance of facial and body, reduces generation time, overcomes high-cost defects, and achieves an efficient balance between identity fidelity, body coordination and text semantics, improving the identity fidelity and consistency of body structure in generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260103A_ABST
    Figure CN120260103A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating an identity holding image and a terminal. The method comprises the following steps: receiving a target person image and text prompt information transmitted by a user; hierarchical extraction of facial features and body global features is carried out on the target person image, and feature fusion is carried out to generate fusion features; generating text embedding features according to the text prompt information; inputting the text embedding feature, the facial feature and the fusion feature into a pre-trained diffusion model to generate an identity maintaining image; wherein the fusion features are only injected into an identity sensitive layer with the maximum identity retention effect in the diffusion model, and the facial features are only injected into other layers except the identity sensitive layer; according to the method, facial and body features are extracted in a layered mode, a selective injection strategy is adopted, core identity features are limited to a key layer, on the basis that the face and the body are kept synchronously, excessive interference to text-driven detail generation is avoided, and the problem that the text control ability is limited is solved; and efficient balance of identity fidelity, body coordination and text semantics is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model-based image generation, and particularly to a method and a terminal for generating identity-preserving images. Background Art

[0002] The technology of identity-preserving image generation has received extensive attention in the fields of computer vision and graphics in recent years. Among them, the identity-preserving image generation method based on the diffusion model, as a new technical path, is gradually showing its unique advantages and potential. Currently, the identity-preserving generation based on the diffusion model is mainly achieved through two types of methods: the fine-tuning-based method and the non-fine-tuning method.

[0003] (1) Fine-tuning-based method The fine-tuning-based text-to-image identity customization aims to enable the pre-trained model to generate images with specific identities according to text descriptions. Textual Inversion optimizes a new word embedding for the provided identity by the user, while Dreambooth fine-tunes the entire generator to further improve the fidelity. Subsequently, various methods have explored different fine-tuning paradigms in the generator and the embedding space to achieve higher identity fidelity and text alignment. Despite these advances, the time-consuming optimization process for each identity (at least several minutes) still limits its wider application.

[0004] (2) Non-fine-tuning method To reduce the resource requirements of the fine-tuning-based method, a series of non-fine-tuning methods have emerged. Such methods use the face image of the reference person as an additional input and can output pictures consistent with the ID characteristics of the input image. This mode effectively separates offline training from online inference, without the need for a large amount of data and training waiting time. Users only need to provide a photo and can directly use the fine-tuned model for image generation. For example, IP-Adapter uses the CLIP image encoder to extract facial features and embeds these features into the generated images through the cross-attention layer. IP-Adapter-FaceID further improves this process by replacing the CLIP image embedding with the face ID embedding obtained from the face recognition model, enhancing the identity consistency. PhotoMaker uses multiple facial image embeddings processed by the CLIP image encoder to achieve more accurate identity preservation. InstantID uses the Face Encoder and the Face ControlNet to jointly regulate face generation. The latest work, PULID, introduces semantic alignment loss and layout alignment loss to achieve effective identity ID customization while minimizing unexpected changes to the behavior of the original model.

[0005] However, based on the above existing technologies, the following disadvantages still exist: (1) The method based on fine-tuning adopts a two-stage mode of "training + generation" when in use. It requires inputting multiple image pictures of reference person IDs, and the effect of this mode is positively correlated with the scale of training data. Therefore, it often requires a large amount of image data support and a certain amount of training time, which also increases the user's usage cost.

[0006] (2) The existing methods without fine-tuning only focus on the preservation of facial identity, while ignoring the synchronous preservation of body identity, resulting in problems such as uncoordinated generated images (the generated person has a relatively small body proportion but a large head, a significant difference in skin color between the face and the body, and does not match the figure of the person in the reference image, etc.).

[0007] (3) The problem of limited text control ability caused by over-reliance on reference features. SUMMARY OF THE INVENTION

[0008] The technical problem to be solved by the present invention is to provide a method and a terminal for generating identity-preserving images, which do not require fine-tuning, have coordinated face and body, and have strong text control ability.

[0009] To solve the above technical problem, the technical solution adopted by the present invention is as follows: A method for generating identity-preserving images, comprising the steps of: S1. Receive a target person image and text prompt information input by the user; S2. Perform hierarchical extraction of facial features and global body features on the target person image, and perform feature fusion to generate fused features; generate text embedding features according to the text prompt information; S3. Input the text embedding features, the facial features, and the fused features into a pre-trained diffusion model to generate identity-preserving images; Wherein, the fused features are only injected into the identity-sensitive layer that has the greatest effect on identity retention in the diffusion model, and the facial features are only injected into other layers outside the identity-sensitive layer.

[0010] To solve the above technical problem, another technical solution adopted by the present invention is as follows: An identity-preserving image generation terminal, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the above-mentioned method for generating identity-preserving images.

[0011] The beneficial effects of the present invention are as follows: A method and a terminal for generating an identity-preserving image according to the present invention extract facial and body features in layers and adopt a selective injection strategy to limit the core identity features (fusion features) to the key layer (identity-sensitive layer). On the basis of ensuring the synchronous preservation of the face and the body, it avoids excessive interference with text-driven detail generation (such as hair color and clothing modification), and solves the problem of limited text control ability; without fine-tuning, only a single reference image is required, and the "one-person-one-training" mode is abandoned, reducing the generation time-consuming, overcoming the high-cost defect of the traditional fine-tuning method that relies on a large amount of data and training time, and achieving an efficient balance among identity fidelity, body coordination, and text semantics. Description of the Drawings

[0012] Figure 1 It is a flow example diagram of a method for generating an identity-preserving image according to an embodiment of the present invention; Figure 2 It is an example diagram of hierarchical feature extraction of a method for generating an identity-preserving image according to an embodiment of the present invention; Figure 3 It is an example diagram of the structure of a body global feature encoder of a method for generating an identity-preserving image according to an embodiment of the present invention; Figure 4 It is an example diagram of feature integration of a method for generating an identity-preserving image according to an embodiment of the present invention; Figure 5 It is an example diagram of diffusion model feature injection of a method for generating an identity-preserving image according to an embodiment of the present invention; Figure 6 It is an example diagram of the structure of a terminal for generating an identity-preserving image according to an embodiment of the present invention; Figure 7 It is an example diagram of the effect comparison between a method for generating an identity-preserving image according to an embodiment of the present invention and the prior art; Figure 8 It is an example diagram of the ablation experiment result of the effect of the hierarchical feature fusion mechanism; Figure 9 It is an example diagram of the ablation experiment result of the effect of the selective attention injection strategy; Label Description: 1. A terminal for generating an identity-preserving image; 2. A processor; 3. A memory. Detailed Embodiments

[0013] To illustrate the technical content, the achieved objectives, and the effects of the present invention in detail, the following is described in conjunction with the embodiments and the accompanying drawings.

[0014] Please refer to Figures 1 to 5 , a method for generating an identity-preserving image, including the steps of: S1. Receive the target person image and text prompt information input by the user; S2. Hierarchically extract the facial features and the global body features of the target person image, and perform feature fusion to generate fused features; generate text embedding features according to the text prompt information; S3. Input the text embedding features, the facial features, and the fused features into a pre-trained diffusion model to generate an identity-preserved image; Among them, the fused features are only injected into the identity-sensitive layer that has the greatest effect on identity retention in the diffusion model, and the facial features are only injected into other layers outside the identity-sensitive layer.

[0015] As can be seen from the above description, the beneficial effects of the present invention are as follows: A method and a terminal for generating an identity-preserved image of the present invention hierarchically extract facial and body features, and adopt a selective injection strategy to limit the core identity features (fused features) to the key layer (identity-sensitive layer). On the basis of ensuring the synchronous retention of the face and the body, it avoids excessive interference with text-driven detail generation (such as hair color and clothing modification), and solves the problem of limited text control ability; without fine-tuning, only a single reference image is required, abandoning the "one-person-one-training" mode, reducing the generation time-consuming, overcoming the high-cost defects of traditional fine-tuning methods that rely on a large amount of data and training time, and achieving an efficient balance of identity fidelity, body coordination, and text semantics.

[0016] Further, the extraction of the facial features and the global body features includes the steps of: Segment the target person image through an image segmentation algorithm to generate a panoramic person image; Use a face detection and recognition model to crop the face region image in the panoramic person image to generate a facial image; Input the panoramic person image and the facial image into a hierarchical identity extraction network, and use the global body feature encoder and the facial feature encoder therein to perform feature encoding on the panoramic person image and the facial image respectively to generate facial features and global body features.

[0017] As can be seen from the above description, panoramic segmentation is used to separate the person from the background, eliminating the interference of complex backgrounds on the learning of identity features; the facial image accurately locates the face region, and combines with the panoramic image to independently extract the microscopic facial features (such as facial features) and the macroscopic body features (such as body shape). This preprocessing process provides a pure input for subsequent hierarchical feature encoding, ensuring the synchronous and accurate capture of body features (such as posture and skin color) and facial features, and solving the core problem of "disproportionate head-to-body ratio and body identity mismatch" in the generated images from the source.

[0018] Further, the composition of the facial feature encoder includes the ArcFace algorithm and the CLIP image encoder; The facial feature encoder performing feature encoding on the facial image includes the steps of: The facial images are respectively input into the ArcFace algorithm and the CLIP image encoder, and the output results of the ArcFace algorithm and the CLIP image encoder are added after passing through the MLP layer respectively to obtain the facial features.

[0019] As can be seen from the above description, ArcFace extracts facial biometric features (such as highly recognizable details like cheekbones and nose bridges), and CLIP captures image semantic features (such as lighting and skin color). The two are fused to achieve fine-grained facial representation. Through the non-linear transformation of the MLP layer, the feature interaction is enhanced, making the facial region of the generated image highly consistent with the reference image in terms of structure (such as eye distance and lip shape) and appearance (such as skin texture and hair color), improving facial identity fidelity and reducing the facial distortion problem caused by traditional single encoders.

[0020] Furthermore, the body global feature encoder consists of a preset number of parallel feature stream branches. The preset number of feature stream branches is used to capture multi-granularity features from local texture to global pose, and each feature stream branch includes depthwise separable convolutional layers with different numbers of layers; The body global feature encoder performing feature encoding on the panoramic image of the person includes the steps of: Inputting the panoramic image of the person into the preset number of feature stream branches respectively for hierarchical feature extraction to obtain a preset number of branch features; Introducing an aggregation gate mechanism to integrate the outputs of the feature stream branches, and performing feature enhancement through a residual structure to generate body global features.

[0021] As can be seen from the above description, the multi-branch structure captures multi-granularity features from local texture (such as clothing patterns) to global pose (such as standing posture) in parallel through depthwise separable convolutional layers with different numbers of layers; the aggregation gate mechanism dynamically allocates weights to each branch to adapt to the feature fusion requirements of different body shapes and postures; the residual structure progressively enhances the feature expression. This design enables the model to accurately model macroscopic identity features such as body shape, skin color, and limb proportion, generating a body structure consistent with the reference image and solving the problems of "missing body features and uncoordinated generation" in the prior art.

[0022] Furthermore, the generation of the fused features includes the steps of: Using a learnable linear projection layer to project and transform the body global features into a dimensional space compatible with the facial features; Fusing the facial features and the projected body global features through a feature interaction network; Performing feature interaction through two layers of MLP and introducing a residual connection mechanism to generate fused features.

[0023] As described above, the linear projection layer maps the body global features (low-dimensional global vectors) and the facial features (high-dimensional fine-grained vectors) to a compatible space, avoiding the feature fragmentation caused by dimensional differences; the feature interaction network realizes cross-scale feature fusion through two layers of MLP, and the residual connection retains the original semantic information. This mechanism strengthens the correlation between facial and body features (such as the consistency of facial skin color and body skin color), ensures the integrity of the overall identity representation of the generated person, and improves the global coordination of the image.

[0024] Further, input the text embedding features, the facial features and the fused features into a pre-trained diffusion model, specifically injecting into the diffusion model by using a selective attention injection strategy: ; where, A represents the attention mechanism, l represents the attention layer, l id represents the identity-sensitive layer, and represent the balance coefficients, Q 、K t 、V t 、K f 、V f 、K lo and V lo are defined as follows: ; where, W q 、 、 、 、 and are learnable linear layers, Z is the latent variable of the current step of the diffusion model; E p represents the text embedding feature, E fused represents the fused feature, E f represents the facial feature.

[0025] As described above, through experiments, it is determined that the identity-sensitive layer (down_blocks.2.attentions.1) is the key position for identity retention, and only injecting the fused feature here can lock the core identity information (such as facial contour, body shape); injecting facial features (such as facial feature details) into other layers and combining with text embedding allows text-driven non-core feature adjustment (such as expression, clothing style). This strategy avoids the text out-of-control caused by the traditional "full-layer injection" and realizes a flexible generation mode of "core identity unchanged, details modified as needed".

[0026] Further, the identity-sensitive layer is specifically the second attention layer in the downsampling module of the diffusion model.

[0027] As can be seen from the above description, based on experimental verification, the second attention layer in the downsampling module of the diffusion model plays a decisive role in identity retention, accurately locating the key nodes for feature injection. Compared with blind full-layer injection, focusing on specific layers can reduce interference with the text processing path of the diffusion model, maximize the retention of the control of text prompts over the generated content while ensuring identity fidelity. For example, when modifying the clothing color of a reference person, the identity-sensitive layer ensures that the body shape of the person remains unchanged, and other layers allow text instructions to drive color adjustment to avoid feature conflicts.

[0028] Furthermore, the construction of the training dataset of the diffusion model includes: Obtaining a number of sample person images as targets; Generating corresponding text descriptions for each of the sample person images through a multimodal large language model as text prompt information; Segmenting the sample person images through an image segmentation algorithm to generate a panoramic person image; Cropping the face region image in the panoramic person image using a face detection and recognition model to generate a facial image; Constructing a training dataset from the facial image, the panoramic person image, the text prompt information, and the sample person images; Among them, the facial image is used to generate facial features input to the diffusion model; the panoramic person image is used to generate global body features, and feature fusion is performed with the facial features to generate fused features input to the diffusion model; the text prompt information is used to generate text embedding features.

[0029] As can be seen from the above description, a large amount of multimodal data (image + text) enables the model to learn the common identity representation rules across people (such as the correlation between body shape and facial features), getting rid of the inefficient mode of "one-person-one-fine-tuning"; the separate annotation of the panoramic image and the facial image forces the model to explicitly model the collaborative relationship between body and facial features, improving the identity generalization ability during zero-shot generation. The strong data foundation in the pre-training stage provides underlying support for "single-image generation and fast response" in the inference stage.

[0030] Furthermore, the training of the diffusion model is divided into two stages: In the first stage, only the diffusion loss function is used to train the model to obtain the first-stage model; In the second stage, an identity loss function and the diffusion loss function are introduced to simultaneously perform second-stage training on the first-stage model to obtain the trained diffusion model.

[0031] As described above, in the first stage, only the diffusion loss is optimized to ensure that the model masters the basic image generation capabilities (such as composition and color); in the second stage, the identity loss (face feature constraint based on ArcFace) is superimposed to directionally enhance the identity preservation ability. Progressive training avoids premature overfitting of identity features, and at the same time ensures that the generated images not only conform to the natural image distribution, but also can accurately reproduce the facial and body features of the reference person, improving the stability and practicality of the model.

[0032] A method and a terminal for generating identity-preserving images according to the present invention are applicable to scenarios with the need for generating identity-preserving images, such as AI travel photography, AI photo shoots, animation scene design, etc.

[0033] Please refer to Figures 1 - 5 , and the first embodiment of the present invention is as follows: A method for generating identity-preserving images includes the steps of: S1. Receive a target person image and text prompt information input by the user; S2. Perform hierarchical extraction of facial features and body global features on the target person image, and perform feature fusion to generate fused features; The extraction of the facial features and body global features includes the steps of: Segment the target person image through an image segmentation algorithm to generate a panoramic person image; Use a face detection and recognition model to crop the face region image in the panoramic person image to generate a facial image.

[0034] In this embodiment, reference can be made to Figure 2 As shown, through a panoramic segmentation algorithm, pixel-level separation of the person body and the background is achieved to generate a foreground mask image of the person without background interference (panoramic person image) I1, effectively eliminating the interference of the complex environmental background on the learning of identity features, and using a face detection and recognition model to crop the face region image (facial image) I2. The image size is uniformly set. In this embodiment, it is set to 512×512, and in other equivalent embodiments, it can be set to other fixed sizes.

[0035] Input the panoramic person image and the facial image into a hierarchical identity extraction network, and use the body global feature encoder and the facial feature encoder in it to perform feature encoding on the panoramic person image and the facial image respectively to generate facial features and body global features.

[0036] In this embodiment, the hierarchical identity extraction network is composed of a facial feature encoder and a body feature encoder in parallel.

[0037] The composition of the facial feature encoder includes the ArcFace algorithm and the CLIP image encoder; The facial feature encoder encodes the facial map, including the steps of: Input the facial map into the ArcFace algorithm and the CLIP image encoder respectively, and add the output results of the ArcFace algorithm and the CLIP image encoder after passing through the MLP layer respectively to obtain the facial features.

[0038] In this embodiment, the facial feature encoder consists of the ArcFace and the CLIP image encoder, mainly extracting fine-grained facial features. Input the facial map into the ArcFace algorithm and the CLIP image encoder respectively, and add the output results of the ArcFace algorithm and the CLIP image encoder after passing through the MLP layer respectively to obtain the facial features E f : ; Among them, MLP 1() and MLP 2() respectively represent the MLP layer processing on the output results of the ArcFace and the CLIP image encoder; I 2 represents the facial map, ArcFace () represents the processing of the ArcFace algorithm, C i () represents the processing of the CLIP image encoder.

[0039] The body global feature encoder consists of a preset number of parallel feature stream branches. The preset number of feature stream branches is used to capture multi-grained features from local texture to global pose. Each feature stream branch includes depthwise separable convolutional layers with different numbers of layers; The body global feature encoder encodes the person panorama, including the steps of: Input the person panorama into the preset number of feature stream branches respectively for hierarchical feature extraction to obtain a preset number of branch features; Introduce an aggregation gate mechanism to integrate the outputs of each feature stream branch, and perform feature enhancement through a residual structure to generate body global features.

[0040] In this embodiment, reference can be made to Figure 2 and Figure 3 , where the structure of the body global feature encoder can be referred to Figure 3 . In order to effectively reduce parameters, depthwise separable convolutions are introduced to replace standard convolutions, and a multi-scale feature pyramid is constructed. For the input, T parallel feature stream branches are constructed, and each branch contains depthwise separable convolutional layers with different numbers of layers, aiming to capture multi-grained features from local texture to global pose. Input the person panorama into the preset number of feature stream branches respectively for hierarchical feature extraction to obtain a preset number of branch features: ; Among them, x t represents the branch feature extracted by the t-th feature stream branch represents the depthwise separable convolution operation of the t-th feature stream branch, I 1 represents the panoramic view of the person, T represents the number of the feature stream branches; To achieve the deep learning of full-scale features, the traditional fixed weight allocation method is abandoned, and instead the Aggregation Gate (AG) mechanism is adopted to integrate the outputs of each feature stream in a dynamic and flexible manner. The Aggregation Gate mechanism is introduced to integrate the outputs of each feature stream branch, and feature enhancement is performed through a residual structure to generate the body global feature E b : ; Among them, G() represents the weight generation function for the Aggregation Gate mechanism to fuse the branch features, ⊙ represents element-wise multiplication, and x represents the original input feature.

[0041] The generation of the fused feature includes the steps of: Using a learnable linear projection layer to project and transform the body global feature into a dimensional space compatible with the facial feature; Fusing the facial feature and the projected body global feature through a feature interaction network; Performing feature interaction through two layers of MLP and introducing a residual connection mechanism to generate the fused feature.

[0042] In this embodiment, as Figure 4 shown, to achieve the collaborative optimization of the global identity feature and the facial fine-grained feature, a hierarchical feature fusion module is constructed. Compared with the basic method of injecting two types of feature embeddings in parallel through a decoupled cross-attention mechanism, the representational integrity of the identity feature is significantly improved through a two-stage refinement process in this embodiment. First, use the learnable linear projection layer L to project the body global feature into a dimensional space compatible with the facial feature , so as to obtain: ; Among them, represents the body global feature mapped to the same dimension as the facial feature; L() represents the processing of the linear projection layer.

[0043] Then, a feature interaction network is constructed to achieve the deep fusion of the global feature and the facial feature. By connecting the facial identity embedding E f with the projected global feature , obtain: ; Subsequently, feature interaction is achieved through two layers of MLP, and a residual connection mechanism is introduced to preserve the original identity semantics, obtaining fused features : .

[0044] Generate text embedding features according to the text prompt information.

[0045] In this embodiment, the text description P is input into the CLIP text encoder to obtain the text embedding feature E P .

[0046] S3. Input the text embedding feature, the facial feature, and the fused feature into a pre-trained diffusion model to generate an identity-preserving image; Among them, the fused feature is only injected into the identity-sensitive layer in the diffusion model that has the greatest effect on identity preservation, and the facial feature is only injected into other layers outside the identity-sensitive layer.

[0047] In preliminary experiments, an attempt was made to embed the fused feature into all cross-attention layers of the diffusion model. Although this method achieved a high identity fidelity, it led to serious text runaway: the model overly retained the detailed features of the reference image (such as clothing texture, hair color, pose, etc.), thus losing the ability to respond to text prompts. For example, when attempting to modify the hair color and clothing of a character, the model was severely affected by the original features of the reference character and could not follow the input text prompt.

[0048] To solve this problem, the influence of each attention layer on identity feature preservation was analyzed. The experimental results show that the second attention layer (down_blocks.2.attentions.1) in the diffusion Unet downsampling module plays a crucial role in identity preservation. Therefore, in this embodiment, down_blocks.2.attentions.1 is defined as the identity-sensitive layer, and a selective attention injection strategy is designed.

[0049] Input the text embedding feature, the facial feature, and the fused feature into the pre-trained diffusion model, specifically by injecting them into the diffusion model using the selective attention injection strategy: ; where A represents the attention mechanism, l represents the attention layer, l id represents the identity-sensitive layer, and represent the balance coefficients, Q , K t, V t , K f , V f , K lo and V lo are defined as follows: ; where W q , , , , and are learnable linear layers, Z is the latent variable of the current step of the diffusion model; E p represents the text embedding feature, E fused represents the fusion feature, E f represents the facial feature.

[0050] Specifically, refer to Figure 5 , the fusion feature E fused is only injected into the identity-sensitive layer, while the facial feature E f is injected into other layers. This strategy effectively avoids excessive interference from the original detailed features of the person while retaining the identity information, thus achieving a delicate balance between identity retention and text fidelity.

[0051] The selective attention injection strategy injects E p , E f and E fused into the diffusion model to achieve identity-preserving image generation.

[0052] Embodiment 2 of the present invention is: A method for generating identity-preserving images, which is different from Embodiment 1 in that the training of the diffusion model is described in this embodiment.

[0053] The construction of the training dataset of the diffusion model includes: Obtaining a number of sample person images as targets; Generating corresponding text descriptions for each of the sample person images through a multimodal large language model as text prompt information; Segmenting the sample person images through an image segmentation algorithm to generate a person panorama; Cropping the face region image in the person panorama by using a face detection and recognition model to generate a facial image; Constructing a training dataset from the facial image, the person panorama, the text prompt information, and the sample person images; Among them, the facial image is used to generate facial features for inputting into the diffusion model; the full-body panoramic image of the person is used to generate global body features, and feature fusion is performed with the facial features to generate fused features for inputting into the diffusion model; the text prompt information is used to generate text embedding features.

[0054] In this embodiment, to train the identity-preserving image generation model, a large number of full-body and half-body color portrait images are collected as targets (sample size ≥ 500,000), and the corresponding text description P is generated for each image through the BLIP-2 model. Pixel-level separation of the person's body from the background is achieved through the panoramic segmentation algorithm to generate an interference-free foreground mask image of the person (full-body panoramic image of the person) I1, effectively eliminating the interference of complex backgrounds on the learning of identity features, and a face detection and recognition model is used to crop out the face area image (facial image) I2. Finally, the image size is uniformly set to 512×512 to form a training dataset for deep learning.

[0055] After the training dataset undergoes hierarchical feature extraction and even feature fusion, the diffusion model is input for training.

[0056] The training of the diffusion model is divided into two stages: In the first stage, only the diffusion loss function is used for model training to obtain the first-stage model; In the second stage, an identity loss function and the diffusion loss function are introduced to simultaneously perform second-stage training on the first-stage model to obtain the trained diffusion model.

[0057] In this embodiment, cross-modal feature interaction will inevitably lead to a decrease in the facial detail fidelity in the generated images. To solve this problem, a facial id loss function is introduced during the training process. The traditional method of calculating the facial id loss during diffusion training usually involves one-step prediction to obtain the x 0 at the t-th time step; however, the x 0 prediction generated by this method has a large amount of noise, seriously affecting the accuracy of the facial id loss calculation.

[0058] To improve this situation, the SDXL-Lightning model is adopted in this embodiment, which can quickly generate accurate and identity-conditioned x 0 from pure noise in only 4 steps. Calculating the facial id loss on the x 0 that is highly similar to the true data distribution can significantly improve the accuracy. Specifically, the facial id loss L id is defined as: ; Among them, xRepresents the facial region of the reference image, x’ Represents the facial region of the model-generated image, f Represents the Arcface face recognition network, CosSim Represents the cosine similarity between the features extracted from the facial regions of the images.

[0059] Meanwhile, the noise loss function of the diffusion model L noise is used as the optimization loss of the model, expressed as: ; During the forward process of the diffusion model, Gaussian noise is sampled and added to the data sample z 0, resulting in a noisy sample t at the time step z t . During the reverse process, z t , t and the text prompt embedding generated by the CLIP text encoder C are used as the input to the noise prediction model . The noise prediction model of the text-to-image diffusion model is mainly composed of a UNet architecture, which includes residual blocks, self-attention layers, and cross-attention layers.

[0060] Therefore, the complete objective function is: ; where is the balance coefficient. During training, only the learnable linear layers in the body feature encoder, hierarchical feature fusion module, and image-conditioned cross-attention layer are optimized according to this objective, and the rest remains unchanged.

[0061] In this embodiment, during training, 8 A800 GPUs (each with 80GB of memory) were used, the learning rate was set to 1e-4, and the model was trained for 200,000 steps with a batch size of 24. To be consistent with the CFG training method, a random ignoring strategy was introduced, that is, the conditional embeddings related to the person foreground image I1, facial image I2, and text condition P were ignored separately or simultaneously with a probability of 0.02, thereby increasing the variability during the training process. The training process was divided into two stages. In the first stage, only the traditional diffusion loss function was used to train the model. After the first stage was completed, the identity loss was introduced on the basis of the model in the first stage, and the second stage of training was carried out simultaneously with the diffusion loss

[0062] Please refer to Figure 6 , Embodiment 3 of the present invention is: A terminal 1 for generating an identity-preserving image comprises a processor 2, a memory 3 and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, the steps in the method for generating an identity-preserving image described in the first or second embodiment above are implemented.

[0063] In summary, the present invention provides a method and terminal for generating an identity-preserving image, which extracts facial and body features in layers and adopts a selective injection strategy to limit the core identity features (fusion features) to the key layer (identity-sensitive layer). On the basis of ensuring the synchronization of face and body, it avoids excessive interference with text-driven detail generation (such as hair color and clothing modification), thereby solving the problem of limited text control ability. It does not require fine-tuning, only requires a single reference image, abandons the "one person one training" model, reduces generation time, overcomes the high cost defects of traditional fine-tuning methods that rely on large amounts of data and training time, and achieves an efficient balance between identity fidelity, body coordination and text semantics.

[0064] The present invention can achieve the following advantages: 1. Only a single reference image is needed to achieve identity-preserving image generation, without the need for additional data collection or time-consuming fine-tuning, and the time required for a single inference is reduced to seconds. The method of this invention is compared with several state-of-the-art fine-tuning-free methods, and the comparison results are presented in Figure 7 .

[0065] 2. Through the collaborative design of the fine-grained facial encoder and the global body encoder, the rationality of the generation of body structure (such as body shape and posture) is significantly improved while accurately maintaining facial identity features (such as facial features and skin color), solving the core problem of the incoordination between limbs and facial features in traditional methods. Combined with the hierarchical feature fusion mechanism, the integrity of identity representation is further strengthened to ensure that the details such as body proportions and skin texture of the generated character are highly consistent with the real identity. Ablation experiments were conducted, and the results (such as Figure 8 The effectiveness of this design is demonstrated.

[0066] 3. By introducing a selective attention injection strategy, we strengthen the control of text conditions while retaining identity information, and achieve an efficient balance between identity fidelity and text semantics. We conducted an ablation experiment and the results (such as Figure 9 The effectiveness of this design is demonstrated.

[0067] The present invention constructs a universal hierarchical identity feature extraction network, learns the common identity representation rules across people through massive data pre-training, completely abandons the inefficient mode of "one person one fine-tuning" in traditional methods, and realizes rapid personalized generation under zero-sample conditions; The global body encoder explicitly models macroscopic identity features such as body shape and skin color, which complement the microscopic features of the face encoder (such as pupil texture and lip curvature). Through the hierarchical feature fusion module, cross-scale feature alignment is achieved, fundamentally solving the problem of limb identity mismatch caused by single-face encoding. At the same time, based on the face ID loss function of the face recognition model, by constraining the similarity of the deep features between the generated face and the target identity, the generation accuracy of highly recognizable regions such as cheekbones and nasal bridges is directionally enhanced; By conducting experiments, the influence of each attention layer on the retention of identity features was analyzed. It was observed that the identity-sensitive layer plays a crucial role in identity retention, and based on this, a selective attention injection strategy was designed, which not only avoids excessive interference with the text-driven generation process but also ensures the high-fidelity transmission of the core identity features.

[0068] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in the relevant technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for generating an identity-preserving image, characterized in that, Including the steps: S1. Receive the target person image and text prompt information passed in by the user; S2. Perform hierarchical extraction of facial features and body global features on the target person image, and perform feature fusion to generate fused features; Generate text embedding features according to the text prompt information; S3. Input the text embedding features, the facial features, and the fused features into a pre-trained diffusion model to generate an identity-preserving image; Among them, the fused features are only injected into the identity-sensitive layer in the diffusion model that has the greatest effect on identity retention, and the facial features are only injected into other layers outside the identity-sensitive layer.

2. The method for generating an identity-preserved image according to claim 1, wherein The extraction of the facial features and body global features includes the steps: Segment the target person image through an image segmentation algorithm to generate a panoramic person image; Use a face detection and recognition model to crop the face region image in the panoramic person image to generate a facial image; Input the panoramic person image and the facial image into a hierarchical identity extraction network, and use the body global feature encoder and the facial feature encoder in it to perform feature encoding on the panoramic person image and the facial image respectively to generate facial features and body global features.

3. The method for generating an identity-preserving image according to claim 2, wherein The composition of the facial feature encoder includes the ArcFace algorithm and the CLIP image encoder; The facial feature encoder performing feature encoding on the facial image includes the steps: Input the facial image into the ArcFace algorithm and the CLIP image encoder respectively, and add the output results of the ArcFace algorithm and the CLIP image encoder after passing through the MLP layer respectively to obtain the facial features.

4. The method for generating an identity-preserved image according to claim 2, wherein The composition of the body global feature encoder includes a preset number of parallel feature stream branches, and the preset number of feature stream branches are used to capture multi-granularity features from local texture to global pose, and each feature stream branch includes depthwise separable convolutional layers with different numbers of layers; The body global feature encoder performing feature encoding on the panoramic person image includes the steps: Input the panoramic person image into the preset number of feature stream branches respectively, perform hierarchical feature extraction to obtain a preset number of branch features; Introduce an aggregation gate mechanism to integrate the outputs of each feature stream branch, and perform feature enhancement through a residual structure to generate body global features.

5. The method for generating an identity-preserving image according to claim 1, wherein The generation of the fused features includes the steps: Use a learnable linear projection layer to transform and project the body global features into a dimension space compatible with the facial features; Fuse the facial features and the projected body global features through a feature interaction network; Perform feature interaction through two-layer MLP and introduce a residual connection mechanism to generate fused features.

6. The method for generating an identity-preserved image according to claim 1, wherein Input the text embedding features, the facial features, and the fused features into a pre-trained diffusion model, specifically using a selective attention injection strategy to inject into the diffusion model: ; Among them, A represents the attention mechanism, l represents the attention layer, l id represents the identity-sensitive layer, and represents the balance coefficient, Q 、K t 、V t 、K f 、V f 、K lo and V lo are defined as follows: ; Among them, W q , , , , and are learnable linear layers, Z is the latent variable of the diffusion model at the current step; E p represents text embedding features, E fused represents fused features, E f represents facial features.

7. A method for generating an identity-preserving image according to claim 1 or 6, characterized in that, The identity-sensitive layer is specifically the second attention layer in the downsampling module of the diffusion model.

8. The method for generating an identity-preserving image according to claim 1, wherein The construction of the training dataset of the diffusion model includes: Obtain a number of sample person images as targets; Generate corresponding text descriptions for each sample person image through a multi-modal large language model as text prompt information; Segment the sample character image through an image segmentation algorithm to generate a full-body character image; Use a face detection and recognition model to crop the face region image in the full-body character image to generate a facial image; Construct a training dataset from the facial image, the full-body character image, the text prompt information, and the sample character image; Among them, the facial image is used to generate facial features input to the diffusion model; the full-body character image is used to generate global body features, and the global body features are fused with the facial features to generate fused features input to the diffusion model; the text prompt information is used to generate text embedding features.

9. The method for generating an identity-preserving image according to claim 1, wherein The training of the diffusion model is divided into two stages: In the first stage, only use the diffusion loss function to train the model to obtain the first-stage model; In the second stage, introduce the identity loss function and the diffusion loss function to simultaneously perform the second-stage training on the first-stage model to obtain the trained diffusion model.

10. A generation terminal for identity-preserving images, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps in the method for generating an identity-preserved image according to any one of claims 1-9 above.

Citation Information

Patent Citations

  • Identity preserving diffusion model fine tuning method and system

    CN118351576A

  • Virtual model clothing display image intelligent generation method and device based on diffusion model

    CN119131212A

  • Deep learning-based expression editing model training method

    CN119181124A

  • Face identity preservation for image-to-image models using stable diffusion generative model

    US20240265498A1

  • Facial expression-based detection method for deepfake by generative artificial intelligence (AI)

    US20240378921A1

Cited By

  • Image generation identity personalization method based on diffusion model

    CN120953049A