Character image generation method and device, electronic equipment and readable storage medium

By acquiring and processing character description information and face features, and using the self-attention and cross-attention mechanism in the diffusion model to generate target images, the problem of lack of real-time and accuracy in generating portrait images in the prior art is solved, and high-quality and diverse image generation is achieved.

CN120340081APending Publication Date: 2025-07-18SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510182934.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The portrait images generated in the prior art lack real-time or accuracy, resulting in poor user experience.

Method used

By obtaining the text features, face identity features, face key points features and face reference map features of character description information, we use the shared self-attention mechanism, cross attention mechanism and diffusion model of shared feedforward network to generate target image feature vectors, and finally generate high-quality and diverse target images.

Benefits of technology

The generated target image can better combine the appearance and identity information of the portrait reference image to effectively process multimodal information, improving the quality and accuracy of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340081A_ABST
    Figure CN120340081A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and provides a figure image generation method and device, electronic equipment and a readable storage medium. The method comprises the following steps: acquiring text features corresponding to figure description information, and extracting face identity features, face key point features and face reference image features; inputting the text feature and the first image feature into a preset diffusion model to obtain a text processing feature and a first image processing feature; obtaining a second image feature, and inputting the second image feature into a preset diffusion model to obtain a second image processing feature; based on the face identity feature, the face key point feature and the face reference image feature, obtaining a face identity comprehensive feature, and inputting the face identity comprehensive feature into a preset diffusion model to obtain a third image processing feature; and obtaining a target image feature vector based on the text processing feature and the third image processing feature, and generating a target image. Through the technical scheme provided by the invention, the quality and accuracy of image generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to a method, apparatus, electronic device, and readable storage medium for generating a human image. Background Art

[0002] A way to generate a high-fidelity portrait image with specific identity features can be to create a personalized portrait image using a single reference image. Further, text-driven image generation technology has received great attention because it can generate images according to the text prompts received by users.

[0003] In the prior art, the model can be controlled by means of fine-tuning such as trying LoRA and texture flipping to achieve precise control of the specific identity information of the user, or the model can be guided by text prompts to recognize the features of a specific subject. However, the portrait images generated by the above methods lack real-time performance or accuracy, resulting in a poor user experience. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a method, apparatus, electronic device, and readable storage medium for generating a human image to solve the problem that the generated portrait images in the prior art lack real-time performance or accuracy.

[0005] In a first aspect of the embodiments of the present disclosure, a method for generating a human image is provided, including:

[0006] Obtain text features corresponding to human description information, and extract face identity features, face key point features, and face reference map features from a portrait reference map;

[0007] Input the text features and first image features into a preset diffusion model to process the first image features and text features through a shared self-attention mechanism to obtain text processing features and first image processing features, where the first image features are obtained by adding noise to the face reference map features;

[0008] Obtain second image features after linear processing of the face reference map features, and input the second image features into the preset diffusion model to process the first image processing features and the second image features through a first cross-attention mechanism to obtain second image processing features;

[0009] Based on the face identity features, face key point features, and face reference map features, obtain a face identity comprehensive feature, and input the face identity comprehensive feature into the preset diffusion model to process the face identity comprehensive feature and the second image processing features through a second cross-attention mechanism to obtain third image processing features;

[0010] Based on the text processing feature and the third image processing feature, a target image feature vector is obtained through a shared feed-forward network in a preset diffusion model, and a target image is generated based on the target image feature vector.

[0011] In a second aspect of the embodiments of the present disclosure, a device for generating a human image is provided, including:

[0012] An extraction module configured to obtain a text feature corresponding to human description information and extract a face identity feature, a face key point feature, and a face reference image feature from a portrait reference image;

[0013] A first processing module configured to input the text feature and a first image feature into a preset diffusion model to process the first image feature and the text feature through a shared self-attention mechanism, so as to obtain a text processing feature and a first image processing feature, where the first image feature is obtained by performing noise addition processing on the face reference image feature;

[0014] A second processing module configured to obtain a second image feature after linear processing of the face reference image feature, and input the second image feature into a preset diffusion model to process the first image processing feature and the second image feature through a first cross-attention mechanism, so as to obtain a second image processing feature;

[0015] A third processing module configured to obtain a comprehensive face identity feature based on the face identity feature, the face key point feature, and the face reference image feature, and input the comprehensive face identity feature into a preset diffusion model to process the comprehensive face identity feature and the second image processing feature through a second cross-attention mechanism, so as to obtain a third image processing feature;

[0016] A generation module configured to obtain a target image feature vector through a shared feed-forward network in a preset diffusion model based on the text processing feature and the third image processing feature, and generate a target image based on the target image feature vector.

[0017] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0018] In a fourth aspect of the embodiments of the present disclosure, a readable storage medium is provided. The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0019] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: obtaining the text features corresponding to the person description information, and extracting the face identity features, face key point features, and face reference map features from the portrait reference map, inputting the text features and the first image features into a preset diffusion model to process the first image features and the text features through a shared self-attention mechanism to obtain text processing features and first image processing features, where the first image features are obtained by adding noise to the face reference map features, obtaining the second image features after linear processing of the face reference map features, and inputting the second image features into the preset diffusion model to process the first image features and the second image features through a first cross-attention mechanism to obtain second image processing features, obtaining a face identity comprehensive feature according to the face identity features, face key point features, and face reference map features, and inputting the face identity comprehensive feature into the preset diffusion model to process the face identity comprehensive feature and the second image processing features through a second cross-attention mechanism, obtaining a target image feature vector according to the text features and the third image processing features through a shared feed-forward network in the preset diffusion model, and generating a target image according to the target image feature vector. It is possible to better combine the appearance and identity information of the portrait reference map by using the face identity comprehensive feature obtained by fusing the face identity features, face key point features, and face reference map features, and be able to adjust the generated target image based on the text information during the process of generating the target image, more effectively process multi-modal information, and improve the quality and accuracy of image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of a method for generating a person image provided by an embodiment of the present disclosure;

[0022] Figure 2 It is a flowchart of another method for generating a person image provided by an embodiment of the present disclosure;

[0023] Figure 3 It is a schematic diagram of a device for generating a person image provided by an embodiment of the present disclosure;

[0024] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0026] A method and apparatus for generating a human image according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0027] Customized image generation technology creates visual content consistent with the identity of the person in a portrait reference image. In the prior art, generating high-fidelity portrait images with specific identity features (i.e., the target images in this application) is a hot topic in the field of content generation, aiming to create personalized portrait images using a single portrait reference image. Text-driven image generation technology has received great attention because it can generate target images according to text prompts input by users. However, traditional customization methods, which rely on additional controls to ensure that the generated target images meet user requirements, tend to neglect the precise control of the identity information of the person in the portrait reference image and often struggle to strike a balance between training efficiency and identity information protection. For example, using fine-tuning methods such as LoRA and texture inversion to control the model to achieve precise control of specific identity information requires separate training for the identity in each portrait reference image, which is a cumbersome process and lacks real-time inference capabilities. In addition, there are also methods that use text prompts to guide the model to identify the features of specific subjects, but the semantic-based approach has limited accuracy, requires corresponding identity text information for the model, and is difficult to handle identity changes in each portrait reference image. Therefore, the embodiments of this application propose a method for generating a human image that can make the details of the generated target image richer and more delicate, thereby better capturing and reproducing the individual identity features of the portrait reference image, and can accurately control the identity information of each portrait reference image to ensure the high quality and consistency of the generated target image.

[0028] Figure 1 It is a schematic flowchart of a method for generating a human image provided by an embodiment of the present disclosure. Figure 1 The method for generating a human image can be executed by a terminal device or a server. As Figure 1 shown, the method for generating a human image includes:

[0029] S101, obtaining text features corresponding to the human description information, and extracting face identity features, face key point features, and face reference image features from the portrait reference image.

[0030] Specifically, corresponding text features are extracted from the received person description information. Among them, the text features can be extracted from the text information input by the user. The text information can be the text information describing the portrait reference picture or the text information of the user's expectation for the target image. For example, "smile" can be input, and then the expression of the portrait reference picture will be changed to a smiling expression based on the portrait reference picture for the target image.

[0031] Among them, the text features can be used to represent information such as the appearance, style, and emotion of the person.

[0032] In addition, it is necessary to extract the face identity feature, face key point feature, and face reference picture feature from the portrait reference picture. Among them, the face identity feature is used to represent the uniqueness of the person in the portrait reference picture, the face key point feature is used to capture the structural information of the person's face in the portrait reference picture, such as the positions of the eyes, mouth, nose, etc., and the face reference picture feature is used to provide the visual information of the portrait reference picture.

[0033] Through the text features extracted from the person description information and the face identity feature, face key feature, and face reference picture feature extracted from the portrait reference picture, rich semantic and visual information can be provided for the generation of the target image. Among them, the face identity feature, face key feature, and face reference picture feature can accurately describe the appearance and structural information of the person in the portrait reference picture.

[0034] S102, input the text features and the first image features into a preset diffusion model to process the first image features and the text features through a shared self-attention mechanism, and obtain text processing features and first image processing features.

[0035] Among them, the first image features are obtained by adding noise to the face reference picture features.

[0036] Specifically, input the text features and the first image features into a preset diffusion model. In the preset diffusion model, the first image features and the text features are processed through a shared self-attention mechanism to obtain text processing features and first image processing features respectively. Among them, the preset diffusion model can be a multimodal diffusion backbone MM-DiT architecture (MultimodalDiffusion Backbone).

[0037] Among them, the first image feature is obtained by adding noise to the face reference image feature, which is used to introduce randomness to the face reference image feature, so that the generated target image has a certain degree of diversity. The preset diffusion model is a generative model that can generate high-quality target images by gradually removing noise; the shared self-attention mechanism is an attention mechanism used to process image features and text features simultaneously, enabling the two to influence each other. Adding noise can introduce random noise into the image features and increase the diversity of the generated target images.

[0038] Among them, the first image feature can be obtained by splicing the noise map and the face reference image feature.

[0039] By adding noise to the face reference image feature, the diversity of the generated target images is increased, avoiding the result of the target image being too single. By using the shared self-attention mechanism to process the first image feature and the text feature, the text feature and the image feature can influence each other, enhancing the fusion ability of the preset diffusion model for semantic information and visual information.

[0040] S103: Obtain the second image feature after linearly processing the face reference image feature, and input the second image feature into the preset diffusion model to process the first image processing feature and the second image feature through the first cross-attention mechanism to obtain the second image processing feature.

[0041] Specifically, linearly process the face reference image feature to obtain the second image feature and input it into the preset diffusion model, so that the first cross-attention mechanism in the preset diffusion model processes the first image processing feature and the second image feature to obtain the second image processing feature.

[0042] Among them, the second image feature retains the original information of the face reference image, and after linear processing, it can better interact with other features in the preset diffusion model.

[0043] Among them, the first cross-attention mechanism can analyze the correlation between the first image processing feature and the second image feature, and weight and adjust the above two features according to the correlation to obtain the second image processing feature. Through the first cross-attention mechanism, the deep fusion between the first image processing feature and the second image feature is realized, improving the expression ability of the features, enabling the preset diffusion model to better understand the semantic and visual information of the image.

[0044] Among them, linear processing is to perform a linear transformation on the face reference image feature to adjust the dimension and distribution of the face reference image feature to make it more suitable for subsequent processing steps; the second image feature is the feature of the face reference image feature after linear processing; the first cross-attention mechanism is an attention mechanism used to process the first image processing feature and the second image feature, enabling the two to influence each other.

[0045] S104. Based on the face identity features, face key-point features, and face reference image features, obtain the comprehensive face identity features, and input the comprehensive face identity features into a preset diffusion model to process the comprehensive face identity features and the second image processing features through the second cross-attention mechanism to obtain the third image processing features.

[0046] Specifically, based on the face identity features, face key-point features, and face reference image features, comprehensively obtain the comprehensive face identity features. The comprehensive face identity features include the identity information of the person and also fuse the facial structure and visual style information.

[0047] In addition, input the comprehensive face identity features into the preset diffusion model, and through the second cross-attention mechanism in the preset diffusion model, make the comprehensive face identity features interact with the second image processing features to obtain the third processing features. The third processing features fuse the identity information of the person, the facial structure information, and the visual style information of the image, providing high-quality input for the subsequent image generation steps.

[0048] Among them, the role of the second cross-attention mechanism is to further fuse the comprehensive face identity features and the second image processing features, enabling the preset diffusion model to better understand the identity and appearance features of the person. The second cross-attention mechanism can analyze the mutual relationship between the above two features and adjust and optimize the features according to this relationship.

[0049] S105. Based on the text processing features and the third image processing features, through the shared feed-forward network in the preset diffusion model, obtain the target image feature vector, and generate the target image based on the target image feature vector.

[0050] Specifically, input the text processing features and the third image processing features into the shared feed-forward network in the preset diffusion model. The shared feed-forward network deeply fuses the text processing features and the third image processing features through a series of non-linear transformations and inter-layer interactions, can retain the semantic information of the text processing features, and combine the visual details of the third image processing features, avoiding information conflict or loss, so that the generated target image feature vector can reflect both the appearance and semantic requirements of the person in the portrait reference image.

[0051] Among them, the shared feed-forward network is a multi-layer neural network structure, and its core function is to fuse and transform features from different sources; the target image feature vector is a high-dimensional representation that synthesizes all the key information of the text description and the portrait reference image and is the core basis for generating the target image.

[0052] In addition, a target image is generated based on the target image feature vector. The target image is not only visually similar to the reference image but also meets the semantic requirements in the text description, such as details of the person's expression, hairstyle, clothing, etc. The guidance of the target image feature vector enables the generated image to not only conform to the person description but also retain the style and details of the reference image.

[0053] According to the technical solution provided by the embodiments of the present disclosure, high-quality and diverse target images can be generated while retaining the identity information and appearance features of the person in the portrait reference image. The target image is similar in style to the reference image but has diversity.

[0054] In some embodiments, based on the face identity feature, face key point feature, and face reference image feature, a face identity comprehensive feature is obtained, including:

[0055] The face identity feature and the face key point feature are input into a preset identity network model to obtain a first intermediate feature;

[0056] The position box mask of the person in the portrait reference image is obtained, and the first intermediate feature is masked by the position box mask to obtain a second intermediate feature;

[0057] The face reference image feature and the second intermediate feature are processed through a third cross-attention mechanism to obtain the face identity comprehensive feature.

[0058] Specifically, the face identity feature and the face key point feature are input into a preset identity network model. The preset identity network model can be IdentityNet, a portrait control network based on a diffusion model. The preset identity network can learn the feature representation related to the structure and identity of the specific face according to the face identity feature and the position information of the face key points, and output an intermediate feature that fuses the identity feature and the face key points, that is, the first intermediate feature. The first intermediate feature contains the identity information of the person in the portrait reference image and also retains the key details of the facial structure.

[0059] In addition, the position box mask of the person is obtained from the portrait reference image. The position box mask is a two-dimensional mask matrix used to mark the position of the person in the image, usually a binary matrix, that is, the value of the person area is 1 and the value of the background area is 0.

[0060] The position box mask is applied to the first intermediate feature. Through the masking process, only the feature information of the person area is retained, while the background part is ignored, so as to ensure that the features are concentrated on the person itself and avoid the interference of background noise on subsequent processing.

[0061] After mask processing, a new feature is obtained, called the second intermediate feature. The second intermediate feature focuses more on the portrait area in the portrait reference image and can more accurately reflect the identity and structural information of the portrait.

[0062] In addition, through the third cross-attention mechanism, the face reference image feature is fused with the second intermediate feature to achieve effective fusion of the two, and a comprehensive face identity feature is obtained.

[0063] Among them, the comprehensive face identity feature combines the advantages of face identity features, face key point features, and face reference image features. It not only contains the identity and structural information of the person in the portrait reference image but also retains the visual details and style of the original image. The comprehensive face identity feature is a comprehensive and accurate representation that can provide high-quality input for face image generation or editing tasks.

[0064] Among them, the third cross-attention mechanism is a feature fusion mechanism that can analyze the mutual relationship between the face reference image feature and the second intermediate feature and adjust the features according to this relationship. Specifically, the cross-attention mechanism calculates the similarity or correlation between the two features and then dynamically weights the features according to this similarity. For example, if the second intermediate feature in a certain area is highly correlated with the face reference image feature, the features in that area will be given higher weights.

[0065] Furthermore, before obtaining the comprehensive face identity feature, the face reference image feature and the second intermediate feature processed by the third cross-attention mechanism are determined as the initial comprehensive face identity feature, and the initial comprehensive face identity feature is input into the query transformer Q-Former. Among them, Q-Former can reprocess and refine the initial comprehensive face identity feature to obtain a more finely tuned and targeted comprehensive face identity feature, which plays a key role in controlling the identity and details of the portrait in the subsequent diffusion model.

[0066] According to the technical solution provided by the embodiments of the present disclosure, a comprehensive and accurate comprehensive face identity feature can be generated, which retains the identity and structural information of the face and also integrates the visual details and style of the reference image, and can significantly improve the generation of the target image, making it have broad application prospects in the fields of portrait synthesis, virtual character creation, image restoration, etc.

[0067] In some embodiments, extracting face identity features, face key point features, and face reference image features from the portrait reference image includes:

[0068] Inputting the portrait reference image into a preset portrait identity feature extraction model to obtain the face identity feature corresponding to the portrait reference image;

[0069] Input the portrait reference image into a preset portrait key point extraction model to obtain the face key point features corresponding to the portrait reference image;

[0070] Input the portrait reference image into a preset portrait image feature extraction model to obtain the face reference image features corresponding to the portrait reference image.

[0071] Specifically, input the portrait reference image into a preset portrait identity feature extraction model. The preset portrait identity feature extraction model is usually a neural network trained through deep learning, which is specifically used to capture the identity information of the face.

[0072] Inside the model, the portrait reference image first undergoes a series of convolutional layers and activation functions. The low-level features are further abstracted and fused to form high-level feature representations, and a compact and discriminative face identity feature vector is output. The face identity feature vector can uniquely represent the identity of the person in the portrait reference image and can remain relatively stable even under different lighting, angles, or expression conditions.

[0073] Among them, the preset portrait identity feature extraction model can be the face recognition ArcFace model. This model can extract representative face identity features, which can be used to distinguish the faces of different individuals and lay a foundation for subsequent control to generate the consistency between the target image and the portrait reference image.

[0074] In addition, input the portrait reference image into a preset portrait key point extraction model to extract the face key point features. The preset portrait key point extraction model is usually a trained deep neural network that can automatically detect and locate the key points in the face.

[0075] Inside the model, the input image first passes through the feature extraction layers, which can identify the local features in the image. Subsequently, the model predicts the position of each key point through regression or classification, and outputs a set of key point coordinates that define the structural layout of the face. Through the key point coordinates, the face key point features can be further extracted to reflect the geometric structure and shape information of the face in the portrait reference image.

[0076] Among them, the preset portrait key point extraction model can be the Dlib model library of machine learning algorithms. The Dlib model library is trained with a large amount of face data and can accurately locate the key point coordinates of each key part of the face, output the key points and key point features of the face. Among them, it is required that the portrait reference image clearly identifies the key parts such as eyes, nose, mouth, etc., which is helpful for more refined portrait control and feature fusion operations based on the key point features.

[0077] In addition, input the portrait reference image into a preset portrait image feature extraction model to extract the face reference image features. The preset portrait image feature extraction model can extract the high-level semantic features of the image.

[0078] Inside the model, the input image is processed through multiple convolutional layers and pooling layers to gradually extract texture, color, and shape information from the image. Subsequently, through the fully connected layer, the above features are integrated into the feature vector of the face reference map, which can comprehensively represent the visual information of the portrait reference map, including the style, texture, and color distribution of the image, etc.

[0079] Among them, the preset portrait image feature extraction model can be the image encoder Vision Transformer (VIT). This encoder can effectively capture various visual features in the portrait reference map, convert them into a vector representation form suitable for subsequent calculations, thereby obtaining the face reference map features, and providing basic image feature information for subsequent image-related processing.

[0080] According to the technical solution provided by the embodiments of the present disclosure, three key features can be extracted from the portrait reference map: face identity features, face key point features, and face reference map features, which describe face information from three aspects of identity, structure, and visual details respectively, provide comprehensive and rich inputs for face image processing and generation tasks, can significantly improve the quality and diversity of target image generation, and at the same time provide strong support for tasks such as face recognition, image editing, and virtual character creation.

[0081] In some embodiments, based on the text processing features and the third image processing features, through the shared feed-forward network in the preset diffusion model, the target image feature vector is obtained, including:

[0082] The shared feed-forward network processes the text processing features and the third image processing features to obtain the feed-forward image features and feed-forward text features corresponding to the current text processing features and the third image processing features;

[0083] Perform iterative processing on the feed-forward image features and the feed-forward text features to obtain the target image feature vector and the target text feature vector, and discard the target text feature vector, retaining the target image feature vector.

[0084] Specifically, the text processing features and the third image processing features are input into the shared feed-forward network in the preset diffusion model. Among them, the shared feed-forward network is a multi-layer neural network structure, and its core function is to fuse and transform features from different sources, usually including multiple fully connected layers (or convolutional layers), activation functions (such as ReLU or GELU), and normalization layers (such as LayerNorm), and can perform complex non-linear transformations on the input features.

[0085] Through a series of inter-layer interactions and non-linear transformations of the shared feed-forward network, the text processing features and the third image features are deeply fused, preserving the semantic information of the text features and combining the visual details of the image features at the same time, so that the generated feed-forward image features and feed-forward text features corresponding to the current text processing features and the third image processing features can reflect the requirements of the text description and the portrait reference image simultaneously.

[0086] Furthermore, the essence of the feed-forward image features and the feed-forward text features is the same. Since different processing methods need to be performed on the feed-forward image features and the feed-forward text features in the next iteration process, they are named separately.

[0087] In addition, after obtaining the feed-forward image features and the feed-forward text features, iterative processing is performed on them. The purpose of the iterative processing is to make the feature vectors more accurately reflect the requirements of the target image through multiple optimizations and adjustments.

[0088] In each iteration process, the feed-forward image features obtained in the previous iteration round are determined as the first image features, and the feed-forward text features obtained in the previous iteration round are determined as the text features. The steps of repeating the input of the text features and the first image features into the preset diffusion model to process the first image features and the text features through the shared self-attention mechanism to obtain the text processing features and the first image processing features are performed. The first image features are obtained by adding noise to the face reference map features; the second image features after linear processing of the face reference map features are acquired, and the second image features are input into the preset diffusion model to process the first image processing features and the second image features through the first cross-attention mechanism to obtain the second image processing features; based on the face identity features, the face key point features and the face reference map features, the face identity comprehensive features are obtained, and the face identity comprehensive features are input into the preset diffusion model to process the face identity comprehensive features and the second image processing features through the second cross-attention mechanism to obtain the third image processing features; the steps of processing the text processing features and the third image processing features through the shared feed-forward network to obtain the feed-forward image features and the feed-forward text features corresponding to the current text processing features and the third image processing features are performed until the iterative processing is completed. The feed-forward image features obtained in the last iterative processing are determined as the target image feature vector, and the feed-forward text features obtained in the last iterative processing are determined as the target text feature vector, and the target text feature vector is discarded, and the target image feature vector is retained.

[0089] Among them, each iterative processing will fine-tune the feature vectors to make them better fuse the image features and the text features and approach the ideal feature representation of the target image.

[0090] Among them, the target text feature vector plays an auxiliary role in the iterative process but is not directly used when generating the target image. Therefore, after the iterative processing is completed, the target text feature vector is discarded, and the target image feature vector is retained. The target image feature vector is a high-dimensional feature representation that synthesizes all the key information of the text description and the image reference diagram, and can accurately guide the generation of the target image, so that the generated target image can meet the requirements of both the text description and the portrait reference diagram.

[0091] According to the technical solution provided by the embodiments of the present disclosure, the combination of the shared feed-forward network and the iterative processing optimizes both the text processing features and the third image processing feature vector in terms of semantics and vision, obtaining the target image feature vector. The target image feature vector synthesizes the key information of the text description and the portrait reference diagram, so that an accurate target image can be generated based on the target image feature vector, significantly improving the quality and diversity of the target image.

[0092] In some embodiments, iterative processing is performed on the feed-forward image features and the feed-forward text features to obtain a target image feature vector and a target text feature vector, including:

[0093] Taking the current feed-forward image feature as the first image feature in the next iterative process and taking the current feed-forward text feature as the text feature in the next iterative process;

[0094] Performing the step of inputting the text feature and the first image feature into a preset diffusion model until the number of iterations reaches the target number of iterations.

[0095] Specifically, the feed-forward image feature and the feed-forward text feature obtained in the previous iteration are respectively used as the first image feature and the text feature, that is, the current feed-forward image feature is used as the first image feature in the next iterative process, and the current feed-forward text feature is used as the text feature in the next iterative process. These two features will be used as inputs and enter the preset diffusion model for further processing.

[0096] The process of further processing by the preset diffusion model is as follows: repeatedly execute the step of inputting the text feature and the first image feature into the preset diffusion model to process the first image feature and the text feature through the shared self-attention mechanism, and obtain the text processing feature and the first image processing feature. The first image feature is obtained by adding noise to the face reference map feature; obtain the second image feature after linear processing of the face reference map feature, and input the second image feature into the preset diffusion model to process the first image processing feature and the second image feature through the first cross-attention mechanism to obtain the second image processing feature; based on the face identity feature, the face key point feature, and the face reference map feature, obtain the comprehensive face identity feature, and input the comprehensive face identity feature into the preset diffusion model to process the comprehensive face identity feature and the second image processing feature through the second cross-attention mechanism to obtain the third image processing feature; process the text processing feature and the third image processing feature through the shared feed-forward network to obtain the feed-forward image feature and the feed-forward text feature corresponding to the current text processing feature and the third image processing feature, until the iterative processing is completed, and determine the feed-forward image feature obtained in the last iterative processing as the target image feature vector, and determine the feed-forward text feature obtained in the last iterative processing as the target text feature vector.

[0097] That is, the shared self-attention mechanism, the first cross-attention mechanism, the second cross-attention mechanism, and the shared feed-forward network can be determined as a sub-module in the preset diffusion model. Each sub-module represents an iterative process. There are multiple sub-modules in the preset diffusion model, and the number of sub-modules is the same as the number of iterative processes.

[0098] After each iteration is completed, check whether the current iteration count reaches the target iteration count. The target iteration count is a pre-set parameter used to control the depth and accuracy of the iterative process. If the current iteration count does not reach the target iteration count, continue to execute the iterative process. If the current iteration count reaches the target iteration count, stop the iterative process.

[0099] According to the technical solution provided by the embodiments of the present disclosure, multiple iterative processes can enable the obtained target image feature vector to integrate all key information of the text description and the portrait reference map, so as to accurately generate the target image. The introduction of the iterative process optimizes the feature vector in terms of semantics and vision, significantly improving the quality and diversity of the generated image.

[0100] In some embodiments, obtaining the second image feature after linear processing of the face reference map feature includes:

[0101] Obtain the first feature dimension corresponding to the first image processing feature;

[0102] Perform linear processing on the features of the face reference image to make the feature dimension corresponding to the face reference image features the same as the first feature dimension, and determine the face reference image features after adjusting the feature dimension as the second image features.

[0103] Specifically, obtain the first feature dimension corresponding to the first portrait processing feature. Here, the feature dimension refers to the length or number of dimensions of the feature vector, which determines the amount of information that the feature vector can represent. The first image processing feature is the image feature processed through the previous steps, and its feature dimension can be used for subsequent feature alignment and fusion.

[0104] In addition, perform linear processing on the face reference image features to make their feature dimensions the same as those of the first image processing feature. Here, linear processing is a common feature transformation method, that is, adjusting the dimension and distribution of features through matrix operations.

[0105] Among them, the consistency of the feature dimension is the basis for realizing feature fusion, and linear processing can be performed through a linear layer.

[0106] Furthermore, after linear processing, the feature dimension of the face reference image features has become the same as the first feature dimension. Determine the face reference image features after adjusting the feature dimension as the second image features, which can be used for subsequent steps such as feature fusion.

[0107] According to the technical solution provided by the embodiments of the present disclosure, it is possible to adjust the dimension of the face reference image features to be consistent with the dimension of the first image processing feature, ensuring the consistency of the feature dimension, while retaining the important information of the face reference image features, and improving the performance and effect of the entire system.

[0108] In some embodiments, generating a target image based on the target image feature vector includes:

[0109] Input the target image feature vector into a preset feature decoder, so that the feature decoder generates a target image based on the portrait reference image and conforming to the character description information based on the target image feature vector.

[0110] Specifically, obtain the target image feature vector. The target image feature vector is the result of multiple optimizations and fusions, containing rich semantic information and visual details. Input the target image feature vector into the preset feature decoder. Here, the preset feature decoder is a trained neural network, usually composed of multiple decoding layers, which can gradually restore the feature vector to the image content.

[0111] The preset feature decoder maps the target image feature vector to an intermediate feature space through a fully connected layer or a convolutional layer, converting the high-dimensional target image feature vector into a feature representation suitable for image generation. For example, the feature decoder may map the feature vector to a low-resolution feature map, where each pixel contains rich semantic information.

[0112] The preset feature decoder gradually increases the resolution of the feature map through a series of upsampling operations. Among them, the upsampling operation usually includes a transposed convolution layer or an interpolation method (such as bilinear interpolation). Each upsampling operation doubles the resolution of the feature map. At the same time, the content of the feature map is further refined through a convolutional layer, gradually increasing the detail information of the image and making it gradually approach the final image resolution.

[0113] After each upsampling operation, the preset feature decoder refines the feature map through a convolutional layer. These convolutional layers can capture local features and further optimize the details of the image. For example, convolutional layers can be used to smooth image edges, enhance texture details, or adjust color distribution.

[0114] After multiple upsamplings and feature refinements, the preset feature decoder generates a target image that not only conforms to the visual style of the portrait reference image but also meets the semantic requirements in the person description information. For example, if the text description information mentions "a smiling woman with long hair and glasses", the generated image will show a female image that meets these descriptions while retaining the visual style of the portrait reference image.

[0115] According to the technical solution provided by the embodiments of the present disclosure, the visual style of the portrait reference image is utilized, and the semantic requirements of the person description information are also combined, making the generated target image highly accurate and diverse.

[0116] Figure 2 It is a schematic flowchart of another method for generating a person image provided by the embodiments of the present disclosure, as Figure 2As shown in the figure, obtain the text information corresponding to the person description information, input it into the text encoder to obtain text features, obtain a portrait reference image, input the portrait reference image into a preset portrait identity feature extraction model to obtain the face identity features corresponding to the portrait reference image, and input the portrait reference image into a preset portrait key point extraction model to obtain the face key point features corresponding to the portrait reference image. At the same time, input the portrait reference image into a preset portrait image feature extraction model to obtain the face reference image features corresponding to the portrait reference image. Input the face identity features and the face key point features into a preset identity network model to obtain a first intermediate feature. Obtain the position box mask of the portrait in the portrait reference image, and perform mask processing on the first intermediate feature through the position box mask to obtain a second intermediate feature. Process the face reference image features and the second intermediate feature through a third cross-attention mechanism and process them again through Q-Former to obtain face identity comprehensive features; obtain a noise map, splice the noise map with the face reference image features to obtain a first image feature, input the text features and the first image feature into a preset diffusion model, and process the first image feature and the text features through the shared self-attention mechanism in the preset diffusion model to obtain text processing features and first image processing features. Perform linear processing on the face reference image features to obtain a second image feature, and process the first image processing features and the second image feature through the first cross-attention mechanism in the preset diffusion model to obtain second image processing features; input the face identity comprehensive features into the preset diffusion model, and process the face identity comprehensive features and the second image processing features through the second cross-attention mechanism to obtain third image processing features; according to the text processing features and the third image processing features, process the text processing features and the third image processing features through the shared feed-forward network in the preset diffusion model to obtain the feed-forward image features and feed-forward text features corresponding to the current text processing features and the third image processing features; perform iterative processing on the feed-forward image features and the feed-forward text features to obtain a target image feature vector and a target text feature vector, discard the target text feature vector, retain the target image feature vector, and generate a target image based on the target image feature vector.

[0117] Any combination of the above all optional technical solutions can form an optional embodiment of the present disclosure, which will not be elaborated here one by one.

[0118] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For the details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the method of the present disclosure.

[0119] Figure 3 It is a schematic diagram of a device for generating a person image provided by an embodiment of the present disclosure. As Figure 3As shown in the figure, the generation device for the portrait image includes: an extraction module 301, a first processing module 302, a second processing module 303, a third processing module 304, and a generation module 305, where:

[0120] The extraction module 301 is configured to obtain the text features corresponding to the person description information, and extract the face identity features, face key point features, and face reference image features from the portrait reference image;

[0121] The first processing module 302 is configured to input the text features and the first image features into a preset diffusion model, so as to process the first image features and the text features through a shared self-attention mechanism to obtain text processing features and first image processing features, and the first image features are obtained by adding noise to the face reference image features;

[0122] The second processing module 303 is configured to obtain the second image features after linear processing of the face reference image features, and input the second image features into a preset diffusion model, so as to process the first image processing features and the second image features through a first cross-attention mechanism to obtain second image processing features;

[0123] The third processing module 304 is configured to obtain a face identity comprehensive feature based on the face identity features, face key point features, and face reference image features, and input the face identity comprehensive feature into a preset diffusion model, so as to process the face identity comprehensive feature and the second image processing features through a second cross-attention mechanism to obtain third image processing features;

[0124] The generation module 305 is configured to obtain a target image feature vector based on the text processing features and the third image processing features through a shared feed-forward network in the preset diffusion model, and generate a target image based on the target image feature vector.

[0125] In some embodiments, the third processing module 304 is configured to:

[0126] Input the face identity features and the face key point features into a preset identity network model to obtain a first intermediate feature;

[0127] Obtain the position box mask of the portrait in the portrait reference image, and perform mask processing on the first intermediate feature through the position box mask to obtain a second intermediate feature;

[0128] Process the face reference image features and the second intermediate feature through a third cross-attention mechanism to obtain a face identity comprehensive feature.

[0129] In some embodiments, the extraction module 301 is configured to:

[0130] Input the portrait reference image into a preset portrait identity feature extraction model to obtain the face identity features corresponding to the portrait reference image;

[0131] Input the portrait reference image into a preset portrait key point extraction model to obtain the face key point features corresponding to the portrait reference image;

[0132] Input the portrait reference image into a preset portrait image feature extraction model to obtain the face reference image features corresponding to the portrait reference image.

[0133] In some embodiments, the generation module 305 is configured to:

[0134] Process the text processing features and the third image processing features through a shared feed-forward network to obtain the feed-forward image features and feed-forward text features corresponding to the current text processing features and the third image processing features;

[0135] Perform iterative processing on the feed-forward image features and the feed-forward text features to obtain a target image feature vector and a target text feature vector, discard the target text feature vector, and retain the target image feature vector.

[0136] In some embodiments, the generation module 305 is configured to:

[0137] Use the current feed-forward image feature as the first image feature in the next iterative processing, and use the current feed-forward text feature as the text feature in the next iterative processing;

[0138] Execute the step of inputting the text feature and the first image feature into a preset diffusion model until the number of iterations reaches the target number of iterations.

[0139] In some embodiments, the second processing module 303 is configured to:

[0140] Obtain the first feature dimension corresponding to the first image processing feature;

[0141] Linearly process the face reference image features so that the feature dimension corresponding to the face reference image features is the same as the first feature dimension, and determine the face reference image features after adjusting the feature dimension as the second image feature.

[0142] In some embodiments, the generation module 305 is configured to:

[0143] Input the target image feature vector into a preset feature decoder so that the feature decoder generates a target image based on the portrait reference image and conforming to the character description information based on the target image feature vector.

[0144] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.

[0145] Figure 4 is a schematic diagram of the electronic device 4 provided by the embodiments of the present disclosure. As Figure 4 shown, the electronic device 4 in this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.

[0146] The electronic device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the electronic device 4, do not constitute a limitation on the electronic device 4, and may include more or fewer components than shown, or different components.

[0147] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0148] The memory 402 may be an internal storage unit of the electronic device 4, for example, the hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. The memory 402 may also include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0149] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0150] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium (such as a computer-readable storage medium). Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0151] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A method for generating a human image, characterized in that, Including: Obtain the text features corresponding to the person description information, and extract the face identity features, face key point features, and face reference image features from the portrait reference image; Input the text features and the first image features into a preset diffusion model to process the first image features and the text features through a shared self-attention mechanism, and obtain text processing features and first image processing features. The first image features are obtained by adding noise to the face reference image features; Obtain the second image features after linear processing of the face reference image features, and input the second image features into the preset diffusion model to process the first image processing features and the second image features through a first cross-attention mechanism, and obtain second image processing features; Based on the face identity features, face key point features, and face reference image features, obtain a face identity comprehensive feature, and input the face identity comprehensive feature into the preset diffusion model to process the face identity comprehensive feature and the second image processing features through a second cross-attention mechanism, and obtain third image processing features; Based on the text processing features and the third image processing features, through the shared feed-forward network in the preset diffusion model, obtain a target image feature vector, and generate a target image based on the target image feature vector.

2. The method according to claim 1, characterized in that, Based on the face identity features, face key point features, and face reference image features, obtaining a face identity comprehensive feature includes: Input the face identity features and the face key point features into a preset identity network model to obtain a first intermediate feature; Obtain the position box mask of the portrait in the portrait reference image, and perform mask processing on the first intermediate feature through the position box mask to obtain a second intermediate feature; Process the face reference image features and the second intermediate feature through a third cross-attention mechanism to obtain a face identity comprehensive feature.

3. The method according to claim 1, wherein Extracting the face identity features, face key point features, and face reference image features from the portrait reference image includes: Input the portrait reference image into a preset portrait identity feature extraction model to obtain the face identity features corresponding to the portrait reference image; Input the portrait reference image into a preset portrait key point extraction model to obtain the face key point features corresponding to the portrait reference image; Input the portrait reference image into a preset portrait image feature extraction model to obtain the face reference image features corresponding to the portrait reference image.

4. The method according to claim 1, characterized in that, Based on the text processing features and the third image processing features, obtaining a target image feature vector through the shared feed-forward network in the preset diffusion model includes: Process the text processing features and the third image processing features through the shared feed-forward network to obtain the feed-forward image features and feed-forward text features corresponding to the current text processing features and the third image processing features; Perform iterative processing on the feed-forward image features and the feed-forward text features to obtain a target image feature vector and a target text feature vector, and discard the target text feature vector and retain the target image feature vector.

5. The method according to claim 4, wherein Performing iterative processing on the said feedforward image features and the said feedforward text features to obtain a target image feature vector and a target text feature vector, including: Taking the current said feedforward image features as the first image features in the next iterative processing, and taking the current said feedforward text features as the text features in the next iterative processing; Performing the step of inputting the said text features and the first image features into a preset diffusion model until the number of iterations reaches the target number of iterations.

6. The method according to claim 1, characterized in that, Obtaining second image features after linear processing of the said face reference map features, including: Obtaining a first feature dimension corresponding to the first image processing features; Performing linear processing on the said face reference map features to make the feature dimension corresponding to the said face reference map features the same as the first feature dimension, and determining the face reference map features after adjusting the feature dimension as the second image features.

7. The method according to claim 1, wherein Generating a target image based on the said target image feature vector, including: Inputting the said target image feature vector into a preset feature decoder, so that the feature decoder generates a target image based on the portrait reference map and conforming to the said character description information based on the said target image feature vector.

8. A generating device for a human image, characterized in that, Including: An extraction module, configured to obtain text features corresponding to character description information, and extract face identity features, face key point features and face reference map features from a portrait reference map; A first processing module, configured to input the said text features and the first image features into a preset diffusion model to process the first image features and the said text features through a shared self-attention mechanism to obtain text processing features and first image processing features, where the first image features are obtained by adding noise to the face reference map features; A second processing module, configured to obtain second image features after linear processing of the said face reference map features, and input the second image features into the preset diffusion model to process the first image processing features and the second image features through a first cross-attention mechanism to obtain second image processing features; A third processing module, configured to obtain a comprehensive face identity feature based on the said face identity features, face key point features and face reference map features, and input the comprehensive face identity feature into the preset diffusion model to process the comprehensive face identity feature and the second image processing features through a second cross-attention mechanism to obtain third image processing features; A generation module, configured to obtain a target image feature vector based on the said text processing features and the third image processing features through a shared feedforward network in the preset diffusion model, and generate a target image based on the said target image feature vector.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the said computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium storing a computer program, characterized in that, When the said computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Multi-modal data fusion method and system for large model training

    CN121479640A

  • A multimodal data fusion method and system for training large models

    CN121479640B