Image generation method and device, readable storage medium and program product

By extracting and fusing the identity features and image features of reference face images, the problem of poor identity consistency of faces in the prior art is solved, and the consistency requirement for model images in e-commerce and other scenarios is achieved.

CN120125713APending Publication Date: 2025-06-10XIAMEN MEITUZHIJIA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214402.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, when generating images, the consistency of facial identities is poor, making it difficult to meet the high consistency demand for model images in e-commerce and other scenarios.

Method used

By acquiring the reference face image and the original character image, extracting the identity features of the reference face area and multiple image features at different feature levels, performing feature fusion, and generating the target character image of the target face area.

Benefits of technology

The consistency of face identities in the generated image is improved. The generated target face area focuses on both detailed texture information and improves authenticity, as well as identity characteristics, ensuring identity consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125713A_ABST
    Figure CN120125713A_ABST
Patent Text Reader

Abstract

The invention relates to an image generation method and device, a readable storage medium and a program product. The method comprises the following steps: acquiring a reference face image and an original figure image; the reference face image comprises a reference face area, and the original figure image comprises an original face area; extracting identity features of the reference face region based on the reference face image, and extracting a plurality of image features of different feature levels based on the reference face image; fusing the plurality of image features with the identity feature to obtain a plurality of fused features corresponding to the plurality of image features; performing image generation processing based on the plurality of fusion features and the original figure image, so as to edit the original face region according to the reference face region, and generating a target figure image including a target face region obtained by editing; wherein the number of the fusion features is at least two. By adopting the method, the consistency of human face identities in the generated image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image generation, and particularly to an image generation method, apparatus, readable storage medium, and program product. Background Art

[0002] With the development of artificial intelligence technology, through artificial intelligence technology, it is possible to edit the face in the original human image to generate a target human image that meets specific requirements. For example, in the e-commerce scenario, the traditional model shooting scenario involves human resources such as models, photographers, and makeup artists, and is also restricted by shooting time and venue, resulting in a problem of high resource consumption. To solve the above problems, there is a need for AI model image generation. In the prior art, an image generation model is usually used to perform image generation on the original human image based on text prompts or picture prompts.

[0003] However, using the above prior art, the consistency of the human face identity in the generated image is poor. Summary of the Invention

[0004] Based on this, the present application provides an image generation method, apparatus, readable storage medium, and program product, which can improve the consistency of the human face identity in the generated image.

[0005] On the one hand, the present application provides an image generation method, including:

[0006] Obtain a reference face image and an original human image; the reference face image includes a reference face area, and the original human image includes an original face area;

[0007] Extract the identity feature of the reference face area based on the reference face image, and extract multiple image features at different feature levels based on the reference face image;

[0008] Fuse the multiple image features with the identity feature respectively to obtain multiple fusion features corresponding to the multiple image features respectively;

[0009] Perform image generation processing based on the multiple fusion features and the original human image to edit the original face area according to the reference face area, and generate a target human image including the edited target face area;

[0010] Wherein the number of the fusion features is at least two.

[0011] In one embodiment, the fusing the multiple image features with the identity feature respectively to obtain multiple fusion features corresponding to the multiple image features respectively includes:

[0012] For each of the multiple image features, perform attention feature fusion based on the image feature and the identity feature to obtain a first fusion feature;

[0013] Perform a non-linear feature transformation based on the first fusion feature to obtain a non-linear feature;

[0014] Perform feature fusion according to the first fusion feature and the non-linear feature to obtain a fusion feature corresponding to the image feature.

[0015] In one embodiment, the performing attention feature fusion based on the image feature and the identity feature to obtain a first fusion feature includes:

[0016] Respectively use a trained query weight matrix and a value weight matrix to perform a linear transformation on the image feature to obtain a query vector and a value vector;

[0017] Use a trained key weight matrix to perform a linear transformation on the identity feature to obtain a key vector;

[0018] Perform attention feature fusion according to the query vector, the value vector and the key vector to obtain a first fusion feature.

[0019] In one embodiment, the performing a non-linear feature transformation based on the first fusion feature to obtain a non-linear feature includes:

[0020] Perform a feature addition operation on the first fusion feature, the query vector, the value vector and the key vector to obtain a second fusion feature;

[0021] Perform a non-linear feature transformation on the second fusion feature through a trained feed-forward network to obtain a non-linear feature;

[0022] The performing feature fusion according to the first fusion feature and the non-linear feature to obtain a fusion feature corresponding to the image feature includes:

[0023] Perform a feature addition operation on the second fusion feature and the non-linear feature to obtain a fusion feature corresponding to the image feature.

[0024] In one embodiment, the extracting multiple image features of different feature levels based on the reference face image includes:

[0025] Segment the reference face image into multiple image patches, and map the multiple image patches into multiple patch vectors respectively;

[0026] Based on the multiple patch vectors, feature extraction is sequentially performed through multiple encoding layers stacked in the trained image encoder to obtain intermediate features extracted by each of the multiple encoding layers;

[0027] Based on at least two of the intermediate features extracted by each of the multiple encoding layers, multiple image features at different feature levels are determined.

[0028] In one embodiment, the based on the multiple patch vectors, feature extraction is sequentially performed through multiple encoding layers stacked in the trained image encoder to obtain intermediate features extracted by each of the multiple encoding layers, including:

[0029] Based on the pre-trained class learning vector and the multiple patch vectors, feature extraction is sequentially performed through multiple encoding layers stacked in the trained image encoder. When each encoding layer performs feature extraction, feature interaction is performed on the class learning vector and the multiple patch vectors input to the encoding layer to obtain an updated class learning vector and updated multiple patch vectors, and the intermediate features extracted by the encoding layer formed based on the updated class learning vector and the updated multiple patch vectors.

[0030] In one embodiment, the based on at least two of the intermediate features extracted by each of the multiple encoding layers, determining multiple image features at different feature levels, including:

[0031] The intermediate features extracted by an encoding layer that is one encoding layer before the last encoding layer in the stacking order among the multiple encoding layers are determined as the image features at the first feature level;

[0032] The intermediate features extracted by the last encoding layer are determined as the image features at the second feature level;

[0033] The class learning vector in the image features at the second feature level is extracted to obtain the image features at the third feature level.

[0034] On the one hand, the present application also provides an image generation device, including:

[0035] An acquisition module, configured to acquire a reference face image and an original person image; the reference face image includes a reference face area, and the original person image includes an original face area;

[0036] A feature extraction module, configured to extract identity features of the reference face area based on the reference face image, and extract multiple image features at different feature levels based on the reference face image;

[0037] A feature fusion module, configured to fuse the multiple image features with the identity features respectively to obtain multiple fusion features respectively corresponding to the multiple image features;

[0038] An image generation module, configured to perform image generation processing based on the multiple fusion features and the original person image, so as to edit the original face region according to the reference face region, and generate a target person image including the edited target face region; wherein the number of the fusion features is at least two.

[0039] On the one hand, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0040] Obtain a reference face image and an original person image; the reference face image includes a reference face region, and the original person image includes an original face region;

[0041] Extract the identity feature of the reference face region based on the reference face image, and extract multiple image features of different feature levels based on the reference face image;

[0042] Fuse the multiple image features with the identity feature respectively to obtain multiple fusion features respectively corresponding to the multiple image features;

[0043] Perform image generation processing based on the multiple fusion features and the original person image, so as to edit the original face region according to the reference face region, and generate a target person image including the edited target face region;

[0044] Wherein the number of the fusion features is at least two.

[0045] On the one hand, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0046] Obtain a reference face image and an original person image; the reference face image includes a reference face region, and the original person image includes an original face region;

[0047] Extract the identity feature of the reference face region based on the reference face image, and extract multiple image features of different feature levels based on the reference face image;

[0048] Fuse the multiple image features with the identity feature respectively to obtain multiple fusion features respectively corresponding to the multiple image features;

[0049] Perform image generation processing based on the multiple fusion features and the original person image, so as to edit the original face region according to the reference face region, and generate a target person image including the edited target face region;

[0050] Wherein the number of the fusion features is at least two.

[0051] The above image generation method, device, readable storage medium, and program product extract identity features of a reference face region based on a reference face image, and extract multiple image features of different feature levels based on the reference face image. The image features of a higher feature level can focus on the semantic information of the reference face image, while the image features of a lower feature level can focus on detailed texture information such as makeup and skin texture. Then, the multiple image features are respectively fused with the identity features. In this way, the multiple fusion features can take into account the identity features, semantic information, and detailed texture information, and the number of the fusion features is at least two. Therefore, based on the multiple fusion features and the original person image, image generation processing can be performed to achieve editing of the original face region according to the reference face region. Moreover, for the target face region in the generated target person image, both the detailed texture information is concerned to improve the authenticity of the target face region, and the identity features of the reference face region are concerned. When generating multiple target person images or generating the target person image multiple times, the identity consistency of the faces in the generated images is high. Description of the Drawings

[0052] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a schematic flowchart of an image generation method in an embodiment;

[0054] Figure 2 It is a schematic diagram of a fusion network framework in an embodiment;

[0055] Figure 3 It is a schematic diagram of an image generation architecture in an embodiment;

[0056] Figure 4 It is a structural block diagram of an image generation device in an embodiment;

[0057] Figure 5 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0058] In order to make the objectives, technical solutions, and beneficial effects of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] In one embodiment, as Figure 1 shown, an image generation method is provided. In this embodiment, the method is exemplified by being applied to a computer device. The computer device can be a terminal or a server. It can be understood that the method can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. Among them, the terminal can be a personal computer, a laptop, a smartphone, a tablet computer, or others. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps:

[0060] Step 102, obtain a reference face image and an original person image; the reference face image includes a reference face area, and the original person image includes an original face area.

[0061] Among them, the reference face image is an image of a face used as a reference. The reference face image can be a real face image taken of a human. The reference face image can also be a virtual face image generated by an artificial intelligence model. The reference face area is the area in the reference face image representing the face.

[0062] The original person image is a person image whose face area needs to be edited. The original face area is the area in the original person image representing the face. The original person image can be a half-body image or a full-body image of a person. Specifically, the original person image can be a personal photo, a portrait photo, or a model fitting photo. In an e-commerce scenario, there is a need to obtain different model fitting photos of different models wearing the same piece of clothing. Then, a fitting photo of one model wearing the clothing can be taken to obtain the original person image, and then the image generation method provided in this application can be used to generate other model fitting photos of different models wearing the same piece of clothing, so as to reduce the resource consumption in this e-commerce scenario.

[0063] Exemplarily, the computer device can receive an image generation request and obtain the reference face image and the original face image from the image generation request. Among them, the image generation request can be generated based on an image generation interface. For example, the reference face image can be uploaded or selected in the image generation interface, and the original person image can be uploaded or selected in the image generation interface. After triggering the generation trigger control in the image generation interface, an image generation request can be generated based on the reference face image and the original person image. The image generation request can also indicate the number of target person images to be generated.

[0064] Step 104, extract the identity features of the reference face area based on the reference face image, and extract multiple image features at different feature levels based on the reference face image.

[0065] Among them, the identity feature is a feature used to distinguish different faces. It can be understood that different faces may have different identity features. The identity feature can also be referred to as the ID (Identification) feature. The image feature is a feature used to identify and distinguish images. The image feature may include features such as semantics, texture, color, brightness, and shape. The feature hierarchy is the hierarchy of features respectively output by performing feature extraction at different levels on the reference face image. For example, for the reference face image, feature extraction can be sequentially performed through multiple stacked feature extraction layers to perform feature extraction at different levels. Each feature extraction layer can output the features it extracts. The hierarchy of the features extracted by the feature extraction layer earlier in the stacking order is lower than the hierarchy of the features extracted by the feature extraction layer later in the stacking order.

[0066] Exemplarily, the computer device can input the reference face image into the trained face recognition model and obtain the identity feature of the reference face region output by the face recognition model. Among them, the face recognition model can be the Arcface model (a face recognition model proposed by a research team at the Hong Kong University of Science and Technology), the FaceNet model (a face recognition model proposed by a research team at Google), or others. The identity feature extracted by the face recognition model can be the embedding feature (Embeding).

[0067] In one embodiment, the computer device can extract multiple image features at different feature levels based on the reference face image through the trained image encoder. Among them, the image encoder is used to extract image features from image data. The image encoder can be ViT (Vision Transformer, a model that applies the Transformer architecture to image processing tasks), ResNet (Residual Network), or others. The image encoder can be the encoder for processing images in a multimodal model. The multimodal model can be the CLIP model (Contrastive Language-ImagePre-Training, a multimodal pre-trained neural network model), the BLIP model (Bootstrapping Language-Image Pre-training, a multimodal vision-text large language model), or others.

[0068] In one embodiment, the computer device may sequentially perform feature extraction through multiple stacked encoding layers in a trained image encoder based on a reference face image to obtain intermediate features extracted by each of the multiple encoding layers; and determine multiple image features at different feature levels according to at least two of the intermediate features extracted by each of the multiple encoding layers. Herein, stacking means that the multiple encoding layers are sequentially connected. When sequentially performing feature extraction through the multiple stacked encoding layers, the features extracted by the prior encoding layer among two adjacent encoding layers in the stacking order are input to the subsequent encoding layer for feature extraction.

[0069] Step 106: Fuse the multiple image features with the identity features respectively to obtain multiple fused features corresponding to the multiple image features respectively.

[0070] Exemplarily, the computer device may, for each of the multiple image features, fuse the targeted image feature with the identity feature to obtain a fused feature corresponding to the targeted image feature.

[0071] In one embodiment, the computer device may, for each of the multiple image features, perform attention feature fusion on the targeted image feature and the identity feature to obtain a fused feature corresponding to the targeted image feature. Herein, attention feature fusion is a process of performing feature fusion using an attention mechanism.

[0072] Step 108: Perform image generation processing based on the multiple fused features and the original person image to edit the original face region according to the reference face region, and generate a target person image including the edited target face region; wherein the number of fused features is at least two.

[0073] Herein, the number of fused features is at least two. For example, the number of fused features may be two, three, four, or other numbers more than two. The number of fused features may be the same as the number of image features. For example, when the number of image features is three, the number of fused features may also be three. The target face region is the region representing the face in the target person image.

[0074] Exemplarily, the computer device may determine input features based on the original person image, and perform image generation processing based on the multiple fused features and the input features through a trained image generation model to edit the original face region according to the reference face region, and generate a target person image including the edited target face region.

[0075] Among them, the image generation model may include a cross-attention layer. Multiple fused features are input into the cross-attention layer, and the cross-attention layer can guide the image generation process based on the cross-attention mechanism and multiple fused features. The image generation model may be an image generation model based on the DiT (Diffusion Transformer, which combines a diffusion model and a Transformer) architecture, such as Diffusion Transformer, Stable Diffusion 3 (the third-generation image generation model released by Stability AI), or others. The image generation model may also be an image generation model based on the U-Net architecture, such as Stable Diffusion (the first-generation image generation model released by Stability AI), Latent Diffusion Models (Latent Diffusion Models, abbreviated as LDM), or others.

[0076] In one embodiment, the number of original person images may be multiple, and each original person image includes an original face region. In this embodiment, the computer device can perform image generation processing based on multiple fused features and multiple original person images to edit the original face regions included in the multiple original person images respectively according to the reference face region, and generate multiple target person images corresponding to the multiple original person images respectively. Each target person image includes a target face region obtained by editing the face image based on the corresponding original person image.

[0077] In one embodiment, the original person image includes an original clothing region. The computer device can perform image segmentation based on the original person image to determine a region mask representing the original clothing region, and determine input features according to the original person image and the region mask; through the trained image generation model, perform image generation based on multiple fused features and the input features to edit the original face region according to the reference face region and retain the content of the original clothing region, and generate a target person image including the edited target face region and the target clothing region formed by retaining the content.

[0078] Among them, the original clothing region is the region representing the clothing worn by the person in the original person image. Clothing includes, for example, clothes, hats, hair ornaments, etc. The region mask may have the same size as the original person image. In the region mask, each pixel in the first region corresponding to the position of the original clothing region may take a first pixel value, and each pixel in the second region outside the first region may take a second pixel value. The first pixel value may be 1, for example, and the second pixel value may be 0, for example.

[0079] The input features can be obtained by separately mapping the original person image and the region mask to the latent space and then splicing the latent space representations mapped by each. Since the input features contain information about the region mask, when the image generation model generates an image based on the input features, the position information of the original clothing region can be obtained, so that image generation in the original clothing region can be avoided to retain the content of the original clothing region.

[0080] In the above image generation method, the identity features of the reference face region are extracted based on the reference face image, and multiple image features at different feature levels are extracted based on the reference face image. The image features at a higher feature level can focus on the semantic information of the reference face image, while the image features at a lower feature level can focus on the detailed texture information such as makeup and skin texture. Then, the multiple image features are respectively fused with the identity features. In this way, the identity features, semantic information, and detailed texture information can be taken into account in the multiple fused features. The number of fused features is at least two. Therefore, based on the multiple fused features and the original person image for image generation processing, editing the original face region according to the reference face region can be realized. Moreover, for the target face region in the generated target person image, both the detailed texture information is concerned to improve the authenticity of the target face region, and the identity features of the reference face region are concerned. When generating multiple target person images or generating the target person image multiple times, the identity consistency of the faces in the generated images is high.

[0081] In an exemplary embodiment, step 106 may include: for each of the multiple image features, performing attention feature fusion based on the image feature and the identity feature to obtain a first fused feature; performing a non-linear feature transformation on the first fused feature to obtain a non-linear feature; and performing feature fusion based on the first fused feature and the non-linear feature to obtain a fused feature corresponding to the image feature.

[0082] Among them, attention feature fusion is a process of performing feature fusion using an attention mechanism. The attention mechanism can be a cross-attention mechanism or a multi-head cross-attention mechanism. The image feature and the identity feature can be mapped into query vectors, key vectors, and value vectors, and attention feature fusion is performed based on the query vectors, key vectors, and value vectors.

[0083] Non-linear feature transformation is a process of non-linearly transforming features based on activation functions. The activation function can be a Relu function (Linear rectification function), a Sigmoid function (S-shaped curve function), or others. The non-linear feature transformation can be specifically implemented using a Feed Forward Network. The Feed Forward Network can include an input layer, a hidden layer, and an output layer. When processed through the Feed Forward Network, the input data first enters the input layer, and then is passed to the hidden layer through weights and biases. The nodes in the hidden layer perform weighted summation on the input and perform non-linear transformation through the activation function. Finally, the output layer receives the signal processed by the hidden layer and generates the final output.

[0084] In this embodiment, by performing attention feature fusion on the image feature and the identity feature, the key features in the image feature and the identity feature can be more accurately focused on. Based on the first fusion feature, non-linear feature transformation is performed to further improve the expression ability of the features. Then, according to the first fusion feature and the non-linear feature, feature fusion is performed, combining features at different stages, which can improve the richness and expression ability of the features.

[0085] In one embodiment, the steps of obtaining non-linear features by performing non-linear feature transformation based on the first fusion feature may include: The computer device can perform feature fusion on the targeted image feature, identity feature, and the first fusion feature to obtain a second fusion feature, and perform non-linear feature transformation on the second fusion feature through a trained Feed Forward Network to obtain non-linear features.

[0086] In an exemplary embodiment, the steps of obtaining the first fusion feature by performing attention feature fusion on the image feature and the identity feature may include: respectively using a trained query weight matrix and value weight matrix to perform linear transformation on the image feature to obtain a query vector and a value vector; using a trained key weight matrix to perform linear transformation on the identity feature to obtain a key vector; performing attention feature fusion according to the query vector, value vector, and key vector to obtain the first fusion feature.

[0087] Among them, the query weight matrix, value weight matrix, and key weight matrix are learned through the training process, and these weight matrices can be updated through the backpropagation algorithm during the training process. The linear transformation can be matrix multiplication. For example, performing matrix multiplication on the image feature and the query weight matrix to obtain the query vector. When performing attention feature fusion, a Multi-head Cross-attention module can be used for processing, and the Multi-head Cross-attention module can include multiple cross-attention layers.

[0088] In this embodiment, the image features are linearly transformed into query vectors and value vectors, and the identity features are linearly transformed into key vectors. In this way, the transformation of the image features into query vectors can represent the information that needs to be focused on currently, the transformation of the identity features into key vectors can provide context information related to the image features, and the transformation of the image features into value vectors can provide specific content, clarifying the roles of the image features and the identity features in the attention feature fusion process, so that the attention feature fusion can be accurately achieved.

[0089] In an exemplary embodiment, the step of obtaining non-linear features by performing non-linear feature transformation based on the first fusion feature may include: performing a feature addition operation on the first fusion feature, the query vector, the value vector, and the key vector to obtain a second fusion feature; and performing non-linear feature transformation on the second fusion feature through a trained feed-forward network to obtain non-linear features.

[0090] The step of obtaining the fusion feature corresponding to the image feature by performing feature fusion based on the first fusion feature and the non-linear feature may include: performing a feature addition operation on the second fusion feature and the non-linear feature to obtain the fusion feature corresponding to the image feature.

[0091] Among them, the feature addition operation (Addition) adds the features to be added element by element. The feed-forward network can be implemented by a multi-layer perceptron (abbreviated as MLP) or a position-wise feed-forward network.

[0092] In this embodiment, by performing a feature addition operation on the first fusion feature, the query vector, the value vector, and the key vector, the features before and after the attention feature fusion can be secondarily fused to obtain a second fusion feature, and then the non-linear feature transformation is performed on the second fusion feature through a feed-forward network to improve the expression ability of the features. Further, the features before and after the processing of the feed-forward network are subjected to a feature addition operation for further fusion, improving the richness and expression ability of the features.

[0093] In one embodiment, the fusion in step 106 can be implemented by a trained fusion network, and the structure of the fusion network can be referred to as Figure 2Schematic diagram of the fusion network framework shown. The fusion network may include a multi-head cross-attention module and a feed-forward network. Q, K, and V may respectively represent a query vector, a key vector, and a value vector. Step 106 may include: The computer device may fuse the targeted image feature with the identity feature for each image feature through the fusion network. During the fusion process in the fusion network, the trained query weight matrix and value weight matrix are respectively used to perform a linear transformation on the image feature to obtain a query vector and a value vector, and the trained key weight matrix is used to perform a linear transformation on the identity feature to obtain a key vector. The query vector, the value vector, and the key vector are input into the multi-head cross-attention module, and attention feature fusion is performed through the multi-head cross-attention module to obtain a first fusion feature. The first fusion feature, the query vector, the value vector, and the key vector are subjected to a feature addition operation to obtain a second fusion feature. The second fusion feature is input into the feed-forward network, and a non-linear feature transformation is performed through the feed-forward network to obtain a non-linear feature. The second fusion feature and the non-linear feature are subjected to a feature addition operation to obtain a fusion feature corresponding to the image feature.

[0094] In an exemplary embodiment, the step of extracting multiple image features at different feature levels based on the reference face image in step 104 includes: dividing the reference face image into multiple image patches, and respectively mapping the multiple image patches into multiple patch vectors; based on the multiple patch vectors, sequentially performing feature extraction through multiple stacked encoding layers in the trained image encoder to obtain intermediate features extracted by each of the multiple encoding layers; determining multiple image features at different feature levels according to at least two of the intermediate features extracted by each of the multiple encoding layers.

[0095] Among them, an image patch is an image block in the reference face image. The reference face image can be divided into multiple image patches according to a pre-configured patch size. For example, the patch size is 16*16, and the image size of the reference face image is 224*224, then the reference face image can be divided into 196 image patches. A patch vector is a vector representation of an image patch.

[0096] An encoding layer is a unit for feature extraction in the image encoder. For example, the image encoder can be ViT, then the encoding layer can be a Transformer Encoder (Transformer encoding layer). Different versions of ViT may include different numbers of encoding layers. For example, ViT-B / 16 may include 12 layers of Transformer Encoder, and ViT-L / 14 may include 24 layers of Transformer Encoder.

[0097] The intermediate features are the features extracted by the encoding layers. The intermediate features extracted by each of the multiple encoding layers have different feature levels. Among the multiple stacked encoding layers, the intermediate features extracted by the encoding layer with a later order have a higher feature level. It can be understood that among every two adjacent encoding layers in the multiple stacked encoding layers, the intermediate features extracted by the encoding layer with a later stacking order have a higher feature level than the intermediate features extracted by the encoding layer with an earlier stacking order. The higher the feature level, the higher the degree of abstraction of the features, and the more capable of expressing semantic information such as image categories. The lower the feature level, the higher the degree of detail of the features, and the more capable of expressing the detailed information in the image. For example, it can express the facial makeup and skin texture in the reference face image. When generating an image subsequently, the features with a low feature level can help improve the similarity of the makeup and skin texture between the target face area in the generated target person image and the reference face area in the reference face image, thereby improving the authenticity of the target face area.

[0098] In this embodiment, by dividing the reference face image into multiple image patches and mapping the multiple image patches into multiple patch vectors respectively, the data dimension can be reduced, and the efficiency of subsequent feature extraction can be improved. Furthermore, based on the multiple patch vectors, feature extraction is sequentially performed through the multiple stacked encoding layers, and multiple image features with different feature levels can be determined according to at least two of the intermediate features extracted by each of the multiple encoding layers, thereby creating conditions for improving the authenticity of the generated image subsequently.

[0099] In one embodiment, determining multiple image features with different feature levels according to at least two of the intermediate features extracted by each of the multiple encoding layers may include: The computer device may determine at least two of the intermediate features extracted by each of the multiple encoding layers as multiple image features with different feature levels. Among them, the at least two intermediate features may be, for example, two intermediate features, three intermediate features, or others.

[0100] In an exemplary embodiment, the step of sequentially performing feature extraction through the multiple stacked encoding layers in the trained image encoder based on the multiple patch vectors to obtain the intermediate features extracted by each of the multiple encoding layers may include: Based on the pre-trained category learning vector and the multiple patch vectors, sequentially perform feature extraction through the multiple stacked encoding layers in the trained image encoder. When each encoding layer performs feature extraction, perform feature interaction on the category learning vector and the multiple patch vectors input to the encoding layer to obtain an updated category learning vector and updated multiple patch vectors, and form the intermediate features extracted by the encoding layer based on the updated category learning vector and the updated multiple patch vectors.

[0101] Among them, the class learning vector is used to aggregate the information of multiple patch vectors to learn the global class information of the image. The class learning vector can be called the Class Token, which can be abbreviated as CLS for short. When each encoding layer performs feature extraction, the class learning vector and multiple patch vectors are subjected to feature interaction through the multi-head attention mechanism to obtain an updated class learning vector and updated multiple patch vectors.

[0102] When performing feature interaction through the multi-head attention mechanism, each of the class learning vector and multiple patch vectors can generate their respective corresponding query vectors Q, key vectors K, and value vectors V through linear transformation. The query vector Q corresponding to the class learning vector will calculate the similarity with the key vectors K at all positions (including those corresponding to itself and multiple patch vectors), and finally weighted aggregate the value vectors V at all positions. The query vector Q corresponding to the image patch will interact with the key vectors K at all positions to capture local or global information.

[0103] In this embodiment, based on the pre-trained class learning vector and multiple patch vectors, feature extraction is sequentially performed through multiple encoding layers stacked in the image encoder. Feature interaction can be performed in the feature extraction of each layer. Local information is captured through each patch vector, and class information is aggregated through the class learning vector, thereby improving the feature expression ability of the intermediate features, and gradually improving the feature expression ability as the encoding layer deepens.

[0104] In one embodiment, the computer device can obtain the pre-trained class learning vector, add positional encoding to each of the class learning vector and multiple patch vectors to form input data, and input the input data into multiple encoding layers stacked in the trained image encoder to sequentially perform feature extraction. Among them, the intermediate features can include the class learning vector and multiple patch vectors, as well as the positional encoding corresponding to the class learning vector and multiple patch vectors respectively. The positional encoding of the class learning vector can be 0, indicating that it is located at the first position of the sequence formed by the class learning vector and multiple patch vectors. The positional encoding corresponding to each of the multiple patch vectors can be determined according to the relative position of the image patch corresponding to the patch vector in the reference face image.

[0105] In an exemplary embodiment, the steps of determining multiple image features at different feature levels according to at least two intermediate features among the intermediate features extracted by each of the multiple encoding layers may include: determining the intermediate feature extracted by an encoding layer before the last encoding layer arranged in the stacking order among the multiple encoding layers as the image feature at the first feature level; determining the intermediate feature extracted by the last encoding layer as the image feature at the second feature level; extracting the class learning vector in the image feature at the second feature level to obtain the image feature at the third feature level.

[0106] Among them, the first feature level, the second feature level, and the third feature level increase in sequence, the third feature level is higher than the second feature level, and the second feature level is higher than the first feature level. An encoding layer before the last encoding layer refers to any encoding layer among all the encoding layers before the last encoding layer.

[0107] For example, the image encoder can adopt the image encoder in the CLIP model. Specifically, the image encoder is a ViT with 24 layers of Transformer Encoder. The image features at the first feature level can be the intermediate features extracted by the 18th layer of Transformer Encoder, the image features at the second feature level can be the intermediate features extracted by the 24th layer of Transformer Encoder, and the image features at the third feature level can be the class learning vector (Class Token) extracted from the intermediate features extracted by the 24th layer of Transformer Encoder. At this time, the image features at the third feature level can represent the aggregated image category information.

[0108] In this embodiment, the intermediate features extracted by an encoding layer before the last encoding layer are determined as the image features at the first feature level, the intermediate features extracted by the last encoding layer are determined as the image features at the second feature level, and the class learning vector in the image features at the second feature level is used as the image features at the third feature level. In this way, the three feature levels can cover the low-level image detail features to the high-level image semantic category features. Combining with the subsequent steps, the low-level image detail features can supplement the image detail information, and the high-level image semantic category features can enhance the ability of the identity features to express the identity information, creating conditions for improving the authenticity and consistency of the generated images subsequently.

[0109] In a specific embodiment, referring to the schematic diagram of the image generation architecture as shown in Figure 3 the above image generation method may specifically include the following steps.

[0110] The computer device can obtain a reference face image and original person images. Among them, the reference face image contains a reference face area, and the number of original person images can be multiple, and each original person image respectively contains an original face area.

[0111] The computer device can adopt the trained Arcface model to extract the identity features of the reference face area based on the reference face image.

[0112] The computer device can adopt the image encoder in the trained CLIP model to extract multiple image features at different feature levels based on the reference face image. The multiple image features can include the image features at the first feature level (Image Feature 1), the image features at the second feature level (Image Feature 2), and the image features at the third feature level (Image Feature 3).

[0113] The computer device can pass through the trained MergeNet as shown in Figure 2 to fuse the multiple image features with the identity features respectively, and obtain multiple fused features corresponding to the multiple image features respectively.

[0114] The computer device can perform image segmentation on each original person image respectively, determine the region mask representing the original clothing area corresponding to each original person image, and determine the input features corresponding to each original person image according to each original person image and the corresponding region mask.

[0115] The computer device can perform image generation through the trained image generation model based on the multiple fused features and the input features corresponding to each original person image respectively, so as to edit each original face area according to the reference face area and retain the content of each original clothing area, and generate multiple target person images corresponding to the multiple original person images respectively. Each target person image includes a target face area obtained by editing the face image based on the corresponding original person image, and a target clothing area formed by retaining the content based on the corresponding original clothing area.

[0116] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0117] Based on the same inventive concept, the embodiments of the present application also provide an image generation device for implementing the above-mentioned image generation method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the image generation device provided below can refer to the limitations on the image generation method in the above text, and will not be repeated here.

[0118] In an exemplary embodiment, as Figure 4 shown, an image generation device 400 is provided, including: an acquisition module 410, a feature extraction module 420, a feature fusion module 430, and an image generation module 440, where:

[0119] The acquisition module 410 is configured to acquire a reference face image and an original person image; the reference face image includes a reference face region, and the original person image includes an original face region.

[0120] The feature extraction module 420 is configured to extract the identity feature of the reference face region based on the reference face image, and extract multiple image features at different feature levels based on the reference face image.

[0121] The feature fusion module 430 is configured to fuse multiple image features with the identity feature respectively to obtain multiple fusion features corresponding to the multiple image features respectively.

[0122] The image generation module 440 is configured to perform image generation processing based on the multiple fusion features and the original person image to edit the original face region according to the reference face region, and generate a target person image including the edited target face region; where the number of fusion features is at least two.

[0123] In an exemplary embodiment, the feature fusion module 430 is further configured to, for each image feature among the multiple image features, perform attention feature fusion based on the image feature and the identity feature to obtain a first fusion feature; perform a non-linear feature transformation on the first fusion feature to obtain a non-linear feature; and perform feature fusion according to the first fusion feature and the non-linear feature to obtain a fusion feature corresponding to the image feature.

[0124] In an exemplary embodiment, the feature fusion module 430 is further configured to respectively perform a linear transformation on the image feature by using a trained query weight matrix and a value weight matrix to obtain a query vector and a value vector; perform a linear transformation on the identity feature by using a trained key weight matrix to obtain a key vector; and perform attention feature fusion according to the query vector, the value vector, and the key vector to obtain a first fusion feature.

[0125] In an exemplary embodiment, the feature fusion module 430 is further configured to perform a feature addition operation on the first fusion feature, the query vector, the value vector, and the key vector to obtain a second fusion feature; perform a non-linear feature transformation on the second fusion feature through a trained feed-forward network to obtain a non-linear feature; and perform a feature addition operation on the second fusion feature and the non-linear feature to obtain a fusion feature corresponding to the image feature.

[0126] In an exemplary embodiment, the feature extraction module 420 is further configured to segment a reference face image into a plurality of image patches, map the plurality of image patches into a plurality of patch vectors respectively; based on the plurality of patch vectors, sequentially perform feature extraction through a plurality of stacked encoding layers in the trained image encoder to obtain intermediate features extracted by each of the plurality of encoding layers; determine a plurality of image features at different feature levels according to at least two of the intermediate features extracted by each of the plurality of encoding layers.

[0127] In an exemplary embodiment, the feature extraction module 420 is further configured to sequentially perform feature extraction through a plurality of stacked encoding layers in the trained image encoder based on a pre-trained class learning vector and the plurality of patch vectors. When each encoding layer performs feature extraction, perform feature interaction on the class learning vector input to the encoding layer and the plurality of patch vectors to obtain an updated class learning vector and updated plurality of patch vectors, and form intermediate features extracted by the encoding layer based on the updated class learning vector and the updated plurality of patch vectors.

[0128] In an exemplary embodiment, the feature extraction module 420 is further configured to determine the intermediate features extracted by an encoding layer that is one before the last encoding layer in the stacking order among the plurality of encoding layers as the image features at the first feature level; determine the intermediate features extracted by the last encoding layer as the image features at the second feature level; extract the class learning vector in the image features at the second feature level to obtain the image features at the third feature level.

[0129] Each module in the above image generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0130] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data that needs to be stored when executing the above image generation method. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an image generation method.

[0131] Those skilled in the art can understand that Figure 5 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0132] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0133] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0134] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0136] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0137] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.

[0138] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An image generation method, characterized in that: The method comprises: Acquire a reference face image and an original person image; the reference face image includes a reference face region, and the original person image includes an original face region; Extracting identity features of the reference face region based on the reference face image, and extracting multiple image features at different feature levels based on the reference face image; fusing the plurality of image features with the identity features respectively to obtain a plurality of fused features respectively corresponding to the plurality of image features; Performing image generation processing based on the multiple fusion features and the original person image to edit the original face region according to the reference face region to generate a target person image including the edited target face region; The number of the fusion features is at least two.

2. The method according to claim 1, characterized in that The step of fusing the plurality of image features with the identity features to obtain a plurality of fused features corresponding to the plurality of image features, respectively, includes: For each image feature of the multiple image features, performing attention feature fusion based on the image feature and the identity feature to obtain a first fused feature; Performing nonlinear feature transformation based on the first fusion feature to obtain a nonlinear feature; Feature fusion is performed according to the first fusion feature and the nonlinear feature to obtain a fusion feature corresponding to the image feature.

3. The method according to claim 2, characterized in that The fusing of attention features based on the image features and the identity features to obtain a first fused feature includes: Using the trained query weight matrix and value weight matrix respectively, linearly transform the image features to obtain a query vector and a value vector; Using a trained key weight matrix to linearly transform the identity feature to obtain a key vector; Attention features are fused according to the query vector, the value vector and the key vector to obtain a first fused feature.

4. The method according to claim 3, characterized in that The performing nonlinear feature transformation based on the first fusion feature to obtain a nonlinear feature includes: Performing a feature addition operation on the first fused feature, the query vector, the value vector, and the key vector to obtain a second fused feature; Performing nonlinear feature transformation on the second fusion feature through a trained feedforward network to obtain a nonlinear feature; The performing feature fusion according to the first fusion feature and the nonlinear feature to obtain a fusion feature corresponding to the image feature includes: Perform a feature addition operation on the second fusion feature and the nonlinear feature to obtain a fusion feature corresponding to the image feature.

5. The method according to any one of claims 1 to 4, characterized in that: The extracting of multiple image features at different feature levels based on the reference face image includes: Segmenting the reference face image into a plurality of image patches, and mapping the plurality of image patches into a plurality of patch vectors respectively; Based on the multiple patch vectors, sequentially extracting features through multiple encoding layers stacked in a trained image encoder to obtain intermediate features extracted by each of the multiple encoding layers; A plurality of image features at different feature levels are determined according to at least two of the intermediate features extracted from each of the plurality of coding layers.

6. The method according to claim 5, characterized in that The step of extracting features in sequence through a plurality of coding layers stacked in a trained image encoder based on the plurality of patch vectors to obtain intermediate features extracted by each of the plurality of coding layers comprises: Based on the pre-trained category learning vector and the multiple patch vectors, feature extraction is performed in sequence through multiple coding layers stacked in the trained image encoder. When each coding layer performs feature extraction, the category learning vector and the multiple patch vectors input to the coding layer are subjected to feature interaction to obtain updated category learning vectors and updated multiple patch vectors, and intermediate features extracted by the coding layer are formed based on the updated category learning vector and the updated multiple patch vectors.

7. The method according to claim 6, characterized in that The determining of a plurality of image features at different feature levels according to at least two of the intermediate features extracted from the plurality of coding layers comprises: Determine an intermediate feature extracted from a coding layer that is arranged before the last coding layer in the stacking order among the multiple coding layers as an image feature of a first feature level; Determine the intermediate features extracted from the last coding layer as image features of the second feature level; The category learning vectors are extracted from the image features at the second feature level to obtain image features at a third feature level.

8. An image generating device, characterized in that: The device comprises: An acquisition module, used to acquire a reference face image and an original character image; the reference face image includes a reference face region, and the original character image includes an original face region; A feature extraction module, used to extract the identity feature of the reference face area based on the reference face image, and to extract multiple image features of different feature levels based on the reference face image; A feature fusion module, used to fuse the multiple image features with the identity feature respectively to obtain multiple fusion features corresponding to the multiple image features respectively; An image generation module is used to perform image generation processing based on the multiple fusion features and the original character image, so as to edit the original face area according to the reference face area to generate a target character image containing the edited target face area; wherein the number of the fusion features is at least two.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.