Image generation method and device, computer equipment and storage medium

By extracting and fusing facial and textual features within a diffusion model, the method enhances the accuracy and efficiency of image generation, addressing inefficiencies in existing controlled text-to-image techniques.

CN120318346APending Publication Date: 2025-07-15NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410052526.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing controlled image generation methods have low efficiency in training models and need to be improved, making it difficult to generate high-quality face images that conform to the description text.

Method used

By obtaining the initial face image and description text, the face features and text features are extracted, and after compression processing, the target diffusion model is used to perform feature fusion to generate the target face image that conforms to the description text.

Benefits of technology

It improves the accuracy and efficiency of the image generation model and can generate high-quality face images that conform to the description text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318346A_ABST
    Figure CN120318346A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image generation method and device, computer equipment and a storage medium. According to the scheme, the initial face image and the description text are acquired, the face features of the initial face image are extracted, the text features of the description text are extracted, then the initial face image is compressed to obtain the hidden space features of the initial face image, and further, the face features, the text features and the hidden space features are fused to obtain the hidden space features of the initial face image. The fused features are obtained; and based on the fused features, generating a target face image corresponding to the initial face image and conforming to the description text. Therefore, the image generation accuracy of the image generation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an image generation method, apparatus, computer device, and storage medium. Background Art

[0002] With the rise of diffusion models, text-to-image generation has attracted wide attention, especially controllable generation. Diffusion models are a type of generative model that convert Gaussian noise into samples of a known data distribution through an iterative denoising process, and the generated images have good diversity and realism. And text-to-image generation is one of the multimodal tasks, and much work on this task is also constructed based on diffusion models.

[0003] In related technologies, the controllable generation method is based on a fixed faceId, and additional network training needs to be performed on the target face, or an energy function is introduced in the sampling stage for gradient guidance. However, these methods have low efficiency in training the model, and the effect of the trained model needs to be improved. Summary of the Invention

[0004] Embodiments of this application provide an image generation method, apparatus, computer device, and storage medium, which improve the accuracy of image generation of an image generation model.

[0005] Embodiments of this application provide an image generation method, including:

[0006] Obtain an initial face image and a description text;

[0007] Extract the face features of the initial face image and the text features of the description text;

[0008] Perform compression processing on the initial face image to obtain the latent space features of the initial face image;

[0009] Perform fusion processing on the face features, the text features, and the latent space features to obtain fused features;

[0010] Generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0011] Correspondingly, embodiments of this application also provide an image generation apparatus, including:

[0012] A first acquisition unit, configured to obtain an initial face image and a description text;

[0013] A first extraction unit, configured to extract the face features of the initial face image and the text features of the description text;

[0014] A first processing unit for compressing the initial face image to obtain the latent space features of the initial face image;

[0015] A second processing unit for fusing the face features, the text features, and the latent space features to obtain the fused features;

[0016] A generating unit for generating a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0017] Correspondingly, an embodiment of the present application further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the image generation method provided in any embodiment of the present application.

[0018] Correspondingly, an embodiment of the present application further provides a storage medium storing multiple instructions suitable for being loaded by a processor to execute the above image generation method.

[0019] In the embodiment of the present application, by obtaining an initial face image and a description text, extracting the face features of the initial face image, and extracting the text features of the description text, then, compressing the initial face image to obtain the latent space features of the initial face image. Further, fusing the face features, the text features, and the latent space features to obtain the fused features; generating a target face image corresponding to the initial face image that conforms to the description text based on the fused features. In this way, the image generation accuracy of the image generation model can be improved. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of an image generation method provided by an embodiment of the present application.

[0022] Figure 2 It is a schematic diagram of an application scenario of an image generation method provided by an embodiment of the present application.

[0023] Figure 3 It is a schematic diagram of another application scenario of an image generation method provided by an embodiment of the present application.

[0024] Figure 4 It is a schematic diagram of another application scenario of an image generation method provided by an embodiment of the present application.

[0025] Figure 5 This is a structural block diagram of an image generation device provided by an embodiment of the present application.

[0026] Figure 6 This is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0028] The embodiments of the present application provide an information recommendation method, device, storage medium, and computer device. Specifically, the information recommendation method in the embodiments of the present application can be executed by a computer device, where the computer device can be a terminal or a server, etc. The terminal can be a smart phone, a tablet computer, a notebook computer, a touch screen, a personal computer (PC), a personal digital assistant (PDA), and other terminal devices. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0029] For example, the computer device can be a server, and the server can obtain an initial face image and a description text; extract the face features of the initial face image and the text features of the description text; perform compression processing on the initial face image to obtain the latent space features of the initial face image; perform fusion processing on the face features, text features, and latent space features to obtain the fused features; and generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0030] Based on the above problems, the embodiments of the present application provide a first image generation method, device, computer device, and storage medium to improve the image generation accuracy of the image generation model.

[0031] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0032] An embodiment of the present application provides an image generation method. This method can be executed by a terminal or a server. In this embodiment of the present application, the case where the image generation method is executed by the server is taken as an example for illustration.

[0033] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of an image generation method provided by an embodiment of the present application. The specific process of this image generation method can be as follows:

[0034] 101. Obtain an initial face image and a description text.

[0035] Among them, the initial face image refers to a face picture corresponding to a specified face identity. For example, if the specified face identity is the face of user A, the initial face image can be a facial image of user A.

[0036] Among them, the description text refers to the text content describing the image to be generated. For example, the description text can be "black hair, smiling" and so on.

[0037] 102. Extract the face features of the initial face image and the text features of the description text.

[0038] In the embodiment of the present application, after obtaining the initial face image, the first face image can be preprocessed first, and then the face features can be extracted from the preprocessed initial face features to obtain the face features.

[0039] Among them, the image preprocessing can include various processing methods. For example, face alignment, image cropping, image scaling, etc.

[0040] In some embodiments, in order to obtain accurate face features of the face image, the step of "extracting the face features of the initial face image" may include the following operations:

[0041] Perform face detection on the initial face image to determine the face key points in the initial face image;

[0042] Perform alignment processing on the initial face image based on the face key points to obtain an aligned face image;

[0043] Extract features from the aligned face image to obtain face features.

[0044] Among them, face detection of the initial face image can be performed through face detection technology. For example, the initial face image can be detected through MTCNN (Multi-task convolutional neural network).

[0045] Among them, MTCNN is divided into three network structures: P-Net, R-Net, and O-Net. P-Net is used to quickly generate candidate windows, R-Net is used to filter and select candidate windows with high precision, and O-Net is used to generate the final bounding box and facial key points. The working principle of MTCNN is as follows:

[0046] 1) Make an image pyramid: Stack images of different sizes from large to small in a pyramid shape, resize the input image to different sizes, and prepare it for input into the network.

[0047] 2) Input the pyramid images into P-Net (Proposal Network) to obtain Proposal bounding boxes (candidate bounding boxes) containing faces, and remove redundant boxes through the non-maximum suppression (NMS) algorithm to initially obtain some face detection candidate boxes. Among them, non-maximum suppression is to suppress elements that are not maxima. In the field of object detection, this method can be used to quickly remove prediction boxes with high overlap and relatively inaccurate calibration.

[0048] 3) Input the face images output by P-Net into R-Net (Refinement Network) to further refine the coordinates of the face detection boxes, and remove redundant boxes through the NMS algorithm. At this time, the obtained face detection boxes are more accurate and have fewer redundant boxes.

[0049] 4) Input the face images output by R-Net into O-Net (Output Network). On the one hand, further refine the coordinates of the face detection boxes, and on the other hand, output the coordinates of the key points of the face (such as: left eye, right eye, nose, left mouth corner, right mouth corner).

[0050] In the embodiment of this application, the initial face image is input into MTCNN. First, the initial face image is transformed at different scales to construct an image pyramid to adapt to the detection of faces of different sizes; for the image pyramid constructed in the previous step, preliminary feature extraction and calibration of the border are performed through a FCN, and the window is adjusted through Bounding-Box Regression and most windows are filtered through NMS, and many face region windows where faces may exist are output; all the predicted face region windows are sent into R-Net, and a large number of candidate boxes with poor effects are filtered out through R-Net. Finally, the selected candidate boxes are further optimized through Bounding-Box Regression and NMS, and more reliable face region boxes are output; finally, input into O-Net, and the facial region is identified through more supervision, and the facial feature points of people will be regressed, and finally five facial key points of the face are output.

[0051] Among them, for aligning the initial face image based on face key points, face alignment can be performed through affine transformation. The affine transformation can align the face according to the eye coordinates in the face image, calculate the corresponding coordinates after transformation, and implement the face alignment function.

[0052] Among them, for face alignment processing according to face key points, it can specifically include the following steps:

[0053] First, obtain the positions of the left eye corner key point and the right eye corner key point from the face key points, calculate the center point position between the two eyes according to the positions of the left eye corner key point and the right eye corner key point, calculate the coordinate difference Δx on the X-axis between the two eye corners, and calculate the coordinate difference Δy on the Y-axis between the two eye corners. Further, calculate the rotation angle required for face alignment according to the center point position, the coordinate difference Δx, and the coordinate difference Δy.

[0054] Among them, the formula for calculating the center point position between the two eyes according to the positions of the left eye corner key point and the right eye corner key point can be as follows:

[0055]

[0056] Among them, x c is the X coordinate of the center point position, y c is the Y coordinate of the center point position; x r is the X coordinate of the right eye corner key point, y r is the Y coordinate of the right eye corner key point; x1 is the X coordinate of the left eye corner key point, and y1 is the Y coordinate of the left eye corner key point.

[0057] Among them, the formula for calculating the rotation angle according to the center point position, the coordinate difference Δx, and the coordinate difference Δy can be as follows:

[0058]

[0059] Among them, θ is the rotation angle, Δy represents the coordinate difference on the Y-axis between the left eye corner key point and the right eye corner key point; Δx represents the coordinate difference on the X-axis between the left eye corner key point and the right eye corner key point.

[0060] In some embodiments, in order to ensure that the facial features of the extracted initial face image exclude the features of non-face regions in the image, the aligned face image can be cropped, scaled, etc. Specifically, the facial bounding box in the initial face image can be obtained, and the size of the facial bounding box can be adjusted so that the facial bounding box includes the entire face region. Then, image cropping can be performed according to the adjusted facial bounding box to retain the face region in the image. Finally, the cropped image can be scaled and adjusted so that the size of the face image conforms to the image size processed by the model. The finally obtained face image is used as the processed initial face image.

[0061] For example, please refer to Figure 2 , Figure 2 which is a schematic diagram of an application scenario of an image generation method provided by an embodiment of the present application. In Figure 2 the shown initial face image, the face is offset. By performing face alignment processing on the initial face image, the face offset is corrected to obtain an aligned face image. In the aligned face image, the facial features are in the positions of the reference facial features.

[0062] Furthermore, for feature extraction of the processed initial face image, face feature extraction can be performed on the face image through ArcFace.

[0063] Among them, ArcFace is a convolutional neural network for face feature extraction. Compared with traditional neural networks, the neurons in the ArcFace network establish connections with some neurons in the previous layer, which can greatly reduce the complexity, reduce the computational amount, and improve the success rate.

[0064] In some embodiments, when performing feature extraction on the initial face image through ArcFace, the obtained face features can be input into an MLP (Multilayer Perceptron) for mapping to obtain the mapped face features.

[0065] Among them, MLP is a neural network model for classification, regression, and clustering. Its principle is to map the input data to the output space through multiple non-linear transformations and continuously adjust the weights during the training process to improve the accuracy of the model. MLP is a group of multiple perceptrons on each layer. MLP consists of three layers: an input layer, a hidden layer, and an output layer. The input layer only receives the input, the hidden layer processes the input, and the output layer generates the result. Basically, the weights need to be trained for each layer. The multilayer perceptron can learn any non-linear function. MLP can learn the weights for mapping any input to the output.

[0066] In some embodiments, to extract features from the description text, the CLIP (Contrastive Language-Image Pre-Training) can be used to extract features from the description text.

[0067] Among them, CLIP is a model pre-trained with text as the supervision signal, and its characteristics are as follows: a multi-modal model involving text and images; through contrastive learning, it calculates the cosine similarity between text features and image features, enabling the model to learn the matching relationship between text and images, including: positive samples: text and images match; negative samples: text and images do not match; it includes two encoders that encode text and images into a fixed length and then calculate the cosine similarity; among them, one encodes text, generally a model based on Transformer; the other encodes images, which can be a CNN (such as ResNet50) or a ViT model.

[0068] Specifically, the description text can be input into CLIP, and the text encoder of CLIP can be used to encode the description text to obtain the text features corresponding to the description text.

[0069] 103. Compress the initial face image to obtain the latent space features of the initial face image.

[0070] Among them, latent space features refer to the patterns and relationships hidden in the data, and latent space features are very important for machine learning and deep learning tasks. For example, in image recognition, latent space features can represent the texture and shape information in the image. In natural language processing, latent space features can represent the semantic information in the text. In addition, latent space features can also be used in fields such as speech recognition and recommendation systems.

[0071] To extract latent space features, various machine learning and deep learning algorithms can be used. For example, using a Convolutional Neural Network (CNN) can extract features from images and capture local and global patterns in the images; or, using a Recurrent Neural Network (RNN) can model sequential data and capture temporal dependencies and context information; or, using an Autoencoder can compress and reconstruct data and extract its implicit structure.

[0072] In this embodiment, the initial face image can be compressed by an autoencoder to obtain the latent space features of the initial face image.

[0073] In some embodiments, the step of "performing compression processing on the initial face image to obtain the latent space feature of the initial face image" may include the following operations:

[0074] Input the initial face image into an encoder, and through the encoder, map the initial face image to a low-dimensional latent variable to obtain the latent space feature.

[0075] Among them, the encoder can be a Variational Autoencoder (VAE). VAE is a generative model. The core idea of VAE is the encoder and the decoder. The encoder compresses the input data into a vector in a latent space, that is, a latent vector, and the decoder maps the latent vector back to the data in the original space. VAE also includes a loss function for optimizing the parameters of the model to make the data generated by it as close as possible to the real data distribution.

[0076] Specifically, the variational autoencoder is a generative model based on neural networks and is a variant of the autoencoder. Its structure consists of two parts: an encoder and a decoder. The encoder maps the input data into a latent space, and the decoder remaps the vector in the latent space into the output of the original data. In the latent space, VAE samples a random vector z, which is called a "latent variable". The basic principle of VAE is to introduce a latent space on the basis of the autoencoder and transform it into a trainable model by introducing the idea of variational inference.

[0077] In a VAE, the latent space is modeled as a Gaussian distribution, which means that latent vectors can be sampled from this distribution. The mean and standard deviation of this Gaussian distribution are both output by the encoder network. Usually, the latent vectors in a VAE are sampled from a standard normal distribution with a mean of 0 and a variance of 1. In this way, the VAE can sample new data samples from the distribution of the original data, thus generating new data. The inference process of the VAE can be divided into two steps: Encoding and Decoding. The encoder maps the input data x to the latent variable z in the latent space. This mapping consists of two parts: a variational layer for calculating the mean and variance of the latent variable, and a noise vector ε sampled from the standard normal distribution. The variational layer generates the mean and variance based on the input data x, and then calculates the latent vector z by sampling ε. This sampling process uses the reparameterization trick, which reparameterizes the connection between the sampling and the variational layer so that the gradient can be passed through the sampling process. The decoder maps the latent variable z to the output of the original data. The input of the decoder is the latent vector z and some noise, and the role of the noise is to make the decoder more robust, so that it can generate continuous outputs in the space around the latent vector z.

[0078] Specifically, the decoder uses a feed-forward neural network to take the latent vector z and the noise as inputs and outputs the reconstructed data x1. The output of the decoder is scaled by a sigmoid or tanh layer to ensure that the output values are within the range of the original data. During the inference process, first the input data x is fed into the encoder to obtain the mean and variance of the latent vector z. Then a vector ε is sampled from the standard normal distribution, and the latent vector z is calculated using the mean and variance of the latent vector and the vector ε. Finally, z is fed into the decoder to obtain the reconstructed data x1.

[0079] For example, when an initial face image is input into the VAE, the initial face image is compressed into a vector in the latent space by the encoder of the VAE, which can be the latent space features.

[0080] 104. Perform a fusion process on the face features, text features, and latent space features to obtain the fused features.

[0081] In some embodiments, the step "Perform a fusion process on the face features, text features, and latent space features to obtain the fused features" may include the following operations:

[0082] Input the face features, text features, and latent space features into the target diffusion model;

[0083] Process the face features, text features, and latent space features through the cross-attention module of the target diffusion model to obtain the fused features.

[0084] In an embodiment of the present application, the target diffusion model can be used to generate a face image that includes a face identity in the input face image that conforms to the description text according to the input face image and the description text. The target diffusion model may include a cross-attention module and an image generation module. Among them, the cross-attention module can be used to perform cross-attention calculations on features of different modalities in the input; the image generation module can be used to generate a target face image according to the fused features.

[0085] Among them, the cross-attention module uses the CrossAttention mechanism. The CrossAttention mechanism introduces an additional input sequence on the basis of self-attention to fuse information from multiple sources. In machine translation, for example, the source language sentence and the target language sentence are regarded as two different input sequences and affect each other through the CrossAttention mechanism, so as to better capture the dependencies between the two languages.

[0086] Specifically, CrossAttention actually refers to the cross-attention layer between the encoder and the decoder. In this layer, the decoder adjusts the attention of the output of the encoder to obtain encoder information related to the current decoding position. In the encoder-decoder architecture, the encoder is responsible for encoding the input sequence into a series of feature vectors, and the decoder gradually generates the output sequence according to these feature vectors. In order to enable the decoder to effectively model the context of the current generation position, the CrossAttention layer is introduced.

[0087] Among them, the calculation process of CrossAttention can include the following steps: Encoder input (usually the output from the encoder): Usually represented as enc_inputs, with a size of (batch_size, seq_len_enc, hidden_dim); Decoder input (the partially generated sequence): They are usually represented as dec_inputs, with a size of (batch_size, seq_len_dec, hidden_dim). Each position of the decoder generates a query vector (query) to calculate the attention weights at all positions of the encoder. All positions of the encoder generate a set of key vectors (keys) and value vectors (values). Perform a dot product operation using the query vector (query) and the key vectors (keys), and obtain the attention weights through the softmax function. Multiply the attention weights by the value vectors and sum the results to obtain the output adjusted by the encoder.

[0088] Among them, the face feature, text feature, and latent space feature are input into the target diffusion model. First, the cross-attention module in the target diffusion model can perform cross-attention calculations on the input face feature, text feature, and latent space feature to obtain the fused feature.

[0089] In some embodiments, the step of "processing the face feature, text feature, and latent space feature through the cross-attention module of the target diffusion model to obtain the fused feature" may include the following operations:

[0090] Perform cross-attention calculation on the face feature and text feature through the cross-attention module to obtain the first feature;

[0091] Perform cross-attention calculation on the text feature and latent space feature through the cross-attention module to obtain the second feature;

[0092] Perform cross-attention calculation on the image feature and latent space feature through the cross-attention module to obtain the third feature;

[0093] Fuse based on the first feature, second feature, and third feature to obtain the fused feature.

[0094] In the embodiments of the present application, the cross-attention module may include multiple cross-attention sub-modules. For example, it may include a first cross-attention sub-module, a second cross-attention sub-module, and a third cross-attention sub-module.

[0095] Among them, the first cross-attention sub-module can be used to calculate the cross-attention between the face feature and the text feature; the second cross-attention sub-module can be used to calculate the cross-attention between the text feature and the latent space feature; the third cross-attention sub-module can be used to calculate the cross-attention between the face feature and the latent space feature.

[0096] For example, please refer to Figure 3 , Figure 3 is a schematic diagram of an application scenario of another image generation method provided by the embodiments of the present application. Among them, cross-attention module 1 can be the first cross-attention sub-module; cross-attention module 2 can be the second cross-attention sub-module; cross-attention module 3 can be the third cross-attention sub-module.

[0097] Among them, the face feature and the text feature can be input into the first cross-attention sub-module. Through the first cross-attention sub-module, cross-attention calculation is performed on the face feature and the text feature, and the calculated feature is output as the first feature; the text feature and the latent space feature are input into the second cross-attention sub-module. Through the second cross-attention sub-module, cross-attention calculation is performed on the text feature and the latent space feature, and the calculated feature is output as the second feature; the face feature and the latent space feature are input into the third cross-attention sub-module. Through the third cross-attention sub-module, cross-attention calculation is performed on the face feature and the latent space feature, and the calculated feature is output as the third feature.

[0098] Further, feature fusion is performed on the first feature, the second feature, and the third feature to obtain a fused feature. The fused feature can include information of the face feature, information of the text feature, and information of the latent space feature. Among them, feature fusion is a method of fusing multiple features together to obtain more complete and reliable information.

[0099] In some embodiments, to improve the image generation efficiency of the target diffusion model, before the step of "inputting the face feature, the text feature, and the latent space feature into the target diffusion model", the method may further include the following steps:

[0100] Obtain a sample set;

[0101] Based on the sample face images and sample description texts in the sample pairs, construct a target diffusion model.

[0102] Among them, the sample set may include multiple sample pairs. Each sample pair includes a sample face image and a sample description text corresponding to the sample face image.

[0103] For example, the sample description text in a certain sample pair may be "yellow hair, wearing a scarf", then the sample face image in this sample pair may be an image of a face identity with yellow hair and wearing a scarf.

[0104] In the embodiments of the present application, the sample pairs in the sample set can be used to train the target diffusion model.

[0105] Among them, obtaining the sample set may include: first collecting multiple sample face images, then specifically describing the sample face images to obtain sample description texts corresponding to the sample face images, and constructing an image-text pair based on each sample face image and the corresponding sample description text as a sample pair. Thus, a sample set including multiple sample pairs can be obtained.

[0106] Among them, the sampled face images can be obtained by photographing face images or retrieving them from an existing face image database. Among them, multiple sampled face images include face images with different face IDs.

[0107] Among them, the sampled face images to be collected need to be images that can detect faces, and preferably frontal face images. The sample description text should describe the sampled face images as fully as possible, such as including facial expressions and head adornments.

[0108] In some embodiments, to improve the accuracy of the generated images, the step of "constructing a target diffusion model based on the sampled face images and sample description text in the sample pair" may include the following operations:

[0109] Extract features from the sampled face images to obtain the sampled face features corresponding to the sampled face images;

[0110] Perform compression processing on the sampled face images to obtain the sample latent space features corresponding to the sampled face images;

[0111] Extract features from the sample description text to obtain the sample text features corresponding to the sample description text;

[0112] Train a preset diffusion model based on the sampled face features, sample latent space features, and sample text features to obtain a target diffusion model.

[0113] In the embodiments of the present application, in order to eliminate irrelevant information in the images, restore useful real information, enhance the detectability of relevant information, and simplify the data to the greatest extent, thereby improving the reliability of feature extraction, image segmentation, matching, and recognition, the sampled face images in the sample pair can be preprocessed.

[0114] Among them, the preprocessing performed on the sampled face images may include face alignment, cropping, scaling, etc.

[0115] Specifically, the preprocessing process may include: First, use a face detection method to detect the sampled face images, detect the face region from the sampled face images, and at the same time perform key point detection to detect the face key points. Then, according to the detected face key points, perform face alignment processing on the sampled face images. Further, obtain the facial frame in the sampled face images, adjust the size of the facial frame so that the facial frame includes the entire face region, and then, according to the adjusted facial frame, perform image cropping to retain the face region in the sampled face images. Finally, the cropped image can be scaled and adjusted so that the size of the face image meets the image size required for model processing, and the finally obtained face image is used as the preprocessed sampled face image.

[0116] Among them, the feature extraction of the sample face image can be the feature extraction of the processed sample face image. Specifically, the face features are extracted by the ArcFace network, and the face features extracted by the ArcFace network pass through an MLP layer to obtain the sample face features.

[0117] Among them, for the compression processing of the sample face image, the encoder in the VAE can be used to compress the sample face image, so as to obtain the sample latent space features corresponding to the sample face image.

[0118] Among them, for the feature extraction of the sample description text, CLIP can be used to extract the features of the sample description text, so as to obtain the sample text features corresponding to the sample description text.

[0119] Among them, the preset diffusion model can be a Latent Diffusion Model (LDM). The diffusion process occurs in the latent space. The LDM includes an autoencoder that learns to compress image data into a low-dimensional representation. By using the encoder, a full-size image can be encoded into low-dimensional latent data (compressed data). Then, through the used decoder D, the latent data is decoded back into an image. After the codec training is completed, a denoising probability model is trained on the low-dimensional latent data. In this process, the encoded text information is introduced to generate an image corresponding to the text description.

[0120] In some embodiments, the method may further include the following steps:

[0121] Obtain time information, and perform feature extraction on the time information to obtain the time features corresponding to the time information;

[0122] Perform splicing processing on the sample face features and the time features to obtain the spliced features;

[0123] In the embodiments of the present application, the sample face features can be used in two steps. One is to splice with the time features, and the other is to perform cross-attention calculation with the text features and the latent space features.

[0124] Among them, when training the preset diffusion model, a time step from 0 to 999 is initialized, and this time step is mapped into a Timeembedding (time embedding) through an MLP, that is, the time features.

[0125] Among them, time embedding is a method of converting time information into a numerical representation, which is usually used in sequence modeling tasks in natural language processing, such as machine translation, text summarization, sentiment analysis, etc. Time embedding can map time information into a fixed-length vector, facilitating the computer to process and understand time information. In natural language processing, time embedding is often used as an additional dimension of word vectors, which can improve the model's understanding and expression ability of time information.

[0126] In the embodiments of the present application, the reason for adding Time embeding: U-net shares parameters, and adding Timeembeding can generate different outputs according to different inputs. It is hoped that during the reverse diffusion process, U-net can first generate some approximate outlines, very rough coarse images, which do not need to be very clear. As the reverse diffusion progresses, when approaching the restoration of the original image, it can learn the information of the edges and corners of the object, some high-frequency information, so as to make the output image more realistic.

[0127] In some embodiments, the step of "training a preset diffusion model based on the sample face features, sample latent space features, and sample text features to obtain a target diffusion model" may include the following operations:

[0128] Training a preset diffusion model based on the concatenated features, sample latent space features, and sample text features to obtain a target diffusion model.

[0129] Specifically, inputting the concatenated features, sample latent space features, and sample text features into the preset diffusion model for training to obtain a target diffusion model.

[0130] In some embodiments, to improve the model training effect, the step of "training a preset diffusion model based on the sample face features, sample latent space features, and sample text features to obtain a target diffusion model" may include the following operations:

[0131] Processing the sample face features, sample text features, and sample latent space features through the cross-attention module of the preset diffusion model to obtain the fused sample features;

[0132] Generating a generated face image through the image generation module of the preset diffusion model based on the fused sample features;

[0133] Determining the difference information between the sample face image and the generated face image;

[0134] Adjusting the preset diffusion model based on the difference information and the preset loss function to obtain a target diffusion network.

[0135] Among them, processing the sample face features, sample text features, and sample latent space features through the cross-attention module of the preset diffusion model may include: inputting the sample face features and sample text features into the first cross-attention sub-module, performing cross-attention calculation on the sample face features and sample text features through the first cross-attention sub-module, and outputting the calculated features as the first sample features; inputting the sample text features and sample latent space features into the second cross-attention sub-module, performing cross-attention calculation on the sample text features and sample latent space features through the second cross-attention sub-module, and outputting the calculated features as the second sample features; inputting the sample face features and sample latent space features into the third cross-attention sub-module, performing cross-attention calculation on the sample face features and sample latent space features through the third cross-attention sub-module, and outputting the calculated features as the third sample features.

[0136] In the embodiments of the present application, through the above triple cross-attention mechanism, the influence of face features and text features on face image generation is fully considered, so that the trained diffusion model can generate accurate face images including the specified face ID and conforming to the specified description text, realizing fine-grained controllable generation of the diffusion model.

[0137] Furthermore, feature fusion is performed on the first sample features, second sample features, and third sample features to obtain the fused sample features.

[0138] Among them, the image generation model may include a denoising diffusion sub-module and an image decoding sub-module. Among them, the denoising diffusion sub-module may be a Denoising U-net, and the denoising diffusion sub-module may use a Denoising Diffusion Probabilistic Model (DDPM) to perform denoising processing on the fused sample features to obtain the denoised fused features. Among them, the image decoding sub-module may be a decoder, which can be used to restore the denoised fused features into an image.

[0139] Among them, DDPM is a generative model based on the diffusion process. It gradually adds random noise to the data and then learns the inverse diffusion process to construct the required data samples from the noise. The main idea of DDPM is to use a random process to generate a series of noisy images from the data samples under gradually increasing noise conditions, and then restore them to clear images through denoising. This process can be formally represented as a Markov chain and learned using the gradient backpropagation algorithm. DDPM uses a fixed procedure to learn, and the latent variables have the same high dimension as the original data. In addition, DPM only needs to calculate the negative log-likelihood as the loss function, avoiding the problems of strict restrictions on the model structure or dependence on proxy objectives.

[0140] Specifically, the denoising diffusion sub-module performs denoising on the fused sample features to obtain the denoised fused features, and then inputs the denoised fused features into the image decoding sub-module. The image decoding sub-module restores the denoised fused features into an image to obtain the generated face image.

[0141] Among them, determining the difference information between the sample face image and the generated face image may include calculating the difference between the face features of the sample face image and the face features of the generated face image to obtain the feature difference between the sample face image and the generated face image as the difference information.

[0142] In the embodiments of the present application, a loss function is provided, and the loss of the feature difference is calculated through the loss function. Among them, the loss function is as follows:

[0143]

[0144]

[0145] Among them, is the loss function of the denoising diffusion sub-module, Z t represents the sample latent space feature vector, C t represents the sample text feature, W ID represents the sample face feature, t = 1...T, which can be understood as a series of denoising autoencoders with equal weights.

[0146] Among them, L represents the preset loss function of the preset diffusion model, is the face feature of the generated face. Through the loss function L, the loss of the distance between the generated face image and the sample image can be calculated.

[0147] Among them, adjusting the preset diffusion model based on the difference information and the preset loss function may include iteratively training the preset diffusion model with the loss calculated by the preset loss function until the preset diffusion model converges, and a target diffusion network can be obtained.

[0148] In the embodiments of the present application, by collecting face images of multiple different face IDs and extracting the face features of different face IDs to train the diffusion model, it is possible to guide the diffusion model to generate high-resolution and high-quality face images during the training process.

[0149] 105. Generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0150] In some embodiments, the step of "generating a target face image corresponding to the initial face image that conforms to the description text based on the fused features" may include the following operations:

[0151] Perform denoising processing on the fused features to obtain denoised features;

[0152] Perform image generation based on the denoised features to obtain the target face image.

[0153] Specifically, the fused denoised features can be input into the denoising diffusion sub-module, and the denoising diffusion sub-module performs denoising processing on the fused features to obtain denoised features. Then, the denoised features are input into the image decoding sub-module, and the image decoding sub-module restores the denoised fused features into an image to obtain the target face image.

[0154] Among them, the target face image can be a face image that includes the face of user A and conforms to the description text.

[0155] The embodiments of the present application disclose an image generation method, which includes: obtaining an initial face image and a description text; extracting the face features of the initial face image and the text features of the description text; performing compression processing on the initial face image to obtain the latent space features of the initial face image; performing fusion processing on the face features, text features, and latent space features to obtain fused features; generating a target face image corresponding to the initial face image that conforms to the description text based on the fused features. In this way, the image generation accuracy of the image generation model can be improved.

[0156] For example, please refer to Figure 4 , Figure 4 which is a schematic diagram of an application scenario of an image generation method provided by the embodiments of the present application. Figure 4It shows the process of training a preset diffusion model to obtain a target diffusion model. First, sample face images and sample description texts are obtained. The sample face images are subjected to face detection by a face detector to identify the face regions in the sample face images, and then affine transformation of the faces in the face regions is performed, that is, alignment processing of the faces in the face regions is performed to obtain the aligned face images.

[0157] Then, the aligned face images are input into a face feature extraction module, and the face features of the aligned face images are extracted through the face feature extraction module. Furthermore, the face features are input into a multi-layer perceptron (MLP) for mapping to obtain face features.

[0158] At the same time, the sample face images are input into an encoder, and the sample face images are compressed through the encoder to obtain latent space features; and the sample description text (a woman wearing an orange headscarf, smiling) is input into a text encoder, and the sample description text is subjected to feature extraction through the text encoder to obtain text features.

[0159] Furthermore, cross-attention calculation can be performed on the face features, latent space features, and text features to obtain fused features. The fused features are input into a denoising network for denoising processing to obtain denoised features. Finally, the denoised features are input into a decoder, and the decoder restores the denoised features into an image to obtain a generated image. The parameters of the preset diffusion model are adjusted according to the face difference between the generated image and the sample face image until the face difference between the generated image and the sample face image is less than a preset difference, and the training is completed to obtain the target diffusion model.

[0160] In the embodiment of the present application, for the trained diffusion model generated according to the above steps, only the face image of a specified face ID and a description prompt text need to be input. During the inference process, the trained diffusion model can generate a target face image with the specified face ID and conforming to the description prompt text according to the face image of the specified face ID and the description prompt text, without introducing an additional control network, which can improve the generation efficiency of face images.

[0161] To facilitate better implementation of the image generation method provided in the embodiment of the present application, the embodiment of the present application also provides an image generation device based on the above image generation method. The meanings of the nouns are the same as those in the above image generation method, and the specific implementation details can refer to the description in the method embodiment.

[0162] Please refer to Figure 5 , Figure 5 which is a structural block diagram of an image generation device provided in the embodiment of the present application. The device includes:

[0163] The first acquisition unit 301 is configured to acquire an initial face image and a description text;

[0164] The first extraction unit 302 is configured to extract the face features of the initial face image and the text features of the description text;

[0165] The first processing unit 303 is configured to perform compression processing on the initial face image to obtain the latent space features of the initial face image;

[0166] The second processing unit 304 is configured to perform fusion processing on the face features, the text features, and the latent space features to obtain the fused features;

[0167] The generation unit 305 is configured to generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0168] In some embodiments, the generation unit 305 may include:

[0169] The first processing subunit is configured to perform denoising processing on the fused features to obtain the denoised features;

[0170] The first generation subunit is configured to perform image generation based on the denoised features to obtain the target face image.

[0171] In some embodiments, the first extraction unit 302 may include:

[0172] The detection subunit is configured to perform face detection on the initial face image to determine the face key points in the initial face image;

[0173] The second processing subunit is configured to perform alignment processing on the initial face image based on the face key points to obtain the aligned face image;

[0174] The first extraction subunit is configured to perform feature extraction on the aligned face image to obtain the face features.

[0175] In some embodiments, the first processing unit 303 may include:

[0176] The mapping subunit is configured to input the initial face image into an encoder, and map the initial face image to a low-dimensional latent variable through the encoder to obtain the latent space features.

[0177] In some embodiments, the second processing unit 304 may include:

[0178] The input subunit is configured to input the face features, the text features, and the latent space features into a target diffusion model;

[0179] A third processing subunit, configured to process the face feature, the text feature, and the latent space feature through a cross-attention module of the target diffusion model to obtain the fused feature.

[0180] In some embodiments, the third processing subunit may specifically be configured to:

[0181] Perform cross-attention calculation on the face feature and the text feature through the cross-attention module to obtain a first feature;

[0182] Perform cross-attention calculation on the text feature and the latent space feature through the cross-attention module to obtain a second feature;

[0183] Perform cross-attention calculation on the image feature and the latent space feature through the cross-attention module to obtain a third feature;

[0184] Fuse based on the first feature, the second feature, and the third feature to obtain the fused feature.

[0185] In some embodiments, the apparatus may further include:

[0186] A second acquisition unit, configured to acquire a sample set, where the sample set includes a plurality of sample pairs, and each sample pair includes a sample face image and a sample description text corresponding to the sample face image;

[0187] A construction unit, configured to construct the target diffusion model based on the sample face image and the sample description text in the sample pair.

[0188] In some embodiments, the construction unit may include:

[0189] A second extraction subunit, configured to extract features from the sample face image to obtain a sample face feature corresponding to the sample face image;

[0190] A fourth processing subunit, configured to perform compression processing on the sample face image to obtain a sample latent space feature corresponding to the sample face image;

[0191] A second extraction subunit, configured to extract features from the sample description text to obtain a sample text feature corresponding to the sample description text;

[0192] A training subunit, configured to train a preset diffusion model based on the sample face feature, the sample latent space feature, and the sample text feature to obtain the target diffusion model.

[0193] In some embodiments, the apparatus may further include:

[0194] A third acquisition unit, configured to acquire time information, perform feature extraction on the time information, and obtain time features corresponding to the time information;

[0195] A third processing unit, configured to splice the sample face features and the time features to obtain spliced features.

[0196] In some embodiments, the training subunit may specifically be configured to:

[0197] Train a preset diffusion model based on the spliced features, the sample latent space features, and the sample text features to obtain the target diffusion model.

[0198] In some embodiments, the training subunit may specifically be configured to:

[0199] Process the sample face features, the sample text features, and the sample latent space features through the cross-attention module of the preset diffusion model to obtain fused sample features;

[0200] Generate a generated face image through the image generation module of the preset diffusion model based on the fused sample features;

[0201] Determine the difference information between the sample face image and the generated face image;

[0202] Adjust the preset diffusion model based on the difference information and a preset loss function to obtain the target diffusion network.

[0203] An embodiment of the present application discloses an image generation device. The initial face image and the description text are acquired by the first acquisition unit 301; the face features of the initial face image are extracted by the first extraction unit 302, and the text features of the description text are extracted; the initial face image is compressed by the first processing unit 303 to obtain the latent space features of the initial face image; the second processing unit 304 performs fusion processing on the face features, the text features, and the latent space features to obtain fused features; the generation unit 305 generates a target face image corresponding to the initial face image that conforms to the description text based on the fused features. In this way, the image generation accuracy of the image generation model can be improved.

[0204] Correspondingly, an embodiment of the present application further provides a computer device, and the computer device may be a server. As Figure 6 shown, Figure 6Schematic diagram of the structure of the computer device provided by the embodiment of the present application. The computer device 500 includes a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, and a computer program stored in the memory 502 and executable on the processor. Among them, the processor 501 is electrically connected to the memory 502. Those skilled in the art can understand that the structure of the computer device shown in the figure does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or arrange different components.

[0205] The processor 501 is the control center of the computer device 500, connecting various parts of the entire computer device 500 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 502, and calling data stored in the memory 502, it executes various functions of the computer device 500 and processes data, thereby monitoring the entire computer device 500.

[0206] In the embodiment of the present application, the processor 501 in the computer device 500 will load the instructions corresponding to the processes of one or more application programs into the memory 502 according to the following steps, and the processor 501 will run the application programs stored in the memory 502 to implement various functions:

[0207] Obtain an initial face image and a description text; extract the face features of the initial face image and the text features of the description text; perform compression processing on the initial face image to obtain the latent space features of the initial face image; perform fusion processing on the face features, text features, and latent space features to obtain the fused features; generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0208] This embodiment obtains an initial face image and a description text, extracts the face features of the initial face image and the text features of the description text, then performs compression processing on the initial face image to obtain the latent space features of the initial face image. Further, perform fusion processing on the face features, text features, and latent space features to obtain the fused features; generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features. In this way, the image generation accuracy of the image generation model can be improved.

[0209] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0210] Optionally, as Figure 6As shown, the computer device 500 further includes: a touch display screen 503, a radio frequency circuit 504, an audio circuit 505, an input unit 506, and a power supply 507. Among them, the processor 501 is electrically connected to the touch display screen 503, the radio frequency circuit 504, the audio circuit 505, the input unit 506, and the power supply 507 respectively. Those skilled in the art can understand that Figure 6 the computer device structure shown in

[0211] does not limit the computer device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. The touch display screen 503 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 503 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as the user using a finger, a stylus, or any suitable object or accessory to operate on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 501, and can receive and execute the command sent by the processor 501. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits it to the processor 501 to determine the type of touch event. Subsequently, the processor 501 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 503 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 503 can also be used as a part of the input unit 506 to implement the input function.

[0212] The radio frequency circuit 504 can be used to receive and transmit radio frequency signals to establish wireless communication with a network device or other computer devices, and receive and transmit signals between the network device or other computer devices.

[0213] The audio circuit 505 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 505 can convert the received audio data into an electrical signal and transmit it to the speaker, which converts it into a sound signal for output. On the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 505, converted into audio data, and then the audio data is output to the processor 501 for processing. After that, it is sent through the radio frequency circuit 504 to, for example, another computer device, or the audio data is output to the memory 502 for further processing. The audio circuit 505 may also include an earphone jack to provide communication between the peripheral earphone and the computer device.

[0214] The input unit 506 can be used to receive input digital, character information or user characteristic information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0215] The power supply 507 is used to supply power to each component of the computer device 500. Optionally, the power supply 507 can be logically connected to the processor 501 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 507 may also include one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, and any other components.

[0216] Although Figure 6 not shown in the figure, the computer device 500 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.

[0217] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0218] As can be seen from the above, the computer device provided in this embodiment can acquire an initial face image and a description text; extract the face features of the initial face image and the text features of the description text; perform compression processing on the initial face image to obtain the latent space features of the initial face image; perform fusion processing on the face features, text features, and latent space features to obtain the fused features; and generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0219] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0220] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores multiple computer programs that can be loaded by a processor to execute the steps in any of the image generation methods provided by the embodiments of the present application. For example, the computer program can execute the following steps:

[0221] Obtain an initial face image and a description text;

[0222] Extract the face features of the initial face image and the text features of the description text;

[0223] Perform compression processing on the initial face image to obtain the latent space features of the initial face image;

[0224] Perform fusion processing on the face features, text features, and latent space features to obtain the fused features;

[0225] Generate a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

[0226] In this embodiment, by obtaining an initial face image and a description text, extracting the face features of the initial face image, and extracting the text features of the description text, then performing compression processing on the initial face image to obtain the latent space features of the initial face image, and further performing fusion processing on the face features, text features, and latent space features to obtain the fused features; generating a target face image corresponding to the initial face image that conforms to the description text based on the fused features. In this way, the image generation accuracy of the image generation model can be improved.

[0227] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0228] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, optical disc, etc.

[0229] Since the computer programs stored in the storage medium can execute the steps in any of the image generation methods provided by the embodiments of the present application, the beneficial effects that can be achieved by any of the image generation methods provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments, which will not be elaborated here.

[0230] The above has introduced in detail an image generation method, device, storage medium, and computer device provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An image generation method, characterized in that, The method includes: Obtaining an initial face image and a description text; Extracting face features of the initial face image and text features of the description text; Performing compression processing on the initial face image to obtain latent space features of the initial face image; Performing fusion processing on the face features, the text features, and the latent space features to obtain fused features; Based on the fused features, generating a target face image corresponding to the initial face image that conforms to the description text.

2. The method according to claim 1, wherein The generating, based on the fused features, a target face image corresponding to the initial face image that conforms to the description text includes: Performing denoising processing on the fused features to obtain denoised features; Based on the denoised features, performing image generation to obtain the target face image.

3. The method according to claim 1, characterized in that, The extracting face features of the initial face image includes: Performing face detection on the initial face image to determine face key points in the initial face image; Based on the face key points, performing alignment processing on the initial face image to obtain an aligned face image; Performing feature extraction on the aligned face image to obtain the face features.

4. The method according to claim 1, wherein The performing compression processing on the initial face image to obtain latent space features of the initial face image includes: Inputting the initial face image into an encoder, and mapping the initial face image to a low-dimensional latent variable through the encoder to obtain the latent space features.

5. The method according to claim 1, wherein The performing fusion processing on the face features, the text features, and the latent space features to obtain fused features includes: Inputting the face features, the text features, and the latent space features into a target diffusion model; Processing the face features, the text features, and the latent space features through a cross-attention module of the target diffusion model to obtain the fused features.

6. The method according to claim 5, wherein The processing the face features, the text features, and the latent space features through a cross-attention module of the target diffusion model to obtain the fused features includes: Performing cross-attention calculation on the face features and the text features through the cross-attention module to obtain a first feature; Performing cross-attention calculation on the text features and the latent space features through the cross-attention module to obtain a second feature; Performing cross-attention calculation on the image features and the latent space features through the cross-attention module to obtain a third feature; Based on the first feature, the second feature, and the third feature, performing fusion to obtain the fused features.

7. The method according to claim 5, wherein Before inputting the face features, the text features, and the latent space features into the target diffusion model, the method further includes: Obtaining a sample set, where the sample set includes multiple sample pairs, and each sample pair includes a sample face image and a sample description text corresponding to the sample face image; Based on the sample face image and the sample description text in the sample pair, constructing the target diffusion model.

8. The method according to claim 7, wherein The constructing the target diffusion model based on the sample face image and the sample description text in the sample pair includes: Extract features from the sample face image to obtain the sample face features corresponding to the sample face image; Perform compression processing on the sample face image to obtain the sample latent space features corresponding to the sample face image; Extract features from the sample description text to obtain the sample text features corresponding to the sample description text; Train a preset diffusion model based on the sample face features, the sample latent space features, and the sample text features to obtain the target diffusion model.

9. The method according to claim 8, wherein The method further includes: Obtain time information and extract features from the time information to obtain the time features corresponding to the time information; Perform splicing processing on the sample face features and the time features to obtain the spliced features; The training of the preset diffusion model based on the sample face features, the sample latent space features, and the sample text features to obtain the target diffusion model includes: Train a preset diffusion model based on the spliced features, the sample latent space features, and the sample text features to obtain the target diffusion model.

10. The method according to claim 8, characterized in that The training of the preset diffusion model based on the sample face features, the sample latent space features, and the sample text features to obtain the target diffusion model includes: Process the sample face features, the sample text features, and the sample latent space features through the cross-attention module of the preset diffusion model to obtain the fused sample features; Generate a generated face image based on the fused sample features through the image generation module of the preset diffusion model; Determine the difference information between the sample face image and the generated face image; Adjust the preset diffusion model based on the difference information and a preset loss function to obtain the target diffusion network.

11. An image generation device, characterized in that, The device includes: A first acquisition unit for acquiring an initial face image and a description text; A first extraction unit for extracting the face features of the initial face image and extracting the text features of the description text; A first processing unit for performing compression processing on the initial face image to obtain the latent space features of the initial face image; A second processing unit for performing fusion processing on the face features, the text features, and the latent space features to obtain the fused features; A generation unit for generating a target face image corresponding to the initial face image that conforms to the description text based on the fused features.

12. A computer device, including a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the program, it implements the image generation method according to any one of claims 1 to 10.

13. A storage medium, characterized in that, The storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute the image generation method according to any one of claims 1 to 10.