Picture processing method and apparatus, device, and storage medium

By generating conditional embedding information and using cross-attention maps, the problem of distortion in the prior art in the art is solved, and similarity and fidelity are improved.

WO2025102946A1PCT designated stage expired Publication Date: 2025-05-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/117932
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-09-10
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

The prior art When editing portrait pictures based on description text, the similarity between the characters and the original portrait pictures is low, resulting in distortion of the characters, especially when the editing is complicated.

Method used

By generating conditional embedding information, combining the features of the original portrait picture and the description text, and using cross-attention maps to generate expression content of the character's expression features in the facial area, editing and processing portrait pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024117932_22052025_PF_FP_ABST
    Figure CN2024117932_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and discloses a picture processing method and apparatus, a device, and a storage medium. The method comprises: generating conditional embedding information on the basis of an original portrait picture and description text; on the basis of the original portrait picture, the conditional embedding information, and a facial mask map, generating a cross attention map corresponding to the original portrait picture; and on the basis of the conditional embedding information and the cross attention map, performing editing to obtain a processed portrait picture corresponding to the original portrait picture. The method can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. According to the method, a facial feature of a person in the original portrait picture and a person expression feature in the description text are fused, thereby ensuring a high similarity between the person in the generated processed portrait picture and the person in the original portrait picture and significantly improving the person fidelity in an editing process.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device, equipment and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 15, 2023, with application number 202311524401.2 and application name “Image processing method, device, equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of image processing technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Art

[0003] Currently, there is a demand for editing the original portrait image according to the description text, that is, editing the environment, character expression, etc. in the original portrait image according to the content of the description text.

[0004] In the related art, a generative model guided by descriptive text is trained, and the trained generative model is used to edit the original portrait image under the guidance of the descriptive text to generate a processed portrait image.

[0005] When using related technologies to edit the original portrait image, since the editing only focuses on the content of the descriptive text, the problem of low similarity between the person in the generated processed portrait image and the person in the original portrait image occurs, that is, the person is distorted. The more complex the editing, the more serious the distortion of the person.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide a method, apparatus, device, and storage medium for image processing. The technical solution is as follows:

[0008] According to one aspect of an embodiment of the present application, a method for image processing is provided, the method being executed by a computer device, the method comprising:

[0009] Generate conditional embedding information based on the original portrait image and the description text, the conditional embedding information including features of the original portrait image and features of the description text; wherein the original portrait image is an image containing facial features of a person, and the description text is text used to describe facial features of the person;

[0010] Generate a cross-attention map corresponding to the original portrait image according to the original portrait image, the conditional embedding information, and the facial mask map, wherein the facial mask map is used to distinguish a facial region from other regions in the original portrait image, and the cross-attention map is used to generate expression content corresponding to the character expression feature in the facial region;

[0011] Based on the conditional embedding information and the cross-attention map, a processed portrait image corresponding to the original portrait image is edited; wherein the processed portrait image retains the facial features of the original portrait image and includes the facial expression features of the character described in the description text.

[0012] According to one aspect of the present application, a method for training an image processing model is provided. The method is performed by a computer device and includes:

[0013] Obtain at least one training sample, each of the training samples comprising a set of corresponding sample portrait images and sample description texts; wherein the sample portrait images are images containing facial features of a person, and the sample description texts are texts used to describe facial features of the person in the sample portrait images;

[0014] Based on the sample portrait image and the sample description text, generating sample conditional embedding information through the image processing model, the sample conditional embedding information including features of the sample portrait image and features of the sample description text;

[0015] According to the sample portrait image, the sample conditional embedding information and the sample facial mask image, a cross-attention map corresponding to the sample portrait image is generated by the image processing model, wherein the sample facial mask image is used to distinguish the facial region and other regions other than the facial region in the sample portrait image, and the cross-attention map is used to generate expression content corresponding to the character expression feature in the facial region;

[0016] Based on the sample conditional embedding information and the cross attention map, obtaining a processed sample portrait picture corresponding to the sample portrait picture through the picture processing model editing;

[0017] Based on the sample portrait picture and the processed sample portrait picture, the parameters of the picture processing model are adjusted to obtain a trained picture processing model.

[0018] According to one aspect of an embodiment of the present application, a picture processing apparatus is provided, the apparatus being deployed on a computer device, the apparatus comprising:

[0019] An embedding module, configured to generate conditional embedding information based on an original portrait image and a description text, wherein the conditional embedding information includes features of the original portrait image and features of the description text; wherein the original portrait image is an image containing facial features of a person, and the description text is text used to describe facial features of the person;

[0020] An attention module is configured to generate a cross-attention map corresponding to the original portrait image based on the original portrait image, the conditional embedding information, and a facial mask map, wherein the facial mask map is configured to distinguish a facial region from other regions in the original portrait image, and the cross-attention map is configured to generate expression content corresponding to the character's expression features in the facial region;

[0021] An editing module is used to edit, based on the conditional embedding information and the cross-attention map, a processed portrait image corresponding to the original portrait image; wherein the processed portrait image retains the facial features of the original portrait image and includes the facial expression features of the character described in the description text.

[0022] According to one aspect of an embodiment of the present application, a training device for an image processing model is provided. The device is deployed on a computer device and includes:

[0023] A sample acquisition module is used to acquire at least one training sample, each of which includes a set of corresponding sample portrait images and sample description texts; wherein the sample portrait images are images containing facial features of a person, and the sample description texts are texts used to describe the facial features of the person in the sample portrait images;

[0024] a sample embedding module, configured to generate sample conditional embedding information based on the sample portrait image and the sample description text through the image processing model, wherein the sample conditional embedding information includes features of the sample portrait image and features of the sample description text;

[0025] a sample attention module, configured to generate a cross-attention map corresponding to the sample portrait image through the image processing model based on the sample portrait image, the sample conditional embedding information, and the sample facial mask image, wherein the sample facial mask image is used to distinguish the facial region from other regions in the sample portrait image, and the cross-attention map is used to generate expression content corresponding to the character expression feature in the facial region;

[0026] a sample editing module, configured to obtain a processed sample portrait image corresponding to the sample portrait image by editing the image processing model based on the sample conditional embedding information and the cross-attention map;

[0027] A model parameter adjustment module is used to adjust the parameters of the image processing model based on the sample portrait image and the processed sample portrait image to obtain a trained image processing model.

[0028] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned image processing method, or to implement the above-mentioned image processing model training method.

[0029] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned image processing method, or to implement the above-mentioned image processing model training method.

[0030] According to one aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the aforementioned image processing method or implement the aforementioned image processing model training method.

[0031] The technical solutions provided by the embodiments of the present application include at least the following beneficial effects:

[0032] Conditional embedding information is generated based on an original portrait image and a description text. The original portrait image is an image containing a person's facial features, and the description text is text describing the person's facial expressions. This allows the conditional embedding information to combine the facial features of the person in the original portrait image with the facial expressions of the person in the description text. A cross-attention map corresponding to the original portrait image is then generated based on the conditional embedding information, the original portrait image, and a facial mask. The facial mask is used to distinguish the facial region from other regions in the original portrait image, and the cross-attention map is used to generate facial expressions corresponding to the facial features of the person in the facial region. Thus, a processed portrait image can be edited based on the conditional embedding information and the cross-attention map. Because the conditional embedding information combines the facial features of the person in the original portrait image and the facial expressions of the description text, when the cross-attention map indicates that the facial expressions corresponding to the facial features of the person in the facial region should be generated, the facial features of the person in the original portrait image and the facial expressions of the description text can be combined to edit the person's facial expressions. This ensures that the person in the processed portrait image is highly similar to the person in the original portrait image, significantly improving the fidelity of the person during the editing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG1 is a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application;

[0034] FIG2 is a flowchart of an image processing method provided by an embodiment of the present application;

[0035] FIG3 is a schematic diagram of a cross-attention map generation process provided by one embodiment of the present application;

[0036] FIG4 is a schematic diagram of a cross-attention graph provided by one embodiment of the present application;

[0037] FIG5 is a flowchart of a method for training an image processing model provided by one embodiment of the present application;

[0038] FIG6 is a flowchart of a training process of an image processing model provided by one embodiment of the present application;

[0039] FIG7 is a comparison diagram of generated and processed portrait images provided by one embodiment of the present application;

[0040] FIG8 is a comparison diagram of emotion editing provided by an embodiment of the present application;

[0041] FIG9 is a line graph showing the effect of time steps on identity preservation provided by one embodiment of the present application;

[0042] FIG10 is a bar chart showing the effect of truncation hyperparameters on identity preservation and editing capabilities according to an embodiment of the present application;

[0043] FIG11 is a schematic diagram of test results of an image processing model provided by one embodiment of the present application;

[0044] FIG12 is a block diagram of an image processing device provided by an embodiment of the present application;

[0045] FIG13 is a block diagram of a training device for an image processing model provided by one embodiment of the present application;

[0046] FIG14 is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0048] The embodiments of the present application can automatically process images through artificial intelligence (AI) technology. With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0049] The solutions provided in the embodiments of this application involve technologies such as computer vision and machine learning based on artificial intelligence, which are specifically introduced and explained through the following embodiments.

[0050] Please refer to FIG1 , which shows a schematic diagram of a solution implementation environment provided by an embodiment of the present application. The solution implementation environment may include: a model training device 10 and a model use device 20 .

[0051] The model training device 10 is an electronic device with data calculation, processing and storage functions. The model training device 10 can be either a terminal device or a server. The model training device 10 can include, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality (AR) devices, virtual reality (VR) devices, and other electronic devices. The model training device 10 is used to train image processing models.

[0052] In the embodiment of the present application, the image processing model is a diffusion model. Optionally, the model training device 10 can use machine learning to train the image processing model to enable it to have better performance. Optionally, the training process of the image processing model is as follows (this is only a brief description, and the specific training process can be found in the following embodiments): obtaining training samples for training, the training samples including a sample portrait image 11 and a sample description text 12; compressing the sample portrait image 11 into a sample latent space image through a compression model; extracting the facial features of the person in the sample portrait image 11 by a facial recognition model; encoding the sample description text 12 by a text model to obtain text embedding information; splicing the facial features and the embedding information corresponding to the first participle to obtain updated embedding information; processing the updated embedding information through a multi-layer perceptron to obtain the final embedding information corresponding to the first participle; replacing the final embedding information corresponding to the first participle back to the text embedding information to obtain conditional embedding information; generating a processed portrait image 13 through a first neural network based on the sample latent space image and the conditional embedding information; calculating the total loss based on the sample portrait image 11 and the processed sample portrait image 13; adjusting the parameters of the image processing model based on the total loss to obtain a trained image processing model.

[0053] The model-using device 20 is an electronic device with data calculation, processing, and storage capabilities. The model-using device 20 can be either a terminal device or a server. The model-using device 20 can include, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, in-vehicle terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality devices, virtual reality devices, cloud technology platforms, intelligent robots, smart transportation terminal systems, and driving control systems. The model-using device 20 uses the trained image processing model to generate a processed portrait image based on the original portrait image and description text.

[0054] The model training device 10 and the model using device 20 can be two independent devices or the same device.

[0055] For example, the server mentioned above can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, but is not limited to these.

[0056] Please refer to Figure 2, which shows a flowchart of an image processing method provided by an embodiment of the present application. The execution subject of each step of the method can be a computer device, for example, the computer device can be the model using device 20 in the solution implementation environment shown in Figure 1. The method can include at least one of the following steps (210-230):

[0057] Step 210: Generate conditional embedding information based on the original portrait image and the description text, where the conditional embedding information includes features of the original portrait image and features of the description text; wherein the original portrait image is a picture containing facial features of a person, and the description text is text used to describe the facial features of the person.

[0058] The original portrait image includes at least one person, which may display complete or partial facial features. The original portrait image is composed of multiple pixels, each of which contains a portion of the image information. The original portrait image can use a variety of formats and representations, including but not limited to the following: Joint Photographic Experts Group (JPEG), Portable Network Graphics (PNG), Bitmap (BMP), Tagged Image File Format (Tagged Image File Format), Raw Image Format (RAW), Scalable Vector Graphics (SVG), and other image formats. Each pixel uses a numerical value to represent image attributes such as brightness, color, and position. For example, in a color image, three channels—red, green, and blue (RGB)—are used to represent color. By controlling the intensity values ​​of these three channels, the colors in the image can be precisely specified and adjusted. In black and white or grayscale images, a single channel is used to represent brightness or grayscale levels. The resolution of an image refers to the number of pixels in each direction of the image. A higher resolution indicates a clearer image.

[0059] The description text is a text used to guide the generation of corresponding picture content containing character expression features on the original portrait image. In some embodiments, the description text is a text used to describe environmental features and character expression features. In this way, the description text is a text used to guide the generation of corresponding picture content containing environmental features and character expression features on the original portrait image. Environmental features come from words with descriptions such as environmental descriptions, and character expression features come from words with descriptions such as expression descriptions and emotional descriptions. For example, the text "A sad woman under a tree" is a description text, in which "sad" is a word with emotional description, and "under a tree" is a word with environmental description. For example, in the text "There is a woman with a happy smile on her face on the grass", "on the grass" is a word with environmental description, "happy" is a word with emotional description, and "smile" is a word with emotional description.

[0060] The conditional embedding information is an embedding vector containing the features of the original portrait image and the features of the description text. The features of the original portrait image are extracted from the original portrait image and are used to quantitatively represent the content of the original portrait image. They can be represented by feature maps, color histograms, graphic data structures, convolutional feature maps, etc. The features of the description text are extracted from the description text and are used to quantitatively represent the content of the description text. They can be represented by one of the following structured data: sparse matrices, sparse vectors, dense vectors, etc. The features of the original portrait image and the features of the description text can use the same feature representation method or different feature representation methods. The specific feature representation method used is determined by relevant technical personnel and is not limited in this application.

[0061] In some embodiments, facial features of a person contained in an original portrait image are extracted through a facial recognition model; the description text is encoded through a text model to obtain text embedding information; and conditional embedding information is generated based on the facial features of the person and the text embedding information.

[0062] The facial recognition model is a neural network model that can extract facial features of a person from an original portrait image. The neural network model can be a Tface model or other neural network model for extracting facial features of a person, and this application does not limit this.

[0063] A text model is a deep learning model used to convert text data into an embedding vector with a numerical representation. In this application, the embedding vector is also called text embedding information. The deep learning model can be a contrastive language-image pretraining (CLIP) text model, or a recurrent neural network (RNN) or other text model, which is not limited in this application.

[0064] Text embedding information is an embedding vector with a numerical representation that describes the text transformation, which contains information such as text features, word embedding, and sentence embedding.

[0065] Through the above method, the facial features of the person and the text embedding information are integrated into conditional embedding information. When generating personalized content based on the embedding information of the first word segmentation, the impact on the facial features of the person in the original portrait image is reduced, which is conducive to the subsequent work of retaining the facial features of the person.

[0066] In some embodiments, the description text includes at least one segmentation word, and the text embedding information includes embedding information corresponding to each of the at least one segmentation word. In some embodiments, conditional embedding information can be generated based on facial features of a person and text embedding information in the following manner: the embedding information corresponding to the first segmentation word in at least one segmentation word is concatenated with the facial features of the person to obtain updated embedding information corresponding to the first segmentation word; wherein the first segmentation word is a segmentation word pre-set to represent the person; the updated embedding information corresponding to the first segmentation word is processed using a multi-layer perceptron to obtain final embedding information corresponding to the first segmentation word; wherein the final embedding information corresponding to the first segmentation word has the same dimension as the embedding information corresponding to the first segmentation word; and the embedding information corresponding to the first segmentation word in the text embedding information is replaced with the final embedding information corresponding to the first segmentation word to generate conditional embedding information.

[0067] A word segment is a text segment that represents semantic information. It can be a character, a word, or a phrase or sentence containing multiple words.

[0068] In some embodiments, one word segment corresponds to one embedded information. For example, if a description text has five word segments, the corresponding conditional embedded information includes five embedded information, and the embedded information corresponds to the word segment one by one.

[0069] In an embodiment of the present application, dimension may refer to the number of attributes possessed by the embedded information. For example, if the first word segment can describe its embedded information from the perspective of 40 attributes, then the dimension of its corresponding embedded information is 40. The final embedded information corresponding to the first word segment has the same dimension as the embedded information corresponding to the first word segment, that is, the number of attributes possessed by the final embedded information corresponding to the first word segment is the same as the number of attributes possessed by the embedded information corresponding to the first word segment. The attributes possessed by the embedded information may refer to various features or information possessed by the word segment, such as semantic information, grammatical information, contextual information, emotional information, etc.

[0070] In some embodiments, the dimension of the embedded information corresponding to the first word segmentation may be equal to or different from the dimension of the facial features of the person. In some embodiments, when the dimension of the embedded information corresponding to the first word segmentation is equal to the dimension of the facial features of the person, the embedded information corresponding to the first word segmentation is directly concatenated with the facial features of the person.

[0071] In some embodiments, when the dimension of the embedded information corresponding to the first word segmentation is not equal to the dimension of the facial features of the character, the dimension of the embedded information corresponding to the first word segmentation is converted to be equal to the dimension of the facial features of the character by dimensionality reduction or other means, and then spliced.

[0072] The first participle is a pre-set word used to characterize a person. Optionally, the first participle can be a personal pronoun referring to the person, or a word indicating the person's gender. In some embodiments, the first participle includes "man" and "woman". In some embodiments, the first participle includes "she", "he" and "person". The first participle used in this application is only used to illustrate this embodiment. Relevant technicians can set other first participles according to the application scenario, and this application does not limit this.

[0073] Multi-Layer Perceptron (MLP) is an artificial neural network (ANN) architecture used for machine learning tasks.

[0074] For example, assuming that the description text is "a sad woman under a tree", the text model is the CLIP text model, and the facial recognition model is the Tface model, the description text is processed by the CLIP text model to obtain text embedding information, which is a 77×768 matrix. The text embedding information includes the embedding information corresponding to each word in the description text. The matrix refers to 77 word segments and their corresponding 768 features. The number of word segments in the description text may be less than 77. In this case, the CLIP text model will automatically fill in meaningless values. The meaningless values ​​are pre-set words that do not participate in training. The first word corresponding to the description text is "woman", and the embedding corresponding to "woman" is The information is a 1×768 vector, and the facial features of the person extracted by the facial recognition model from the original portrait image are a 1×25088 vector. The embedding information corresponding to the first participle is spliced ​​with the facial features of the person to obtain an updated embedding information of 1×25856. The updated embedding information is processed by the multi-layer perceptron to obtain the final embedding information corresponding to the first participle, that is, the final embedding information corresponding to the first participle is a 1×768 vector. The first neural network needs to perform matrix operations during editing. Therefore, only when the final embedding information corresponding to the first participle is equal to the dimension of the embedding information corresponding to its original embedding information, the conditional embedding information is a matrix composed of multiple vectors with the same dimension.

[0075] By embedding the facial features of the person into the text embedding information in the above manner, the person's features can be better preserved when the original portrait image is subsequently personalized edited according to the description text.

[0076] In some embodiments, a facial region in an original portrait image is captured to obtain a facial image corresponding to the original portrait image; and facial features of the person are extracted from the facial image using a facial recognition model.

[0077] In some embodiments, the facial area in the original portrait image can be manually captured, or a neural network model can be used to capture the facial area in the original portrait image, which is not limited in this application.

[0078] In some embodiments, the intercepted facial image can be a regular shape or an irregular shape, which is not limited in this application. For example, the intercepted facial image is a minimum rectangular frame that includes the complete facial area.

[0079] In this way, the features extracted by the facial recognition model are only related to the person's face, which helps to preserve the person's facial features well when performing complex personalized editing later.

[0080] In step 220, a cross-attention map corresponding to the original portrait image is generated based on the original portrait image, the conditional embedding information, and the facial mask map. The facial mask map is used to distinguish the facial area in the original portrait image from other areas except the facial area. The cross-attention map is used to generate expression content corresponding to the character's expression features in the facial area.

[0081] The Cross-Attention Map is used to describe the degree of response between the embedded information corresponding to different word segmentations at different positions in the image.

[0082] In one possible implementation, the cross-attention map corresponding to the original portrait image, the conditional embedding information, and the facial mask image can be generated by compressing the original portrait image to obtain a latent space image; the latent space image is an image with a lower dimensionality than the original portrait image and retains the facial features of the person. The cross-attention map corresponding to the original portrait image is then generated based on the latent space image, the conditional embedding information, and the facial mask image.

[0083] The latent space image is a specific vector space obtained by mapping the original portrait image. The vectors in the latent space image are continuous and follow a probability distribution. For example, the vectors in the latent space image follow a multidimensional normal distribution.

[0084] It is understandable that in the embodiments of the present application, dimension may refer to the number of features in the data representation or data space. Specifically, the original portrait picture can be represented by a large number of features, such as texture, shape, color distribution, etc., so the original portrait picture can be regarded as a high-dimensional vector space, that is, the number of features representing the original portrait picture is very large. By mapping to the latent space to obtain the latent space picture, the high-dimensional data space is converted into a low-dimensional data space, that is, the latent space picture uses fewer features for representation. The latent space picture is a picture with a lower dimension than the original portrait picture. The dimension here refers to the number of features used to represent the picture, that is, the number of features of the latent space picture is less than the number of features of the original portrait picture, thereby removing redundant information and reducing noise interference, making the picture more compact and easy to process in the latent space, which provides convenience for subsequent processing.

[0085] In some embodiments, the original portrait image is compressed using a compression model to obtain a latent space image. The compression model is a generative deep learning model for learning and generating the latent space image of the original portrait image. The compression model can be a variational autoencoder (VAE).

[0086] By compressing the original portrait image into a latent space image in the above manner, it helps to reduce the storage and computational overhead in the processing of the original portrait image while retaining the facial features of the person in the original portrait image.

[0087] According to the original portrait image, the conditional embedding information and the facial mask image, a method for generating a cross-attention map corresponding to the original portrait image can be to perform a cross-attention operation based on the original portrait image and the conditional embedding information to obtain an initial cross-attention map corresponding to the original portrait image. The initial cross-attention map refers to a cross-attention map that does not mask other areas except the facial area. In this application, the cross-attention map refers to a cross-attention map that masks other areas except the facial area. Therefore, it is possible to continue to mask other areas except the facial area in the initial cross-attention map based on the facial mask image to obtain a cross-attention map corresponding to the original portrait image.

[0088] It should be noted that the embodiment of the present application does not limit the method for obtaining the initial cross-attention map corresponding to the original portrait picture. In one possible implementation, the original portrait picture and the conditional embedding information can be directly subjected to a cross-attention operation to obtain the initial cross-attention map. In another possible implementation, in order to reduce the storage and computational overhead in the processing of the original portrait picture, the original portrait picture can be first compressed to obtain a latent space picture, and then the latent space picture and the conditional embedding information can be subjected to a cross-attention operation to obtain the initial cross-attention map. That is to say, when generating the cross-attention map corresponding to the original portrait picture based on the latent space picture, the conditional embedding information and the facial mask map, the latent space picture and the conditional embedding information can be subjected to a cross-attention operation to obtain the initial cross-attention map, and then, based on the facial mask map, the other areas except the facial area in the initial cross-attention map can be shielded to obtain the cross-attention map corresponding to the original portrait picture.

[0089] In some embodiments, a cross attention operation is applied to the latent space image and the conditional embedding information to obtain an initial cross attention map. The formula for the cross attention operation is:

[0090] Among them, M represents the initial cross attention map, Q represents the latent space image, K is the conditional embedding information, K T is the transpose of K, and d is the dimension of the latent space image.

[0091] Exemplarily, as shown in Figure 3, it shows a schematic diagram of the generation process of the cross-attention map provided by an embodiment of the present application. The latent space image is used as Q in the cross-attention mechanism, and the conditional embedding information is used as K in the cross-attention mechanism. According to the above-mentioned cross-attention operation formula, the initial cross-attention map M is calculated, and then the initial cross-attention map M is multiplied by the facial mask map F to finally obtain the cross-attention map A.

[0092] The initial cross-attention map refers to a cross-attention map that does not mask other areas except the face area. In this application, the cross-attention map refers to a cross-attention map that masks other areas except the face area.

[0093] Optionally, one cross-attention map corresponds to the embedding information corresponding to one word segment. The embedding information corresponding to one word segment can correspond to at least one cross-attention map. The number of cross-attention maps that the embedding information corresponding to one word segment can correspond to is determined by the first neural network. For example, assuming that the first neural network has m cross-attention layers and the conditional embedding information corresponds to n word segments, there are a total of m×n cross-attention maps, where m and n are both positive integers.

[0094] In some embodiments, the first neural network includes a cross-attention layer, so the method of generating a cross-attention map corresponding to the original portrait image based on the latent space image, conditional embedding information and facial mask image can be to generate an initial cross-attention map corresponding to the original portrait image through the first neural network based on the latent space image and the conditional embedding information, that is, to use the cross-attention layer in the first neural network to perform a cross-attention operation on the latent space image and the conditional embedding information to obtain the initial cross-attention map; based on the facial mask image, other areas except the facial area in the initial cross-attention map are shielded to obtain the cross-attention map corresponding to the original portrait image.

[0095] Exemplarily, as shown in FIG4 , it shows a schematic diagram of a cross-attention map provided by an embodiment of the present application. The text description is “a bear”, and the corresponding participles are “one” and “bear”. The participle “one” corresponds to the cross-attention map (shown by mark 41 in FIG4 ), and “bear” corresponds to the cross-attention map (shown by mark 42 in FIG4 ).

[0096] A facial mask is a mask in which the facial area is the response area and the areas other than the facial area are the shielding areas. The mask is an image or matrix of the same size as the original portrait image, in which each pixel or element is used to mark a specific area in the original portrait image, namely the facial area and the areas other than the facial area in the original portrait image. There are many ways to represent the pixel values ​​in the mask, including but not limited to the following: binarization, multi-category, probability, and floating-point numbers. For example, the facial mask is represented by binarization, where 1 represents the facial area and 0 represents the areas other than the facial area.

[0097] Through the above method, when editing personalized expressions, the editing of the original portrait image only responds to the facial area, making the response of different segmentations to the expression editing in the facial area more accurate.

[0098] Step 230: Based on the conditional embedding information and the cross-attention map, edit to obtain a processed portrait image corresponding to the original portrait image; wherein the processed portrait image retains the facial features of the original portrait image and includes the facial expression features of the character described by the description text.

[0099] In some embodiments, when the description text is text used to describe environmental features and character expression features, the processed portrait image retains the facial features of the original portrait image and includes the environmental features and character expression features described by the description text.

[0100] It should be noted that if the implementation of step 220 is different, step 230 may also be implemented differently. If in step 220, the cross-attention map is directly obtained in the space of the original portrait image, then the implementation of step 230 can be to edit the original portrait image based on the conditional embedding information and the cross-attention map to obtain a processed portrait image corresponding to the original portrait image.

[0101] If, in step 220, a cross-attention map is obtained in the latent space, that is, step 220 generates a cross-attention map corresponding to the original portrait image based on the latent space image, the conditional embedding information, and the facial mask, step 230 can be implemented by editing the latent space image based on the conditional embedding information and the cross-attention map to generate a processed portrait image. Editing the latent space image helps reduce storage and computational overhead during the editing process because the latent space image has a lower dimension than the original portrait image, while preserving the facial features of the person in the original portrait image.

[0102] In some embodiments, based on the conditional embedding information and the cross-attention map, the latent space image is edited, and the way to generate the processed portrait image can be to add T times of noise to the latent space image through the first neural network to obtain a noise image, where T is a positive integer; according to the conditional embedding information and the cross-attention map, perform T times of noise prediction and denoising operations on the noise image through the first neural network to obtain a denoised image; and restore the denoised image to the dimension corresponding to the original portrait image to generate a processed portrait image.

[0103] The first neural network refers to a fully convolutional neural network for performing semantic segmentation tasks. The first neural network can be a U-Net or other fully convolutional neural network for performing semantic segmentation tasks. This application does not limit this.

[0104] In some embodiments, at any time of noise addition, a noise intensity is first randomly obtained, and noise of the noise intensity is randomly added to the latent space image through the first neural network. Noise intensity is a numerical value used to describe the noise level or intensity, indicating the degree of disturbance of the noise to the original portrait image. The higher the noise intensity, the greater the impact of the noise on the original portrait image. There are many ways to measure noise intensity, which may include but are not limited to the following methods: root mean square error, signal-to-noise ratio, standard deviation, and peak signal-to-noise ratio, etc., which are not limited in this application. In some embodiments, the denoised image is restored to the dimension corresponding to the original portrait image through the above-mentioned compression model to generate a processed portrait image.

[0105] It should be noted that the embodiments of the present application do not limit the manner in which T noise prediction and denoising operations are performed on a noisy image. It is sufficient that the number of noise predictions and denoising operations is the same. In some embodiments, performing T noise prediction and denoising operations on a noisy image may mean that each time a noise prediction is performed, a corresponding predicted noise is obtained, and then a denoising operation is performed using the predicted noise, and this cycle is repeated until the Tth noise prediction is completed, the corresponding predicted noise is obtained, and the predicted noise obtained from the Tth noise prediction is used to perform the Tth denoising operation, thereby obtaining the final denoised image.

[0106] In other embodiments, performing T noise prediction and denoising operations on a noisy image may mean performing noise prediction multiple times (for example, 2 times, 3 times, ..., T times) each time to obtain corresponding predicted noise, and then using the obtained predicted noise to perform denoising operations the same number of times as the noise prediction, and repeating this cycle until T noise predictions and T denoising operations are completed to obtain the final denoised image.

[0107] Based on this, in some embodiments, T noise prediction and denoising operations are performed on the noisy picture through the first neural network to obtain the denoised picture. The method can be to predict the noise added to the latent space picture for the nth time based on the conditional embedding information and the cross-attention map through the first neural network to obtain the nth predicted noise, where the initial value of n is 1 and n is a positive integer less than or equal to T; remove the nth predicted noise from the noisy picture to obtain an updated noisy picture; wherein, the initialized noisy picture is the noisy picture; when n is less than T, the value of n plus 1 is used as the updated n, and the step of predicting the noise added to the latent space picture for the nth time based on the conditional embedding information and the cross-attention map through the first neural network to obtain the nth predicted noise is started again; when n is equal to T, the updated noisy picture is determined to be the denoised picture.

[0108] For example, suppose x0 is the original portrait image, z is the latent space image, and z t is a noisy image, T is the total number of steps to add noise, C is the conditional embedding information, and for time step t, the denoised image can be obtained according to the following formula:

[0109] Among them, ∈ θ That is ∈ θ (z t ,t,C), represents the prediction noise predicted based on the noisy image, the current time step t and the conditional embedding information; δ is a hyperparameter that controls the inversion time; Indicates the cumulative coefficient.

[0110] It should be noted that the sizes of the above-mentioned noisy images, noisy images, denoised images and latent space images are equal.

[0111] Through the above method, personalized content can be generated on the original portrait image.

[0112] In some embodiments, the image processing method can be completed by an image processing model, which is a diffusion model for generating a processed portrait image based on a sample portrait image and a sample description text, and includes a first neural network, a text model, a facial recognition model, a multi-layer perceptron, and a compression model. For example, the original portrait image is compressed into a latent space image by the compression model; the facial features of the person in the original portrait image are extracted using the facial recognition model; the text model encodes the description text to obtain text embedding information; the facial features and the embedding information corresponding to the first participle are concatenated to obtain updated embedding information; the updated embedding information is processed by a multi-layer perceptron to obtain the final embedding information corresponding to the first participle; the final embedding information corresponding to the first participle is replaced with the text embedding information to obtain conditional embedding information; and based on the latent space image and the conditional embedding information, the corresponding processed portrait image is generated by the first neural network.

[0113] In summary, the technical solution provided in the embodiment of the present application generates conditional embedding information based on the original portrait image and the description text. The original portrait image is a picture containing facial features of a person, and the description text is a text used to describe the facial features of the person, so that the conditional embedding information integrates the facial features of the person in the original portrait image and the facial features of the person in the description text. Then, based on the above-mentioned conditional embedding information, the original portrait image and the facial mask map, a cross-attention map corresponding to the original portrait image is generated. The facial mask map is used to distinguish the facial area in the original portrait image from other areas other than the facial area, and the cross-attention map is used to generate facial expression content corresponding to the facial features of the person in the facial area. In this way, based on the conditional embedding information and the cross-attention map, the processed portrait image can be edited. Since the conditional embedding information integrates the facial features of the person in the original portrait image and the expression features of the person in the description text, when the cross-attention map indicates to generate expression content corresponding to the expression features of the person in the facial area, the facial features of the person in the original portrait image and the expression features of the person in the description text can be integrated to edit the person's expression, ensuring that the person in the generated processed portrait image is highly similar to the person in the original portrait image, and the person fidelity is significantly improved during the editing process.

[0114] The image processing model usage process has been described above. The following describes the training process of the image processing model. The image processing model usage process and the training process method steps correspond to each other. For details not shown in the training process corresponding to the embodiment, please refer to the corresponding part of the usage process embodiment.

[0115] Please refer to Figure 5, which shows a flowchart of a method for training an image processing model provided by one embodiment of the present application. The execution subject of each step of the method can be a computer device, for example, the computer device can be the model training device 10 in the solution implementation environment shown in Figure 1. The method can include at least one of the following steps (510-550):

[0116] Step 510, obtaining at least one training sample, each training sample including a set of corresponding sample portrait images and sample description texts; wherein the sample portrait images are images containing facial features of a person, and the sample description texts are texts used to describe the facial features of the person in the sample portrait images.

[0117] The "sample portrait image" and "sample description text" correspond to the "original portrait image" and "description text" in steps 210 to 240, respectively. The "sample description text" is used to describe the content of the "sample portrait image," but the content described by the "description text" may not be completely related to the content of the "original portrait image." For example, if the content of the "sample portrait image" is a woman laughing, and the corresponding "sample description text" is "a woman laughing," the content of the "sample description text" is consistent with the content of the "sample portrait image," that is, the "sample description text" describes the content of the "sample portrait image." For example, if the content of the "original portrait image" is a woman laughing, but the "description text" is "a woman crying," the content described by the "description text" is not completely related to the content of the "original portrait image." Otherwise, the usage of the "sample portrait image" and "sample description text" in this embodiment is the same as the usage of the "original portrait image" and "description text" in the embodiments corresponding to steps 210 to 240. It should be noted that, during the use and training of image processing models, except for the training process which requires calculating losses to adjust the parameters of the image processing model, the other processing processes are the same. Therefore, in the following text, "sample latent space image" corresponds to "latent space image" above. These two terms have the same function in the image processing model except for the different names. The different names are only used to distinguish whether they are in the use process or the training process. The same applies to other terms and are not listed here one by one.

[0118] The image processing model is a diffusion model used to generate processed portrait images based on sample portrait images and sample description text. It includes a first neural network, a text model, a facial recognition model, a multi-layer perceptron, and a compression model. The diffusion model is a deep generative model that generates high-quality images by adding noise and gradually increasing the diffusion process.

[0119] In some embodiments, the training samples may come from an open source dataset or a dataset manually collected and annotated.

[0120] In some embodiments, when the training samples are manually collected and annotated datasets, the sample portrait images of the training samples can be from an open source dataset, or can be manually acquired from the Internet, manually photographed, or collected. Sample description text corresponding to the sample portrait images can be generated by manually describing the content of the acquired sample portrait images, or by extracting the image content from the acquired sample portrait images using an image-to-text model to generate sample description text corresponding to the sample portrait images. The image-to-text model is a deep learning model used to generate descriptive text based on the content of the sample portrait images.

[0121] Obtaining high-quality training samples in the above manner is conducive to obtaining an image processing model with stronger generalization ability, robustness, and performance.

[0122] Step 520 : Based on the sample portrait image and the sample description text, generate sample conditional embedding information through the image processing model. The sample conditional embedding information includes features of the sample portrait image and features of the sample description text.

[0123] In some embodiments, facial features of a person contained in a sample portrait image are extracted through a facial recognition model; the sample description text is encoded through a text model included in an image processing model to obtain sample text embedding information; and sample conditional embedding information is generated based on the facial features of the person and the sample text embedding information.

[0124] In some embodiments, the sample description text includes at least one segmentation, and the sample text embedding information includes embedding information corresponding to each of the at least one segmentation; the embedding information corresponding to the first segmentation in the at least one segmentation is spliced ​​with the facial features of the character to obtain updated embedding information corresponding to the first segmentation; wherein, the first segmentation refers to a pre-set segmentation for representing the character; a multi-layer perceptron is used to process the updated embedding information corresponding to the first segmentation to obtain final embedding information corresponding to the first segmentation; wherein, the final embedding information corresponding to the first segmentation has the same dimension as the embedding information corresponding to the first segmentation; the embedding information corresponding to the first segmentation in the sample text embedding information is replaced with the final embedding information corresponding to the first segmentation to generate sample conditional embedding information.

[0125] Step 520 corresponds to step 210 in the above embodiment. They are similar in processing flow and effects, so they will not be described here in detail.

[0126] Step 530, based on the sample portrait image, the sample conditional embedding information and the sample facial mask image, a cross-attention map corresponding to the sample portrait image is generated through the image processing model. The sample facial mask image is used to distinguish the facial area and other areas except the facial area in the sample portrait image, and the cross-attention map is used to generate expression content corresponding to the character's expression features in the facial area.

[0127] Step 530 corresponds to step 220 in the above embodiment. They are similar in processing flow and effects, so they will not be described in detail here.

[0128] In one possible implementation, step 530 may be implemented by compressing the sample portrait image using an image processing model to obtain a sample latent space image; the sample latent space image is an image with a lower dimensionality than the sample portrait image and retains the facial features of the person. Based on the sample latent space image, the sample conditional embedding information, and the sample facial mask image, the image processing model generates a cross-attention map corresponding to the sample portrait image.

[0129] Similarly, the dimension here refers to the number of features of the image, that is, the number of features of the sample latent space image is smaller than the number of features of the sample portrait image, thereby removing redundant information and reducing noise interference, making the image more compact and easier to process in the latent space, which facilitates subsequent processing.

[0130] According to the sample portrait picture, the sample conditional embedding information and the sample facial mask picture, the method of generating the cross-attention map corresponding to the sample portrait picture through the image processing model can be to perform a cross-attention operation based on the sample portrait picture and the sample conditional embedding information to obtain the initial cross-attention map corresponding to the sample portrait picture. The initial cross-attention map corresponding to the sample portrait picture refers to a cross-attention map that does not mask other areas except the facial area. In this application, the cross-attention map corresponding to the sample portrait picture refers to a cross-attention map that masks other areas except the facial area. Therefore, it is possible to continue to mask other areas except the facial area in the initial cross-attention map corresponding to the sample portrait picture based on the sample facial mask picture to obtain the cross-attention map corresponding to the sample portrait picture.

[0131] It should be noted that the embodiment of the present application does not limit the method for obtaining the initial cross-attention map corresponding to the sample portrait image. In one possible implementation, the sample portrait image and the sample conditional embedding information can be directly subjected to a cross-attention operation to obtain the initial cross-attention map corresponding to the sample portrait image. In another possible implementation, in order to reduce the storage and calculation overhead during the processing of the sample portrait image, the sample portrait image can be first compressed to obtain a sample latent space image, and then the sample latent space image and the sample conditional embedding information can be subjected to a cross-attention operation to obtain the initial cross-attention map corresponding to the sample portrait image. That is, when the cross-attention map corresponding to the sample portrait image is generated by the image processing model based on the sample latent space image, the sample conditional embedding information and the sample facial mask image, the sample latent space image and the sample conditional embedding information can be subjected to a cross-attention operation to obtain the initial cross-attention map corresponding to the sample portrait image, and then, based on the sample facial mask image, the other areas except the facial area in the initial cross-attention map corresponding to the sample portrait image can be shielded to obtain the cross-attention map corresponding to the sample portrait image. It should be noted that the calculation formula of the cross-attention operation can be found in the introduction of step 220 and will not be repeated here.

[0132] In some embodiments, the first neural network includes a cross-attention layer, so the method of generating a cross-attention map corresponding to the sample portrait picture through the image processing model according to the sample latent space picture, the sample conditional embedding information and the sample facial mask picture can be to generate an initial cross-attention map corresponding to the sample portrait picture through the first neural network included in the image processing model according to the sample latent space picture and the conditional embedding information, that is, to use the cross-attention layer in the first neural network to perform a cross-attention operation on the sample latent space picture and the sample conditional embedding information to obtain the initial cross-attention map corresponding to the sample portrait picture; based on the sample facial mask picture, other areas except the facial area in the initial cross-attention map are shielded to obtain the cross-attention map corresponding to the sample portrait picture.

[0133] Step 540 , based on the sample conditional embedding information and the cross-attention map, obtain a processed sample portrait image corresponding to the sample portrait image through image processing model editing.

[0134] It should be noted that if step 530 is implemented differently, step 540 may also be implemented differently. If, in step 530, a cross-attention map is directly obtained in the space of the sample portrait image, then step 5400 may be implemented by editing the sample portrait image based on the sample conditional embedding information and the cross-attention map corresponding to the sample portrait image to obtain a processed sample portrait image corresponding to the sample portrait image.

[0135] If in step 530, a cross-attention map is obtained in the latent space, that is, step 530 generates a cross-attention map corresponding to the sample portrait picture through the image processing model based on the sample latent space picture, the sample conditional embedding information and the sample facial mask picture, then the implementation method of step 540 can be to edit the sample latent space picture through the image processing model based on the sample conditional embedding information and the cross-attention map to generate a processed sample portrait picture corresponding to the sample portrait picture.

[0136] In some embodiments, based on the sample conditional embedding information and the cross-attention map, the sample latent space image is edited through the image processing model to generate a processed sample portrait image corresponding to the sample portrait image. The method can be to add T times of noise to the sample latent space image through the first neural network included in the image processing model according to the conditional embedding information and the cross-attention map to obtain a sample noise image, where T is a positive integer; perform T times of noise prediction and denoising operations on the sample noise image through the first neural network to obtain a sample denoised image; and restore the sample denoised image to the dimension corresponding to the sample portrait image to generate a processed portrait image.

[0137] It should be noted that the embodiments of the present application do not limit the manner in which T noise prediction and denoising operations are performed on the sample noise image. It is sufficient that the number of noise predictions and denoising operations is the same. In some embodiments, performing T noise prediction and denoising operations on the sample noise image may mean that each time a noise prediction is performed, a corresponding predicted noise is obtained, and then a denoising operation is performed using the predicted noise, and this cycle is repeated until the Tth noise prediction is completed, the corresponding predicted noise is obtained, and the predicted noise obtained from the Tth noise prediction is used to perform the Tth denoising operation, thereby obtaining the final sample denoised image.

[0138] In other embodiments, performing T noise prediction and denoising operations on a sample noise image may mean performing noise prediction multiple times (for example, 2 times, 3 times, ..., T times) each time to obtain corresponding predicted noise, and then using the obtained predicted noise to perform denoising operations the same number of times as the noise prediction, and repeating this cycle until T noise predictions and T denoising operations are completed to obtain the final sample denoised image.

[0139] Based on this, in some embodiments, the first neural network performs T noise prediction and denoising operations on the sample noisy picture to obtain the sample denoised picture. The implementation method can be to predict the noise added to the sample latent space picture for the nth time according to the sample conditional embedding information and the cross-attention map through the first neural network to obtain the nth predicted noise, where the initial value of n is 1 and n is a positive integer less than or equal to T; remove the nth predicted noise from the sample noisy picture to obtain an updated sample noisy picture; wherein the initialized sample noisy picture is the sample noise picture; when n is less than T, the value of n plus 1 is used as the updated n, and the step of predicting the noise added to the sample latent space picture for the nth time according to the sample conditional embedding information and the cross-attention map through the first neural network to obtain the nth predicted noise is started again; when n is equal to T, the updated sample noisy picture is determined as the sample denoised picture.

[0140] In some embodiments, based on the noise added to the sample latent space image for the nth time and the nth predicted noise, a noise loss is determined, where the noise loss is used to measure the difference between the added noise and the predicted noise; and based on the noise loss, the parameters of the first neural network are adjusted.

[0141] For example, after obtaining the noise added to the sample latent space image for the tth time and the nth predicted noise, the noise loss is calculated by calculating the distance between the noise added to the sample latent space image for the tth time and the tth predicted noise. The noise loss function can be defined as:

[0142] Among them, ε is the variational self-encoder, Z is the latent space image, C is the conditional embedded information, ∈ θ The denoiser used to predict the noise θ, ∈ is a random sample drawn from a standard normal distribution, z t is the sample noisy image after denoising, z t is the sample noisy image after t-1 denoising, t represents the current time step, α t is a predefined coefficient sequence used to control the variance change during model training; the parameters of the first neural network are adjusted according to the noise loss; the probability distribution p(z t |z0) in closed form:

[0143] in, Indicates the cumulative coefficient.

[0144] By adjusting the parameters of the first neural network according to the noise loss in the above manner, the first neural network can be made to predict noise more accurately.

[0145] Step 550 : Based on the sample portrait image and the processed sample portrait image, adjust the parameters of the image processing model to obtain a trained image processing model.

[0146] In some embodiments, based on the sample portrait picture and the processed sample portrait picture, a total loss is determined, and the total loss includes noise loss, identity loss and facial position loss; wherein, the noise loss is used to measure the difference between the added noise and the predicted noise, the identity loss is used to measure the difference between the feature representation of the sample portrait picture and the processed sample portrait picture in the facial area, and the facial position loss is used to measure the control ability of the cross-attention map in the facial area; the parameters of the image processing model are adjusted according to the total loss to obtain the trained image processing model.

[0147] In some embodiments, gradient descent, random search, grid search, etc. can be used to adjust the parameters of the image processing model according to the total loss, which is not limited in this application.

[0148] In some embodiments, when adjusting the parameters of the image processing model, only the parameters of the first neural network and the multi-layer perceptron in the image processing model can be adjusted, and the other parts are not adjusted. The specific part of the parameters to be adjusted is selected by relevant technical personnel according to needs, and this application does not limit this.

[0149] The total loss of the image processing model is determined by noise loss, identity loss and face location loss.

[0150] For example, assuming that the image processing model completes one training at this time and obtains a processed sample portrait image corresponding to the sample portrait image, the total loss is calculated by estimating the difference between the sample portrait image and the corresponding processed sample portrait image. The total loss is defined as:

[0151] in, Refers to the total loss of this training image processing model, represents the face position loss, represents the noise loss, Indicates loss of identity.

[0152] In some embodiments, after the image processing model completes training on a batch of training samples, the parameters of the multilayer perceptron are adjusted based on the total loss. The batch of training samples may include one training sample or multiple training samples.

[0153] For example, when a batch of training samples includes multiple training samples, the description texts in the training samples are the same text, while the sample portraits in the training samples have differences in clarity, amount of noise, etc.

[0154] By comprehensively considering the losses in the three aspects in the above way, it is helpful to adjust the image processing model to suitable parameters during the training process, so that the image processing model has higher accuracy and generalization ability.

[0155] In some embodiments, noise loss is determined based on the noise added and the predicted noise during the editing process of the sample latent space image; identity loss is determined based on the sample portrait image and the processed sample portrait image; facial position loss is determined based on the sample facial mask image and the cross-attention map; and total loss is determined based on the noise loss, identity loss and facial position loss.

[0156] In the above process of editing the sample latent space image, the noise loss is calculated.

[0157] The cross-attention map is used to control the image processing model to only respond to the image area corresponding to the cross-attention map when editing the sample portrait image. The facial position loss is used to measure the accuracy of the position indicated by the cross-attention map in the sample portrait image.

[0158] Since the cross-attention map corresponding to the first participle and the cross-attention maps corresponding to other participles will be restricted to respond in the same area, the sum of their attention weights is 1, which means that the cross-attention map corresponding to the first participle and the cross-attention maps corresponding to other participles will compete when responding to editing. Therefore, other constraints need to be introduced to control the weights of the cross-attention map corresponding to the first participle and the cross-attention maps corresponding to other participles to be maintained within a certain threshold.

[0159] For example, assuming that the number of cross-attention layers in the first neural network is N, the number of word segmentations corresponding to the conditional embedding information is n, the face position loss is calculated by the absolute error loss, which is defined as:

[0160] in, is the cross attention map corresponding to the i-th word in the cross attention layer l, And i is a positive integer; is the cross attention map corresponding to the jth word in the cross attention layer l, And j is a positive integer, M refers to the facial mask map; β and γ represent the control strength of the cross-attention map on the first segmentation and expression segmentation, respectively; λ and μ are the positioning loss rates of identity and expression markers, which are 0.001 and 0.01, respectively; relu(*) means that if the value of formula * is less than 0, the value of relu(*) is 0, and if the value of formula * is greater than 0, the value of relu(*) is equal to the value of formula *.

[0161] In the above manner, by calculating the facial position loss, when personalizing the image, it is helpful for the image processing model to more accurately generate personalized content on the sample portrait image.

[0162] In some embodiments, facial features of a sample portrait image are extracted using a facial recognition model to obtain a first facial feature; facial features of the processed sample portrait image are extracted using a facial recognition model to obtain a second facial feature; and identity loss is determined based on the similarity between the first facial feature and the second facial feature.

[0163] There are many ways to calculate the similarity between the first facial feature and the second facial feature, including but not limited to the following methods: cosine similarity, mean squared error (MSE), structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), and feature vector similarity.

[0164] For example, the image processing model generates the corresponding processed sample portrait image f according to the sample portrait image f and the sample description text. The cosine similarity is used to calculate the sample portrait image f and the processed portrait image The similarity, identity loss function can be defined as:

[0165] Among them, cos sim(a,b) is the cosine similarity calculation formula, that is, calculating the cosine similarity of a and b; represents the facial recognition model, Indicates passing Features extracted from *.

[0166] To sum up, the technical solution provided by the embodiment of the present application obtains at least one high-quality training sample, generates sample conditional embedding information based on the sample portrait picture and sample description text in the training sample, and then generates a cross-attention map corresponding to the sample portrait picture through the image processing model according to the sample portrait picture, the sample conditional embedding information and the sample facial mask map. According to the sample conditional embedding information and the cross-attention map, the processed sample portrait picture is edited. Based on the sample portrait picture and the processed sample portrait picture, the parameters of the image processing model are adjusted to obtain a trained image processing model, so that the image processing model is more robust and generalized, and can complete more diverse personalized editing while also retaining the facial features of the characters in the sample portrait picture in the generated processed sample portrait picture.

[0167] The following describes a specific scenario of a training method using an image processing model:

[0168] Please refer to Figure 6, which shows a flowchart of the training process of the image processing model provided by an embodiment of the present application.

[0169] The image processing model includes the Tface face recognition model, the CLIP text editor, a variational autoencoder, a two-layer multi-layer perceptron (MLP), and a U-Net model. Figure 6 includes four modules: the main text embedding enhancement component, the latent space component, the emotion-aware cross-attention control component, and the identity preservation component.

[0170] The main text embedding enhancement part is performed by the Tface face recognition model, CLIP text editor and two-layer multi-layer perceptron to perform the following steps:

[0171] (1) Extract facial features of sample portrait images through the Tface face recognition model;

[0172] (2) extracting text embedding information of the sample description text through the CLIP text editor; determining the embedding information corresponding to the first word from the text embedding information, and concatenating the embedding information corresponding to the first word with the facial features to obtain updated embedding information;

[0173] (3) The updated embedding information is restored to the dimension of the sample embedding information through a two-layer multi-layer perceptron and replaced back into the text embedding information to obtain the sample conditional embedding information.

[0174] The emotion-aware cross-attention control part is performed by the U-Net model in the following steps:

[0175] (1) Generate a cross-attention map corresponding to the sample portrait image through the U-Net model based on the sample latent space image, sample conditional embedding information and sample facial mask image;

[0176] (2) Calculate the face position loss based on the sample face mask and the cross attention map.

[0177] The latent space part is composed of the variational autoencoder and the U-Net model, which performs the following steps:

[0178] (1) Compress the sample portrait image into a sample latent space image through a variational autoencoder;

[0179] (2) Add noise to the sample latent space image through the U-Net model to obtain the sample noise image;

[0180] (3) Using the U-Net model to predict and remove noise from the sample noisy image, a sample denoised image is obtained;

[0181] (4) Calculate the noise loss based on the sample noisy image and the sample denoised image, and adjust the parameters of the U-Net model according to the noise loss;

[0182] (5) The sample denoised image is restored to the dimension corresponding to the sample portrait image through the variational autoencoder to obtain the processed sample portrait image.

[0183] The identity preservation part is performed by the Tface face recognition model to perform the following steps:

[0184] (1) The Tface face recognition model samples the facial features of the portrait image and its corresponding processed sample portrait image, and calculates the identity loss through cosine similarity;

[0185] (2) Calculate the total loss based on noise loss, face position loss and identity loss, and adjust the parameters of the two-layer multilayer perceptron according to the total loss.

[0186] An example of an experiment using the method of the present application is described below:

[0187] (1) Dataset

[0188] The training dataset for the experiment was constructed using the CelebV-T dataset, which includes 70,000 videos and provides additional text descriptions. The first and last frames of each video in the CelebV-T dataset were extracted as sample portrait images, and the Recognize Anything model was used to generate sample description text corresponding to each sample portrait image. The resulting training dataset is a dataset that pairs sample portrait images and description text one by one. A frame was randomly selected from the middle part of each video in the CelebV-T dataset, and its facial area was cropped as a reference facial image. A pre-trained face parsing model was then used to generate corresponding facial masks for all sample portrait images.

[0189] (2) Experimental details

[0190] The Tface face recognition model was used to extract facial features from the sample portrait images. During the training process, only the U-Net model and the two-layer multi-layer perceptron were trained, and the CLIP text editor was not trained. The image processing model was iteratively trained 150,000 times on 6 NVIDIA V100 (6V100 is approximately 3A100) GPUs, with the learning rate set to 10 -5, the batch size is set to 2. 10% of the training samples are used for text conditioning training to maintain the pure text generation capability of the image processing model. In order to facilitate classifier-free guidance sampling, 10% of the samples are used for training without other constraints. During the training process, half of the samples are used to train the facial area of ​​the sample portrait image to improve the generation quality of personalized content in the facial area. The truncated cross attention control involves 7 emotional words, such as happiness, anger, sadness, etc. The truncated cross attention control here refers to the introduction of truncation parameters to control the intensity of the response of the emotional segmentation in the facial area.

[0191] (3) Evaluation indicators

[0192] Following the methodology used in FastComposer, the quality of generated sample portrait images after processing was evaluated based on facial feature preservation and CLIP-TI (Contrastive Language-Image Pretraining Text-Image) consistency. Facial feature preservation was determined by using MTCNN to detect faces in reference facial images and processed sample portrait images, and then FaceNet was used to compute pairwise identity similarity. The average CLIP-L / 14 image-text similarity was used to assess the similarity between the individual portraits and the sample description text. The efficiency of the image processing model was evaluated using the total fine-tuning and inference time. The evaluation also considered different numbers of Graphics Processing Units (GPUs). By default, all methods were run with standard hyperparameters, using Euler sampling with a step size of 50 and a classifier-free guidance degree of 5 for all methods.

[0193] (4) Processed image generation experiment

[0194] The single-subject evaluation method employed in FastComposer was used to evaluate the performance of the image processing model against existing methods, including DreamBooth, Textual-Inversion, Custom Diffusion, and SubjectDiffusion. Since Face0 does not provide open source code, only publicly available hardware resources were used for comparison. Stable Diffusion was set as the text-only baseline method. The training set used for testing consisted of 15 human subjects and 30 text passages. The sample description texts used for testing covered a variety of scenarios, such as changing image style, image decoration, and character action. Five sample portrait images were used for each subject to fine-tune the optimization-based method. For the one-shot learning-based method, a sample portrait image corresponding to each subject was randomly selected for testing. Table 1 shows the comparison results between our method (hereinafter referred to as "PortraitBooth") and a baseline method for single-subject image generation. PortraitBooth significantly outperformed the other baseline methods in preserving facial features. Please refer to Figure 7, which shows a comparison of sample portrait images generated and processed according to one embodiment of the present application. FastComposer corresponds to baseline method 1, and Subject-Diffusion corresponds to baseline method 2. PortraitBooth is slightly weaker than FastComposer in terms of sample description text consistency. This shortcoming may be because PortraitBooth tends to prioritize the fidelity of facial features, thereby giving up some attention to complex word segmentation.

[0195] Table 1 Comparison between PortraitBooth and other baseline methods for single-subject image generation

[0196] (5) Emotional editing

[0197] In terms of emotion editing, only FastComposer is used as the benchmark method. Please refer to Figure 8, which shows a comparison chart of emotion editing provided by an embodiment of the present application. FastComposer corresponds to benchmark method 1. It can be seen that PortraitBooth has diversity in emotion editing.

[0198] (6) Ablation experiment

[0199] 1) The impact of the main text embedding enhancement on the image processing model

[0200] To investigate the impact of facial features obtained from a pre-trained face recognition model and image encoder, an ablation experiment was conducted. The main text embedding enhancement component was removed from the image processing model, and the CLIP image encoder was used to train and extract facial features to enhance the text embedding information. As shown in Table 2(a), the experimental results clearly demonstrate that utilizing a facial feature extractor trained on a large-scale dataset is significantly more effective than training the image encoder.

[0201] 2) The impact of identity preservation on image processing models

[0202] Table 2(b) shows the results of an ablation experiment on the identity-preserving part, which is removed from the image processing model and trained under the same settings. The results show that the identity-preserving part is beneficial to the fidelity of facial features.

[0203] 3) The impact of emotion perception cross-attention control on image processing models

[0204] To allow the image processing model to focus on semantically relevant facial regions in the cross-attention module, a cross-attention map is used for control. As shown in Table 2(c), based on the results of the ablation experiment on the emotion-aware cross-attention control, the use of the cross-attention map significantly improves identity preservation and CLIP-TI consistency, enabling the image processing model to focus on specific regions. This results in a consistent and harmonious effect of the conditional embedding information, which is a combination of facial features and text embedding information, on the processed sample portrait images.

[0205] Table 2 Results of different ablation experiments

[0206] 4) The impact of hyperparameter t on image processing models

[0207] Figure 9 shows a line graph of the impact of time steps on identity preservation, as provided by one embodiment of the present application. As shown in Figure 9, as the time step (t) used to adjust identity loss increases, the image processing model's ability to preserve facial features increases, but its editing capabilities decrease. Therefore, when t is 250, the image processing model achieves a good balance between facial feature fidelity preservation and editing capabilities.

[0208] 5) Truncation of hyperparameters β and γ

[0209] Based on the analysis of Table 2(c), it is found that unrestricting local regions leads to better editing capabilities, but identity preservation deteriorates. To minimize the impact on identity preservation, the relationship between identity preservation and editing capabilities is only studied when the truncation hyperparameter β is involved. As shown in Figure 10, which shows a bar chart of the impact of the truncation hyperparameter on identity preservation and editing capabilities provided by one embodiment of the present application, when the truncation hyperparameter β is in the range of [0.8, 1], the impact of the truncation hyperparameter on identity preservation is not significant, but it has a certain impact on editing capabilities. However, when the truncation hyperparameter β is less than 0.8, identity preservation increases sharply because the conditional embedding information only has a significant impact on the facial region. As shown in Table 3, which shows the experimental results for facial masks and body masks, this confirms the conclusion that the conditional embedding information only has a significant impact on the facial region. Therefore, the truncation hyperparameter β was set to 0.8. Similarly, experiments were conducted with the truncation hyperparameter γ at 0.1 and 0.2. As shown in Table 4, the truncation hyperparameter β and the truncation hyperparameter γ have little difference in their impact on identity preservation, but a significant difference in editing ability. This is because facial edit responses include not only expressions but also features such as facial hair and accessories. Therefore, the truncation hyperparameter γ is set to 0.1.

[0210] Table 3 Different effects of mask images

[0211] Table 4 Effects of different combinations of truncation hyperparameters β and γ

[0212] As shown in Table 5, it shows the performance comparison of PortraitBooth and other benchmark methods, 1 means that the corresponding aspect of the method performs well, and 0 means that the corresponding aspect of the method performs poorly. PortraitBooth is the only method that can simultaneously solve the problems of high efficiency, robust identity preservation and diversified expression editing.

[0213] Table 5 Performance comparison of PortraitBooth and other baseline methods

[0214] As shown in FIG11 , it shows a schematic diagram of the test results of the image processing model provided in one embodiment of the present application, where the image processing model is obtained based on PortraitBooth training.

[0215] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0216] Please refer to Figure 12, which shows a block diagram of an image processing device provided by one embodiment of the present application. The device has the function of implementing the above-mentioned method example, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the model use device 10 described above, or it can be set in the model use device 10. As shown in Figure 12, the device 1200 can include an embedding module 1210, an attention module 1230, and an editing module 1240.

[0217] The embedding module 1210 is used to generate conditional embedding information based on the original portrait image and the description text, wherein the conditional embedding information includes the features of the original portrait image and the features of the description text; wherein the original portrait image is a picture containing facial features of a person, and the description text is text used to describe the facial features of the person.

[0218] The attention module 1230 is used to generate a cross-attention map corresponding to the original portrait image based on the original portrait image, the conditional embedding information and the facial mask map, wherein the facial mask map is used to distinguish the facial area in the original portrait image from other areas except the facial area, and the cross-attention map is used to generate the expression content corresponding to the character's expression features in the facial area.

[0219] The editing module 1240 is used to edit the processed portrait image corresponding to the original portrait image based on the conditional embedding information and the cross-attention map; wherein the processed portrait image retains the facial features of the original portrait image and includes the character expression features described by the description text.

[0220] In some embodiments, the device 1200 may further include a compression module 1220, as shown in the dotted box in Figure 12. The compression module 1220 is used to compress the original portrait image to obtain a latent space image; wherein the latent space image is a picture with a lower dimension than the original portrait image and retains the facial features of the character. The attention module 1230 is used to generate a cross-attention map corresponding to the original portrait image based on the latent space image, the conditional embedding information and the facial mask map. In some embodiments, the editing module 1240 includes: a denoising submodule, a denoising submodule and a restoration submodule (not shown in Figure 12).

[0221] The noise adding submodule is used to add T noises to the latent space image through the first neural network to obtain a noise image, where T is a positive integer.

[0222] A denoising submodule is configured to perform T noise prediction and denoising operations on the noisy image through the first neural network according to the conditional embedding information and the cross attention map to obtain a denoised image.

[0223] The restoration submodule is used to restore the denoised image to the dimension corresponding to the original portrait image to generate the processed portrait image.

[0224] In some embodiments, the denoising submodule is used to: predict the noise added to the latent space image for the nth time through the first neural network according to the conditional embedding information and the cross-attention map, and obtain the nth predicted noise, where the initial value of n is 1 and n is a positive integer less than or equal to T; remove the nth predicted noise from the noisy image to obtain an updated noisy image; wherein the initialized noisy image is the noisy image; when n is less than T, the value of n plus 1 is used as the updated n, and the step of predicting the noise added to the latent space image for the nth time according to the conditional embedding information and the cross-attention map through the first neural network to obtain the nth predicted noise is started again; when n is equal to T, the updated noisy image is determined as the denoised image.

[0225] In some embodiments, the embedding module 1210 includes an extraction submodule, an encoding submodule, and a generation submodule (not shown in FIG. 12 ).

[0226] The extraction submodule is used to extract the facial features of the person contained in the original portrait image through a facial recognition model.

[0227] The encoding submodule is used to encode the description text through a text model to obtain text embedding information.

[0228] A generating submodule is used to generate the conditional embedding information based on the facial features of the character and the text embedding information.

[0229] In some embodiments, the description text includes at least one segmentation word, and the text embedding information includes the embedding information corresponding to each of the at least one segmentation word; the generation submodule is used to: splice the embedding information corresponding to the first segmentation word in the at least one segmentation word with the facial features of the character to obtain the updated embedding information corresponding to the first segmentation word; wherein, the first segmentation word refers to a pre-set segmentation word used to represent the character; a multi-layer perceptron is used to process the updated embedding information corresponding to the first segmentation word to obtain the final embedding information corresponding to the first segmentation word; wherein, the final embedding information corresponding to the first segmentation word has the same dimension as the embedding information corresponding to the first segmentation word; the embedding information corresponding to the first segmentation word in the text embedding information is replaced with the final embedding information corresponding to the first segmentation word to generate the conditional embedding information.

[0230] In some embodiments, the extraction submodule is used to: capture the facial area in the original portrait image to obtain a facial image corresponding to the original portrait image; and extract the facial features of the person from the facial image using the facial recognition model.

[0231] In some embodiments, the attention module 1230 is used to: generate an initial cross-attention map corresponding to the original portrait image through a first neural network based on the latent space image and the conditional embedding information; and based on the facial mask image, mask other areas in the initial cross-attention map except the facial area to obtain a cross-attention map corresponding to the original portrait image.

[0232] In summary, the technical solution provided in the embodiment of the present application generates conditional embedding information based on the original portrait image and the description text. The original portrait image is a picture containing facial features of a person, and the description text is a text used to describe the facial features of the person, so that the conditional embedding information integrates the facial features of the person in the original portrait image and the facial features of the person in the description text. Then, based on the above-mentioned conditional embedding information, the original portrait image and the facial mask map, a cross-attention map corresponding to the original portrait image is generated. The facial mask map is used to distinguish the facial area in the original portrait image from other areas other than the facial area, and the cross-attention map is used to generate facial expression content corresponding to the facial features of the person in the facial area. In this way, based on the conditional embedding information and the cross-attention map, the processed portrait image can be edited. Since the conditional embedding information integrates the facial features of the person in the original portrait image and the expression features of the person in the description text, when the cross-attention map indicates to generate expression content corresponding to the expression features of the person in the facial area, the facial features of the person in the original portrait image and the expression features of the person in the description text can be integrated to edit the person's expression, ensuring that the person in the generated processed portrait image is highly similar to the person in the original portrait image, and the person fidelity is significantly improved during the editing process.

[0233] Please refer to Figure 13, which shows a block diagram of a training device for an image processing model provided by an embodiment of the present application. The device has the function of implementing the above-mentioned method example, and the function can be implemented by hardware or by hardware executing corresponding software. As shown in Figure 13, the device 1300 may include: a sample acquisition module 1310, a sample embedding module 1320, a sample attention module 1340, a sample editing module 1350 and a model parameter adjustment module 1360.

[0234] The sample acquisition module 1310 is used to obtain at least one training sample, each of which includes a set of corresponding sample portrait pictures and sample description texts; wherein the sample portrait pictures are pictures containing facial features of a person, and the sample description texts are texts used to describe the facial features of the person in the sample portrait pictures.

[0235] The sample embedding module 1320 is used to generate sample conditional embedding information based on the sample portrait image and the sample description text through the image processing model, where the sample conditional embedding information includes features of the sample portrait image and features of the sample description text.

[0236] The sample attention module 1340 is used to generate a cross-attention map corresponding to the sample portrait image through the image processing model based on the sample portrait image, the sample conditional embedding information and the sample facial mask map, wherein the sample facial mask map is used to distinguish the facial area in the sample portrait image from other areas except the facial area, and the cross-attention map is used to generate the expression content corresponding to the character's expression features in the facial area.

[0237] The sample editing module 1350 is used to edit the sample portrait image corresponding to the sample portrait image through the image processing model based on the sample conditional embedding information and the cross-attention map.

[0238] The model parameter adjustment module 1360 is used to adjust the parameters of the image processing model based on the sample portrait image and the processed sample portrait image to obtain a trained image processing model.

[0239] In some embodiments, the apparatus 1300 further includes a sample compression module 1330, as shown in the dotted box in FIG13 . The sample compression module 1330 is configured to compress the sample portrait image using the image processing model to obtain a sample latent space image; wherein the sample latent space image is an image having a lower dimension than the sample portrait image and retaining the facial features of the person. The sample attention module 1340 is configured to edit the sample latent space image using the image processing model based on the sample conditional embedding information and the cross-attention map to generate a processed sample portrait image corresponding to the sample portrait image.

[0240] In some embodiments, the sample editing module 1350 includes: a sample denoising submodule, a sample denoising submodule, and a sample restoration submodule (not shown in FIG. 13 ).

[0241] The sample noise adding submodule is used to add T noises to the sample latent space image through the first neural network included in the image processing model to obtain a sample noise image, where T is a positive integer.

[0242] A sample denoising submodule is used to perform T noise prediction and denoising operations on the sample noisy image through the first neural network according to the conditional embedding information and the cross attention map to obtain a sample denoised image.

[0243] The sample restoration submodule is used to restore the sample denoised image to the dimension corresponding to the sample portrait image to generate the processed sample portrait image.

[0244] In some embodiments, the sample denoising submodule includes a prediction unit, a denoising unit, a circulation unit, and a determination unit (not shown in FIG. 13 ).

[0245] A prediction unit is used to predict the noise added to the sample latent space image for the nth time through the first neural network based on the sample conditional embedding information and the cross-attention map, to obtain the nth predicted noise, where the initial value of n is 1 and n is a positive integer less than or equal to T.

[0246] The denoising unit is configured to remove the n-th predicted noise from the sample noisy picture to obtain an updated sample noisy picture; wherein the initialized sample noisy picture is the sample noisy picture.

[0247] A cyclic unit is used to, when n is less than T, use the value of n plus 1 as the updated n, and again start from the step of predicting the noise added to the sample latent space image for the nth time according to the sample condition embedding information and the cross attention map through the first neural network to obtain the nth predicted noise.

[0248] The determining unit is configured to, when n is equal to T, determine the updated sample noisy image as the sample denoised image.

[0249] In some embodiments, the prediction unit is further used to: determine a noise loss based on the noise added to the sample latent space image for the nth time and the nth predicted noise, wherein the noise loss is used to measure the difference between the added noise and the predicted noise; and adjust the parameters of the first neural network based on the noise loss.

[0250] In some embodiments, the sample embedding module 1320 includes a sample extraction submodule, a sample encoding submodule, and a sample generation submodule (not shown in FIG. 13 ).

[0251] The sample extraction submodule is used to extract the facial features of the person contained in the sample portrait image through a facial recognition model.

[0252] The sample encoding submodule is used to encode the sample description text through the text model included in the image processing model to obtain sample text embedding information.

[0253] The sample generation submodule is used to generate the sample conditional embedding information based on the facial features of the character and the sample text embedding information.

[0254] In some embodiments, the sample description text includes at least one segmentation word, and the sample text embedding information includes the embedding information corresponding to each of the at least one segmentation word; the sample generation submodule is used to: splice the embedding information corresponding to the first segmentation word in the at least one segmentation word with the facial features of the character to obtain the updated embedding information corresponding to the first segmentation word; wherein, the first segmentation word refers to a pre-set segmentation word used to represent the character; a multi-layer perceptron is used to process the updated embedding information corresponding to the first segmentation word to obtain the final embedding information corresponding to the first segmentation word; wherein, the final embedding information corresponding to the first segmentation word has the same dimension as the embedding information corresponding to the first segmentation word; the embedding information corresponding to the first segmentation word in the sample text embedding information is replaced with the final embedding information corresponding to the first segmentation word to generate the sample conditional embedding information.

[0255] In some embodiments, the sample extraction submodule is used to: capture the facial area in the sample portrait picture to obtain a facial picture corresponding to the sample portrait picture; and extract the facial features of the person from the facial picture using the facial recognition model.

[0256] In some embodiments, the sample compression module 1330 is used to: generate an initial cross-attention map corresponding to the sample portrait image through the first neural network included in the image processing model based on the sample latent space image and the conditional embedding information; and based on the sample facial mask image, mask other areas in the initial cross-attention map except the facial area to obtain a cross-attention map corresponding to the sample portrait image.

[0257] In some embodiments, the model parameter adjustment module 1360 includes a loss submodule and a parameter adjustment submodule (not shown in FIG. 13 ).

[0258] A loss submodule is used to determine a total loss based on the sample portrait image and the processed sample portrait image, where the total loss includes noise loss, identity loss, and facial position loss; wherein the noise loss is used to measure the difference between the added noise and the predicted noise, the identity loss is used to measure the difference between the feature representations of the sample portrait image and the processed sample portrait image in the facial area, and the facial position loss is used to measure the control ability of the cross-attention map in the facial area.

[0259] A parameter adjustment submodule is used to adjust the parameters of the image processing model according to the total loss to obtain the trained image processing model.

[0260] In some embodiments, the loss submodule includes a noise loss unit, an identity loss unit, a face position loss unit, and a total loss unit (not shown in FIG. 13 ).

[0261] The noise loss unit is configured to determine a noise loss based on the noise added and the predicted noise during the editing of the sample latent space image.

[0262] The identity loss unit is configured to determine the identity loss based on the sample portrait picture and the processed sample portrait picture.

[0263] A face position loss unit is configured to determine a face position loss based on the sample face mask map and the cross-attention map.

[0264] A total loss unit is configured to determine the total loss according to the noise loss, the identity loss, and the face position loss.

[0265] In some embodiments, the identity loss unit is used to: extract facial features of the sample portrait image through a facial recognition model to obtain a first facial feature; extract facial features of the processed sample portrait image through the facial recognition model to obtain a second facial feature; and determine the identity loss based on the similarity between the first facial feature and the second facial feature.

[0266] To sum up, the technical solution provided by the embodiment of the present application obtains at least one high-quality training sample, generates sample conditional embedding information based on the sample portrait picture and sample description text in the training sample, and then generates a cross-attention map corresponding to the sample portrait picture through the image processing model according to the sample portrait picture, the sample conditional embedding information and the sample facial mask map. According to the sample conditional embedding information and the cross-attention map, the processed sample portrait picture is edited. Based on the sample portrait picture and the processed sample portrait picture, the parameters of the image processing model are adjusted to obtain a trained image processing model, so that the image processing model is more robust and generalized, and can complete more diverse personalized editing while also retaining the facial features of the characters in the sample portrait picture in the generated processed sample portrait picture.

[0267] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0268] Please refer to Figure 14, which shows a block diagram of a computer device 1400 provided in one embodiment of the present application. The computer device 1400 may be the computer device 10 in the implementation environment shown in Figure 1, and is used to implement the image processing method or image processing model training method provided in the above embodiments. Specifically:

[0269] Typically, the computer device 1400 includes a processor 1410 and a memory 1420 .

[0270] The processor 1410 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 1410 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 1410 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1410 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1410 may also include an AI processor for processing computing operations related to machine learning.

[0271] The memory 1420 may include one or more computer-readable storage media, which may be non-transitory. The memory 1420 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1420 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned image processing method or image processing model training method.

[0272] Those skilled in the art will appreciate that the structure shown in FIG. 14 does not limit the computer device 1400 , and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0273] In an exemplary embodiment, a computer-readable storage medium is further provided, wherein a computer program is stored in the storage medium, and when the computer program is executed by the processor, the computer program is used to implement the above-mentioned image processing method or the training method of the image processing model. Optionally, the computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD) or an optical disc, etc. Among them, the random access memory may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM).

[0274] In an exemplary embodiment, a computer program product is further provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the above-described image processing method or image processing model training method.

[0275] It should be noted that the collection and processing of relevant data (such as images, etc.) in this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in practice, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0276] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.

[0277] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for processing an image, the method being executed by a computer device, the method comprising: Based on the original portrait image and the description text, conditional embedding information is generated, wherein the conditional embedding information includes features of the original portrait image and features of the description text; wherein the original portrait image is a picture containing facial features of a person, and the description text is a text used to describe facial features of the person; Generate a cross-attention map corresponding to the original portrait image according to the original portrait image, the conditional embedding information and the facial mask map, wherein the facial mask map is used to distinguish a facial region in the original portrait image from other regions except the facial region, and the cross-attention map is used to generate expression content corresponding to the character expression feature in the facial region; Based on the conditional embedding information and the cross-attention map, a processed portrait image corresponding to the original portrait image is edited; wherein the processed portrait image retains the facial features of the original portrait image and includes the facial expression features of the character described in the description text.

2. The method according to claim 1, wherein generating a cross-attention map corresponding to the original portrait image according to the original portrait image, the conditional embedding information and the facial mask map comprises: Compressing the original portrait image to obtain a latent space image; wherein the latent space image is an image with a dimension lower than that of the original portrait image and retains the facial features of the person; Generating a cross-attention map corresponding to the original portrait image according to the latent space image, the conditional embedding information and the facial mask image; The step of editing to obtain a processed portrait picture corresponding to the original portrait picture based on the conditional embedding information and the cross attention map includes: Based on the conditional embedding information and the cross-attention map, the latent space image is edited to generate the processed portrait image.

3. The method according to claim 2, wherein the editing of the latent space image based on the conditional embedding information and the cross attention map to generate the processed portrait image comprises: Add T noises to the latent space image through the first neural network to obtain a noise image, where T is a positive integer; According to the conditional embedding information and the cross attention map, performing T noise prediction and denoising operations on the noisy image through the first neural network to obtain a denoised image; The denoised image is restored to the dimension corresponding to the original portrait image to generate the processed portrait image.

4. The method according to claim 3, wherein the step of performing T noise prediction and denoising operations on the noisy image through the first neural network according to the conditional embedding information and the cross attention map to obtain a denoised image comprises: According to the conditional embedding information and the cross attention map, predicting the noise added to the latent space image for the nth time through the first neural network to obtain an nth predicted noise, where an initial value of n is 1 and n is a positive integer less than or equal to T; Removing the nth predicted noise from the noisy picture to obtain an updated noisy picture; wherein the initialized noisy picture is the noisy picture; When n is less than T, taking the value of n plus 1 as the updated n, and starting from the step of predicting the noise added to the latent space image for the nth time according to the conditional embedding information and the cross attention map by the first neural network to obtain the nth predicted noise; When n is equal to T, the updated noisy picture is determined as the denoised picture.

5. The method according to any one of claims 1 to 4, wherein generating conditional embedding information based on the original portrait image and the description text comprises: Extracting the facial features of the person contained in the original portrait image by using a facial recognition model; Encoding the description text through a text model to obtain text embedding information; The conditional embedding information is generated based on the facial features of the character and the text embedding information.

6. The method according to claim 5, wherein the description text comprises at least one segmented word, and the text embedding information comprises embedding information corresponding to each of the at least one segmented word; The step of generating the conditional embedding information based on the facial features of the person and the text embedding information includes: splicing the embedding information corresponding to the first participle in the at least one participle with the facial features of the person to obtain updated embedding information corresponding to the first participle; wherein the first participle refers to a pre-set participle used to represent the person; Using a multi-layer perceptron to process the updated embedding information corresponding to the first word segmentation, to obtain final embedding information corresponding to the first word segmentation; wherein the final embedding information corresponding to the first word segmentation has the same dimension as the embedding information corresponding to the first word segmentation; The embedding information corresponding to the first word segment in the text embedding information is replaced with the final embedding information corresponding to the first word segment to generate the conditional embedding information.

7. The method according to claim 5, wherein extracting the facial features of the person contained in the original portrait image by using a facial recognition model comprises: intercepting a facial region in the original portrait image to obtain a facial image corresponding to the original portrait image; The facial features of the person are extracted from the facial image using the facial recognition model.

8. The method according to any one of claims 2 to 7, wherein generating a cross-attention map corresponding to the original portrait image according to the latent space image, the conditional embedding information and the facial mask image comprises: Generate an initial cross-attention map corresponding to the original portrait image through a first neural network according to the latent space image and the conditional embedding information; Based on the facial mask image, other areas except the facial area in the initial cross-attention map are shielded to obtain the cross-attention map corresponding to the original portrait image.

9. A method for training an image processing model, the method being executed by a computer device, the method comprising: Acquire at least one training sample, each of which includes a set of corresponding sample portrait images and sample description texts; wherein the sample portrait images are images containing facial features of a person, and the sample description texts are texts used to describe facial features of the person in the sample portrait images; Based on the sample portrait picture and the sample description text, generating sample conditional embedding information through the picture processing model, the sample conditional embedding information including features of the sample portrait picture and features of the sample description text; According to the sample portrait picture, the sample conditional embedding information and the sample facial mask picture, a cross-attention map corresponding to the sample portrait picture is generated by the picture processing model, the sample facial mask picture is used to distinguish the facial area in the sample portrait picture from other areas except the facial area, and the cross-attention map is used to generate the expression content corresponding to the character expression feature in the facial area; Based on the sample conditional embedding information and the cross attention map, obtaining a processed sample portrait picture corresponding to the sample portrait picture through the picture processing model editing; Based on the sample portrait picture and the processed sample portrait picture, the parameters of the picture processing model are adjusted to obtain a trained picture processing model.

10. The method according to claim 9, wherein generating a cross-attention map corresponding to the sample portrait image through the image processing model according to the sample portrait image, the sample conditional embedding information and the sample facial mask image comprises: The sample portrait image is compressed by the image processing model to obtain a sample latent space image; wherein the sample latent space image is an image with a lower dimension than the sample portrait image and retains the facial features of the person; Generate a cross-attention map corresponding to the sample portrait image through the image processing model according to the sample latent space image, the sample conditional embedding information and the sample facial mask image; The step of obtaining a processed sample portrait picture corresponding to the sample portrait picture by editing the sample conditional embedding information and the cross attention map through the picture processing model includes: Based on the sample conditional embedding information and the cross-attention map, the sample latent space image is edited through the image processing model to generate a processed sample portrait image corresponding to the sample portrait image.

11. The method according to claim 10, wherein the editing of the sample latent space image by the image processing model based on the sample conditional embedding information and the cross attention map to generate a processed sample portrait image corresponding to the sample portrait image comprises: Adding T noises to the sample latent space image through the first neural network included in the image processing model to obtain a sample noise image, where T is a positive integer; According to the conditional embedding information and the cross attention map, performing T noise prediction and denoising operations on the sample noise image through the first neural network to obtain a sample denoised image; The sample denoised image is restored to a dimension corresponding to the sample portrait image to generate the processed sample portrait image.

12. The method according to claim 11, wherein the step of performing T noise prediction and denoising operations on the sample noise image through the first neural network according to the conditional embedding information and the cross attention map to obtain a sample denoised image comprises: According to the sample conditional embedding information and the cross attention map, predicting the noise added to the sample latent space image for the nth time through the first neural network to obtain the nth predicted noise, where the initial value of n is 1 and n is a positive integer less than or equal to T; Removing the nth predicted noise from the sample noisy picture to obtain an updated sample noisy picture; wherein the initialized sample noisy picture is the sample noisy picture; When n is less than T, taking the value of n plus 1 as the updated n, and starting from the step of predicting the noise added to the sample latent space image for the nth time according to the sample conditional embedding information and the cross attention map by the first neural network to obtain the nth predicted noise; When n is equal to T, the updated sample noisy picture is determined as the sample denoised picture.

13. The method according to any one of claims 9 to 12, wherein generating sample conditional embedding information through the image processing model based on the sample portrait image and the sample description text comprises: Extracting the facial features of the person contained in the sample portrait image through a facial recognition model; Encoding the sample description text by using a text model included in the image processing model to obtain sample text embedding information; The sample conditional embedding information is generated based on the facial features of the person and the sample text embedding information.

14. The method according to claim 13, wherein the sample description text comprises at least one segmented word, and the sample text embedding information comprises embedding information corresponding to each of the at least one segmented word; The step of generating the sample conditional embedding information based on the facial features of the person and the sample text embedding information includes: splicing the embedding information corresponding to the first participle in the at least one participle with the facial features of the person to obtain updated embedding information corresponding to the first participle; wherein the first participle refers to a pre-set participle used to represent the person; Using a multi-layer perceptron to process the updated embedding information corresponding to the first word segmentation, to obtain final embedding information corresponding to the first word segmentation; wherein the final embedding information corresponding to the first word segmentation has the same dimension as the embedding information corresponding to the first word segmentation; The embedding information corresponding to the first word segmentation in the sample text embedding information is replaced with the final embedding information corresponding to the first word segmentation to generate the sample conditional embedding information.

15. The method according to any one of claims 10 to 14, wherein generating a cross-attention map corresponding to the sample portrait image through the image processing model according to the sample latent space image, the sample conditional embedding information and the sample facial mask image comprises: According to the sample latent space image and the conditional embedding information, generating an initial cross-attention map corresponding to the sample portrait image through a first neural network included in the image processing model; Based on the sample facial mask image, other areas except the facial area in the initial cross-attention map are shielded to obtain the cross-attention map corresponding to the sample portrait image.

16. The method according to any one of claims 9 to 15, wherein the adjusting the parameters of the image processing model based on the sample portrait image and the processed sample portrait image to obtain the trained image processing model comprises: Based on the sample portrait picture and the processed sample portrait picture, a total loss is determined, the total loss including noise loss, identity loss and face position loss; wherein the noise loss is used to measure the difference between the added noise and the predicted noise, the identity loss is used to measure the difference between the feature representations of the sample portrait picture and the processed sample portrait picture in the face area, and the face position loss is used to measure the control ability of the cross attention map in the face area; The parameters of the image processing model are adjusted according to the total loss to obtain the trained image processing model.

17. The method according to claim 16, wherein determining the total loss based on the sample portrait picture and the processed sample portrait picture comprises: Determining a noise loss based on the noise added and the predicted noise during the editing of the sample latent space image; determining identity loss based on the sample portrait image and the processed sample portrait image; Determining a facial position loss based on the sample facial mask map and the cross-attention map; The total loss is determined based on the noise loss, the identity loss and the face position loss.

18. The method according to claim 17, wherein determining the identity loss based on the sample portrait image and the processed sample portrait image comprises: Extracting facial features of the sample portrait image through a facial recognition model to obtain first facial features; Extracting facial features of the processed sample portrait image through the facial recognition model to obtain second facial features; The identity loss is determined based on a similarity between the first facial feature and the second facial feature.

19. A picture processing device, the device being deployed on a computer device, the device comprising: An embedding module, used to generate conditional embedding information based on an original portrait image and a description text, wherein the conditional embedding information includes features of the original portrait image and features of the description text; wherein the original portrait image is a picture containing facial features of a person, and the description text is a text used to describe facial features of the person; An attention module, for generating a cross-attention map corresponding to the original portrait image according to the original portrait image, the conditional embedding information and a facial mask map, wherein the facial mask map is used to distinguish a facial region in the original portrait image from other regions except the facial region, and the cross-attention map is used to generate expression content corresponding to the character expression feature in the facial region; An editing module is used to edit, based on the conditional embedding information and the cross-attention map, a processed portrait image corresponding to the original portrait image; wherein the processed portrait image retains the facial features of the original portrait image and includes the facial expression features of the character described in the description text.

20. A training device for an image processing model, the device being deployed on a computer device, the device comprising: A sample acquisition module, used to acquire at least one training sample, each of which includes a set of corresponding sample portrait pictures and sample description texts; wherein the sample portrait pictures are pictures containing facial features of a person, and the sample description texts are texts used to describe the facial features of the person in the sample portrait pictures; A sample embedding module, configured to generate sample conditional embedding information based on the sample portrait image and the sample description text through the image processing model, wherein the sample conditional embedding information includes features of the sample portrait image and features of the sample description text; A sample attention module, used for generating a cross-attention map corresponding to the sample portrait image through the image processing model according to the sample portrait image, the sample conditional embedding information and the sample facial mask map, wherein the sample facial mask map is used to distinguish the facial area in the sample portrait image from other areas except the facial area, and the cross-attention map is used to generate expression content corresponding to the character expression feature in the facial area; A sample editing module, configured to obtain a processed sample portrait picture corresponding to the sample portrait picture by editing the picture processing model based on the sample conditional embedding information and the cross attention map; A model parameter adjustment module is used to adjust the parameters of the image processing model based on the sample portrait image and the processed sample portrait image to obtain a trained image processing model.

21. A computer device, comprising a processor and a memory, wherein a computer program is stored in the memory, and wherein the processor executes the computer program to implement the method according to any one of claims 1 to 8, or implements the method according to any one of claims 9 to 18.

22. A computer-readable storage medium, wherein a computer program is stored in the storage medium, wherein the computer program is configured to be executed by a processor to implement the method according to any one of claims 1 to 8, or to implement the method according to any one of claims 9 to 18.

23. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 8, or implements the method according to any one of claims 9 to 18.

Citation Information

Patent Citations

  • Interactive image editing method and device, readable storage medium and electronic equipment

    CN113448477A

  • Model training method and device, image processing method and device, medium and equipment

    CN116935166A

  • Image processing model training method and device, electronic equipment and storage medium

    CN116958325A

  • Image generation method and device, image model construction method and device, equipment and storage medium

    CN117037179A

  • Picture processing method and device, equipment and storage medium

    CN117557686A