Text-driven face editing method based on multi-modal fusion

By improving the generator network structure and adopting multimodal fusion technology, the problem of poor performance of face editing in the existing technology when the attribute changes are large, high-quality fine-grained face editing and attribute decoupling are achieved, and the consistency of face identity is maintained.

CN119941925APending Publication Date: 2025-05-06KASHGAR ELECTRONIC INFORMATION IND TECH RES INST

Patent Information

Application Number
CN202411995433.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing GAN-based face editing algorithm has poor editing effect or even fails when facial attributes change greatly, and there is coupling between different facial attributes, resulting in inaccurate control.

Method used

The text-driven face editing method based on multimodal fusion is adopted. By improving the generator network structure, multimodal fusion and identity feature constraint strategy, fine-grained face editing is achieved, and high-quality face images are generated to achieve attribute decoupling and maintain the consistency of face identity.

Benefits of technology

It realizes high-quality face editing in general scenarios, solves the problem of attribute coupling, and maintains the editing effect when the operation is intense, ensuring the high quality of the generated image and the consistency of the face identity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941925A_ABST
    Figure CN119941925A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and discloses a text-driven face editing method based on multi-modal fusion, and the method comprises the steps: taking a to-be-processed source image and a text prompt as input; mapping the initial implicit code to a vector space through a mapping network to obtain an intermediate implicit code; respectively inputting the intermediate implicit code and the implicit code corresponding to the source image into a generator to obtain a first generated image and a second generated image; and constructing total loss by using the text loss, the style loss and the face loss, optimizing the generated image hidden code by using the total loss, and generating a final face image. According to the method, a StyleGAN semantic network is improved, text and image features are aligned through a CLIP pre-training model, and face image features before and after editing are aligned by using a face recognition network, so that a face editing image with high quality and good effect is generated, attribute decoupling is realized, and consistency of face identities is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, relates to multimodal fusion image processing technology, and in particular to a text-driven face editing method based on multimodal fusion. Background Art

[0002] With the development of deep learning technology, face editing technology based on generative adversarial networks (GAN) has been widely studied and applied. GAN is a network framework that includes a generator and a discriminator. It achieves the purpose of high-quality image generation by making the generated images more and more realistic. At present, face editing methods based on style generative adversarial networks (StyleGAN) are more prominent, such as StyleGAN and StyleClip. StyleGAN uses mapping networks and style transfer technology to generate high-resolution and realistic face images. StyleClip adds the CLIP model on the basis of StyleGAN, and realizes the control of face image editing through text instructions, such as "make this person look 10 years younger", and outputs the edited image.

[0003] However, existing GAN-based face editing algorithms also have some problems: (1) Because the image changes are too drastic, StyleClip will lead to poor editing effects or even failure when the facial attributes change greatly; (2) There is coupling between different facial attributes such as expression, age, and gender, resulting in inaccurate control.

[0004] Patent application CN202310945896.X discloses a text-driven face editing adversarial attack method, device and medium. The method determines a fixed semantic mapping model based on text samples; inputs the original image sample into the image inverse model to obtain the original image inverse features; then inputs it into the fixed semantic mapping model, and its output is weightedly combined with the original image inverse features to obtain the edited image inverse features, and the weighting factor is an optimizable hyperparameter; the edited image is generated according to the edited image inverse features. This method ensures the attack effectiveness of adversarial samples, enhances the invisibility of adversarial interference, improves the image quality of adversarial samples, etc. However, due to the use of the traditional StyleGAN algorithm, the problem of poor editing effect or even failure when facial attributes change greatly is not solved; and the problem of decoupling facial attributes in text editing at a fine-grained level is still not solved by mapping the hidden code to the style space. Summary of the invention

[0005] In view of the problems of low image quality, attribute coupling and failure under intense operation caused by existing face editing methods, the purpose of the present invention is to provide a text-driven face editing method based on multimodal fusion. Through the improvement of the generator network structure, multimodal fusion and identity feature constraint strategy, fine-grained face editing is achieved with text, and high-quality face images are generated, achieving attribute decoupling and maintaining consistent face identity.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions.

[0007] The text-driven face editing method based on multimodal fusion provided by this method comprises the following steps:

[0008] S1 takes the source image to be processed and the text prompt as input;

[0009] S2 maps the initial latent code to the vector space through the mapping network to obtain the intermediate latent code;

[0010] S3 inputs the intermediate hidden code and the hidden code corresponding to the source image into the generator respectively to obtain a first generated image and a second generated image;

[0011] S4 constructs the total loss, uses the total loss to optimize the generated image hidden code, and then returns to step S3 until the total loss is minimized; this step includes the following sub-steps:

[0012] S41 inputs the first generated image and the text prompt into the discriminator to obtain the text loss;

[0013] S42 obtains the style loss based on the deviation between the intermediate hidden code and the hidden code corresponding to the source image;

[0014] S43 inputs the first generated image and the second generated image into a face recognition network to obtain a face loss;

[0015] The sum of S44 text loss, style loss and face loss constitutes the total loss;

[0016] S45 uses the total loss to optimize the intermediate hidden code.

[0017] The mapping network used in the present invention is composed of several fully connected layers arranged in sequence.

[0018] The generator used in the present invention is improved based on the StyleGAN semantic network, which includes an input module, a synthesis module and an image output module. The input module is used to input Fourier features; the synthesis module includes several synthesis layers with the same structure, and the synthesis layer is used to fuse the input intermediate hidden code with the output of the previous synthesis layer or input module; the image output module is used to generate an image according to the output result of the synthesis module. The number of the synthesis layers is 14. The output of the mapping network is input to the corresponding synthesis layer of the generator through an affine transformation network. The synthesis layer first inputs the input intermediate hidden code through modulation (Mod function) and demodulation (Demod function) and inputs it into the convolution layer, and convolves it with the result obtained by averaging the output of the previous synthesis layer or input module (EMA function); the processing result is superimposed with the introduced bias (b), and then further upsampled, activated (Leaky Relu function), downsampled and cropped (Crop function) to obtain the output result of this layer.

[0019] The discriminator used in the present invention is a CLIP model. The CLIP model uses a dual encoder structure, including an image encoder and a text encoder.

[0020] In the above step S2, the initial latent code may be a noise vector satisfying a normal distribution or a randomly generated noise vector.

[0021] In the above step S3, the latent code of the source image to be processed is obtained by image embedding.

[0022] In the above step S4, the total loss is obtained based on the principle mechanism of multimodal fusion, and then the intermediate hidden code is optimized by gradient descent through back propagation.

[0023] In the above step S41, the first generated image and the text prompt are encoded respectively by the image encoder and the text encoder of the CLIP model, and then the cosine distance between the two parameter encodings is calculated as the text loss, which is recorded as D CLIP (G(w),t), G(w) represents the first generated image, G(·) represents the generator, w represents the intermediate hidden code, and t represents the text prompt.

[0024] In the above step S42, the L2 loss function is used to calculate the deviation ∥ww between the intermediate hidden code and the hidden code corresponding to the source image s ∥2 is used as the style loss to ensure that the style of the generated image is similar to that of the source image.

[0025] In the above step S43, based on the face recognition network structure, an additive angle interval loss function is added, that is, the face loss L iD(w). In the present invention, the face recognition network R is used to calculate the cosine distance between the source image and the generated image, so that the modified face maintains a high similarity with the original face. The face recognition network can be Resnet50, ArcFace, VGGFace2, DeepFace, etc. Specifically, the face loss L ID (w) is calculated as follows:

[0026] L ID (w) = 1 - <R(G(w s )),R(G(w))>

[0027] In the formula, R(·) represents the face recognition network, G(w s ) represents the second generated image (i.e., the face image generated by hidden coding of the source image), w s Represents the hidden code corresponding to the source image. <R(G(w s )),R(G(w))> represents the cosine distance between the two.

[0028] In the above step S44, the total loss is specifically expressed as follows:

[0029]

[0030] In the formula, λ L2 represents the style loss weight, λ ID Represents the face loss weight.

[0031] In the above step S45, the Adam optimization algorithm is used to optimize the intermediate latent code. When steps S3-S4 are iterated for 300 to 400 steps, the total loss tends to be stable and reaches the minimum.

[0032] Compared with the prior art, the text-driven face editing method based on multimodal fusion provided by the present invention has the following beneficial effects:

[0033] (1) The present invention is aimed at text-driven face editing in general scenarios and belongs to a multimodal fusion image generation technology. The StyleGAN semantic network is improved, and the text and image features are aligned through the CLIP pre-training model. At the same time, the face recognition network is used to align the face image features before and after editing to generate high-quality and effective face editing images, and to achieve attribute decoupling and maintain consistent face identity.

[0034] (2) The present invention maps text prompts to the style space to obtain a single global attribute direction, thereby improving the editing effect of facial attributes. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Schematic diagram of the generator network structure.

[0036] Figure 2 A schematic flow chart of the text-driven face editing method based on multimodal fusion provided by the present invention.

[0037] Figure 3 Schematic diagram of the principle of text-driven face editing method. DETAILED DESCRIPTION

[0038] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0039] Example

[0040] This embodiment provides a text-driven face editing method based on multimodal fusion, which is implemented using a generative adversarial network improved based on StyleGAN2.

[0041] like Figure 1 As shown, the generative adversarial network used in this embodiment includes a mapping network, a generator, a discriminator and a face recognition network.

[0042] In this embodiment, the initial latent code in a given latent space Z (e.g., a noise vector satisfying a normal distribution or a randomly generated noise vector) is mapped to a vector space to obtain an intermediate latent code through a mapping network. The mapping network is composed of 8 layers of fully connected layers arranged in sequence. The output result of the mapping network is input into the generator after affine transformation.

[0043] The generator is used to generate a face image based on the input intermediate latent code, which includes an input module, a synthesis module and an image output module.

[0044] The input module is used to input Fourier features; the input module includes a Fourier transform unit and a convolution layer. The Fourier transform unit uniformly samples the data input after affine transformation within the frequency range to obtain Fourier features, and convolves the data through a 1×1 convolution layer and inputs the Fourier features into the synthesis module.

[0045] The synthesis module includes 14 layers of synthesis layers with the same structure. The output of the mapping network is input into the corresponding synthesis layer of the generator through an affine transformation network. Compared with the traditional 18-layer feature extraction structure, the synthesis module has been significantly simplified, and the noise input has been deleted, making the network structure more lightweight. The synthesis layer is used to fuse the input intermediate hidden code with the output of the previous synthesis layer or input module. The synthesis layer first modulates the input intermediate hidden code with the given weight ω through the Mod function, and then demodulates it through the Demod function. After that, it is input into the convolution layer (3×3 convolution layer or 1×1 convolution layer), and convolved with the result obtained by averaging the output of the previous synthesis layer or input module through the EMA function; the processing result is superimposed with the introduced bias (b), and then further upsampled, activated (Leaky Relu function), downsampled and cropped (Crop function) to obtain the output result of this layer.

[0046] The image output module (toRGB) is used to generate an RGB image based on the output of the synthesis module. The image output module uses the conventional toRGB setting in the traditional StyleGAN2.

[0047] The discriminator is used to make similar images and text prompts close together in the feature space, and dissimilar ones far apart. In this embodiment, the discriminator is a CLIP model. The CLIP model uses a dual encoder structure, including an image encoder and a text encoder.

[0048] The face recognition network is used to identify the face image generated by the generator and the source image, so that the modified face maintains a high degree of similarity with the face in the source image. In this embodiment, the face recognition network uses Resnet50, for details, see ArcFace: Additive Angular Margin Loss for Deep Face Recognition. The additive angular margin loss function is added to the traditional Resnet50 network.

[0049] The above generator, discriminator and face recognition network can be pre-trained by conventional methods.

[0050] Based on the above analysis, the text-driven face editing method based on multimodal fusion provided in this embodiment is as follows: Figure 2-3 As shown, it includes the following steps:

[0051] S1 takes as input a source image to be processed and a textual prompt.

[0052] Figure 3 The source image to be processed and the text prompt "big eyes" are given.

[0053] S2 maps the initial latent code to the vector space through the mapping network to obtain the intermediate latent code.

[0054] In this step, the initial latent code can be a noise vector that satisfies the normal distribution or a randomly generated noise vector. Through the mapping network, the initial latent code in the latent space Z is mapped to the vector space W+ to obtain the intermediate latent code w.

[0055] S3 inputs the intermediate hidden code and the hidden code corresponding to the source image into the generator respectively to obtain the first generated image and the second generated image.

[0056] In this step, the latent code of the source image to be processed (w s , w s ∈W+) is obtained through image embedding processing. The specific operation is: input the source image to be processed into the pre-trained e4e (encoder4editing) model, and the output is the hidden code of the source image to be processed.

[0057] The intermediate hidden code w and the hidden code w corresponding to the source image s Input the generator respectively to obtain the first generated image G(w) and the second generated image G(w s ).

[0058] S4 constructs the total loss, uses the total loss to optimize the generated image hidden code, and then returns to step S3 until the total loss is minimized.

[0059] This step is based on the principle mechanism of multimodal fusion to obtain the total loss, and then performs gradient descent optimization on the intermediate hidden code through back propagation.

[0060] This step includes the following sub-steps:

[0061] S41 inputs the first generated image and text prompt into the discriminator to obtain the text loss.

[0062] In this step, the first generated image and text prompt are encoded by the image encoder and text encoder of the CLIP model respectively, and then the cosine distance between the two parameter encodings is calculated as the text loss, denoted as D CLIP (G(w), t), G(·) represents the generator, G(w) represents the first generated image (that is, the edited face image), w represents the intermediate hidden code (that is, the hidden code corresponding to the generated face image, which will be iteratively updated through gradient descent), and t represents the text prompt.

[0063] S42 obtains the style loss based on the deviation between the intermediate hidden code and the hidden code corresponding to the source image.

[0064] In this step, the L2 loss function is used to calculate the deviation between the intermediate hidden code (i.e., the iterative hidden code) and the hidden code corresponding to the source image ∥ww s ∥2 is used as the style loss to ensure that the style of the generated image is similar to that of the source image.

[0065] S43 inputs the first generated image and the second generated image into a face recognition network to obtain face loss.

[0066] In this step, the face recognition network is used to recognize the first generated image and the second generated image respectively, and then the face loss L is calculated according to the following formula: ID (w):

[0067] L ID (w) = 1 - <R(G(w s )), R(G(w))>

[0068] In the formula, R(·) represents the face recognition network, G(w s ) represents the second generated image (i.e., the face image generated by hidden coding of the source image), w s Represents the hidden code corresponding to the source image. <R(G(w s )), R(G(w))> represents the cosine distance between the two.

[0069] The sum of S44 text loss, style loss and face loss constitutes the total loss;

[0070] According to steps S41-S43, the total loss is specifically expressed as follows:

[0071]

[0072] In the formula, λ L2 represents the style loss weight, λ ID Represents the face loss weight.

[0073] S45 uses the total loss to optimize the intermediate hidden code.

[0074] According to the calculated total loss, the Adam optimization algorithm is used to optimize the intermediate latent code.

[0075] Therefore, this embodiment completes text-driven face editing by minimizing the total loss. Through continuous iteration, the editing and correction of the source image based on text prompts is completed while ensuring that the generated face image is highly similar to the source image and the style is consistent. When steps S3-S4 are iterated for 300 to 400 steps, the total loss tends to be stable and reaches the minimum.

[0076] from Figure 3It can be seen that the generative adversarial network provided by the present invention generates a face image with enlarged eyes according to the text prompt "big eyes", and the final generated image is consistent with the style and face identity of the source image, which meets the requirements of text-driven face editing.

[0077] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0078] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0080] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

[0081] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A text-driven face editing method based on multimodal fusion, characterized in that: The following steps are involved: S1 takes the source image to be processed and the text prompt as input; S2 maps the initial latent code to the vector space through the mapping network to obtain the intermediate latent code; S3 inputs the intermediate hidden code and the hidden code corresponding to the source image into the generator respectively to obtain a first generated image and a second generated image; S4 constructs the total loss, uses the total loss to optimize the generated image hidden code, and then returns to step S3 until the total loss is minimized; this step includes the following sub-steps: S41 inputs the first generated image and the text prompt into the discriminator to obtain the text loss; S42 obtains the style loss based on the deviation between the intermediate hidden code and the hidden code corresponding to the source image; S43 inputs the first generated image and the second generated image into a face recognition network to obtain a face loss; The sum of S44 text loss, style loss and face loss constitutes the total loss; S45 uses the total loss to optimize the intermediate hidden code.

2. The text-driven face editing method based on multimodal fusion according to claim 1, characterized in that: The mapping network consists of several fully connected layers arranged in sequence.

3. The text-driven face editing method based on multimodal fusion according to claim 2 is characterized in that: The generator is improved based on the StyleGAN semantic network, which includes an input module, a synthesis module and an image output module. The input module is used to input Fourier features; the synthesis module includes several synthesis layers with the same structure, and the synthesis layer is used to fuse the input intermediate hidden code with the output of the previous synthesis layer or input module; the image output module is used to generate an image according to the output result of the synthesis module; the number of synthesis layers is 14; the output of the mapping network is input into the corresponding synthesis layer of the generator through an affine transformation network.

4. The text-driven face editing method based on multimodal fusion according to claim 3 is characterized in that: The synthesis layer first inputs the intermediate hidden code of the input into the convolution layer after modulation and demodulation, and performs convolution processing with the result obtained by averaging the output of the previous synthesis layer or input module; the processing result is superimposed with the introduced bias and further upsampled, activated, downsampled and cropped to obtain the output result of this layer.

5. The text-driven face editing method based on multimodal fusion according to claim 1, characterized in that: The discriminator is the CLIP model.

6. The text-driven face editing method based on multimodal fusion according to claim 1, characterized in that: In step S2, the initial latent code is a noise vector satisfying a normal distribution or a randomly generated noise vector; in step S3, the latent code of the source image to be processed is obtained by image embedding processing.

7. The text-driven face editing method based on multimodal fusion according to any one of claims 1 to 6, characterized in that: In step S41, the first generated image and the text prompt are encoded by the image encoder and the text encoder of the CLIP model respectively, and then the cosine distance between the two parameter encodings is calculated as the text loss, which is recorded as D CLIP (G(w),t), G(w) represents the first generated image, G(·) represents the generator, w represents the intermediate hidden code, and t represents the text prompt.

8. The text-driven face editing method based on multimodal fusion according to claim 7, characterized in that: In step S42, the L2 function is used to calculate the deviation between the intermediate hidden code and the hidden code corresponding to the source image ∥ww s ∥2 as the style loss.

9. The text-driven face editing method based on multimodal fusion according to claim 8, characterized in that: In step S43, the face loss L ID (w) is calculated as follows: L ID (w)=1-<R(G(w s )),R(G(w))> In the formula, R(·) represents the face recognition network, G(w s ) represents the second generated image, w s Represents the hidden code corresponding to the source image.

10. The text-driven face editing method based on multimodal fusion according to claim 9, characterized in that: In step S44, the total loss is specifically expressed as follows: In the formula, λ L2 represents the style loss weight, λ ID Represents the face loss weight.

Citation Information

Patent Citations

  • Text-driven face editing anti-attack method and device and medium

    CN116796829A

Cited By

  • Method and device for migrating voice styles

    CN120808752A

  • AI digital human-oriented multi-modal text-guided image editing optimization method and AI digital human-oriented multi-modal text-guided image editing optimization system

    CN121095376A