Facial expression synthesis method and system based on AU control diffusion model

Through a diffusion model based on AU control, combined with AU estimator, fine-tuning strategy and inversion technology, facial expression images that conform to the target action unit are generated, solving the problem of difficulty in capturing the continuity and complexity of facial expressions in the prior art, and achieving high-precision and natural expression editing.

CN120107057AInactive Publication Date: 2025-06-06UNIV OF SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510220663.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing facial expression generation technology is difficult to achieve a generation method based on action unit (AU), so it is impossible to effectively capture the continuity and complexity of facial expressions.

Method used

Using a diffusion model based on AU control, the original image is analyzed through a pre-trained AU estimator, the initial AU vector is extracted, and the target AU vector is set. Then, the diffusion model is optimized through AU-perceived fine-tuning strategy, combined with inversion techniques and cross-attention mapping, and high-quality facial expression images that match the target AU.

Benefits of technology

It realizes high fidelity, high precision and natural expression editing, supports continuous interpolation and precise control of expressions, and is suitable for a variety of human-computer interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107057A_ABST
    Figure CN120107057A_ABST
Patent Text Reader

Abstract

The invention discloses a facial expression synthesis method and system based on an AU control diffusion model, and the method comprises the steps: obtaining an original image which comprises any expression; analyzing the original image by using a pre-trained AU estimator, extracting a current AU vector as an initial AU vector, and setting a target AU vector; according to the method, the diffusion process of an original image is inverted by adopting an optimization-based inversion technology, a sample is extracted from standard Gaussian distribution to serve as a starting point, the back diffusion process is started, noise is gradually removed, and finally a high-quality target image which not only conforms to a target AU vector but also keeps the space structure of the original image is generated. According to the method, high-fidelity, high-precision and natural expression editing can be realized. The method not only supports continuous interpolation and accurate control of expressions, but also can ensure high fidelity and detail retention of the generated image, realizes accurate quantification and editing control of the expressions through the AU vector, and is suitable for various man-machine interaction scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of facial expression editing, in particular to a facial expression synthesis method and system based on an AU controlled diffusion model. Background Art

[0002] In today's rapidly developing technological environment, human-computer interaction (HCI) has become a key area in the design of intelligent systems. With the continuous advancement of robotics and automatic control systems, how to improve the interaction experience between users and these systems has become the focus of researchers. Traditional human-computer interaction methods usually rely on simple instructions and feedback, lacking the integration of emotional dimensions, resulting in a monotonous and boring interaction experience.

[0003] Automatically generating facial expressions from a single image provides new possibilities for improving human-computer interaction. In recent years, with the widespread application of generative adversarial networks (GANs), significant progress has been made in the field of facial expression generation. For example, StarGAN and StyleGAN, etc. In today's rapidly developing technological environment, human-computer interaction (HCI) has become a key area in the design of intelligent systems. With the continuous advancement of robotics and automatic control systems, how to improve the interactive experience between users and these systems has become the focus of researchers. Traditional human-computer interaction methods usually rely on simple instructions and feedback, lack the integration of emotional dimensions, and lead to the emergence of a single architecture for interactive experience, which can not only generate new facial expressions, but also edit multiple attributes of the face, such as age, hair color, and gender. However, the GAN method is dependent on the number of attributes annotated in the dataset, and can only edit a limited number of facial features or expressions. In recent years, the rise of diffusion models has further promoted the development of image generation and editing. Diffusion models not only perform well in generation quality, but also support editing based on multimodal conditions such as text, which makes up for the deficiency of traditional methods that rely too heavily on reference images. However, the generation results of existing methods, whether based on GAN or diffusion models, are still limited to the limited attribute range defined in the dataset. For example, when StyleGAN is trained on a dataset with 8 expression labels, the generated results will be limited to these 8 expressions and cannot reflect the continuity and diversity of expressions.

[0004] Facial expressions are the result of complex coordinated movements of facial muscles, which are continuous and complex, while emotional images are usually simplified to discrete category labels or intensity levels. To capture this complexity, the Facial Action Coding System (FACS) was developed to describe facial expressions through action units (AUs), which are associated with specific muscle contractions. The FACS system is able to accurately quantify expressions, with 30 AUs directly corresponding to muscle activations, and more than 7,000 unique AU combinations have been observed in real scenes. For example, happy expressions are often associated with "cheek lift" (AU 6) and "mouth corner stretch" (AU 12), while disgust expressions involve "nose wrinkles" (AU 9), "mouth corner droop" (AU 15), and "chin lift" (AU 17). Facial expressions can show subtle differences through changes in the intensity of different AU activations. Compared with simple category or intensity labels, the FACS system provides a more accurate and detailed quantification method for emotional analysis of facial expressions. Therefore, how to synthesize facial expressions from a single image using an action unit-based generation method remains a huge challenge. Summary of the invention

[0005] In this embodiment, a facial expression synthesis method, system, electronic device and storage medium based on the AU controlled diffusion model are provided to realize the synthesis of facial expressions from a single image based on the generation method of action units.

[0006] In a first aspect, an embodiment of the present invention provides a facial expression synthesis method based on an AU controlled diffusion model, the facial expression synthesis method based on an AU controlled diffusion model comprising: Acquire an original image, wherein the original image includes any expression; Analyze the original image using a pre-trained AU estimator, extract a current AU vector as an initial AU vector, and set a target AU vector; The pre-trained diffusion model is optimized and adjusted through AU-aware fine-tuning strategies to improve the model's ability to distinguish expression-related AUs from irrelevant features. An optimization-based inversion technique is used to invert the diffusion process of the original image to obtain its latent representation and generate specific textual cues based on the content of the input image and the target AU vector; Replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU; Samples are drawn from the standard Gaussian distribution as the starting point to start the back diffusion process, gradually remove the noise, and finally generate a high-quality target image that not only conforms to the target AU vector but also maintains the spatial structure of the original image.

[0007] In an optional embodiment, using a pre-trained AU estimator to analyze the original image, extracting a current AU vector as an initial AU vector, and setting a target AU vector includes: Preprocess the original image; Input the processed original image into the pre-trained AU estimator; The AU estimator outputs a series of AU intensity values ​​according to the input image, indicating the activation degree of each AU; according to the result of the AU estimator, an initial AU vector is constructed, and a target vector is set.

[0008] In an optional embodiment, the pre-trained diffusion model is optimized and adjusted through an AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features, including: Obtaining a pre-trained diffusion model, wherein the pre-trained diffusion model is operable in the latent space of the autoencoder; The input image is converted into a latent representation through an autoencoder. In the latent space, Gaussian noise is gradually added to the latent representation to simulate the gradual degradation process. The gradually degraded latent representation is used to simulate the distribution change from a clear image to completely random noise. Through iterative optimization, the diffusion model gradually learns to separate noise from meaningful features in the latent space to capture the underlying structure of the data.

[0009] In an optional embodiment, the pre-trained diffusion model is optimized and adjusted through an AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features, and further includes: Construct the corresponding text prompt according to the content of the input image and the target AU vector; Combine text cues with image features and use the attention mechanism to align the target AU vector with the text cues. For situations without explicit text guidance, use a fixed empty text embedding for unconditional generation. Specific image-cue pairs are designed to introduce AU-specific cues so that the diffusion model learns to associate the visual features of the image with the corresponding AU attributes.

[0010] In an optional embodiment, an optimization-based inversion technique is used to invert the diffusion process of the original image to obtain its latent representation, and a specific text prompt is generated according to the content of the input image and the target AU vector, including: By adjusting the null text embedding at each step, we ensure that the input image is correctly reconstructed during the forward diffusion process; Generate specific textual cues based on the content of the input image and the target AU vector.

[0011] In an optional embodiment, the initial AU vector is replaced with the target AU vector, the original cross-attention map is combined and its value is updated according to the new text prompt to ensure that the generated content meets the target AU, including: Replace the initial AU vector with the target AU vector, instructing the model to generate images with new expressions; Combining the original criss-cross attention maps and updating their values ​​according to the new cues ensures that the generated image both conforms to the target AU vector and maintains the spatial structure of the original image.

[0012] In an optional embodiment, a sample is extracted from a standard Gaussian distribution as a starting point to start a reverse diffusion process, and finally a high-quality target image is generated that not only conforms to the target AU vector but also maintains the spatial structure of the original image, including: Draw a random noise vector from a standard Gaussian distribution as a starting point; Start the back-diffusion process to gradually remove noise, including predicting noise, updating latent representation, and iterative optimization; The decoder transforms the latent representation back to the original image space to generate the target image.

[0013] Compared with the prior art, the facial expression synthesis method based on the AU controlled diffusion model of the present invention has the following beneficial effects: In the modulation stage, the present invention optimizes the pre-trained diffusion-based language-image model through an AU-aware fine-tuning strategy, improving the model's ability to distinguish expression-related motion units from irrelevant features. In the editing stage, the input image is converted into latent noise through an inversion technique, and combined with text prompts containing initial and target AU vectors, attention control is used to achieve text-driven editing of local emotions; The final generated target image not only retains the overall structure of the input image, but also presents a new expression defined by the target AU vector. This method can achieve high-fidelity, high-precision and natural expression editing. This method not only supports continuous interpolation and precise control of expressions, but also ensures high fidelity and detail retention of the generated image. It realizes precise quantification and editing control of expressions through AU vectors, and is suitable for a variety of human-computer interaction scenarios.

[0014] In a second aspect, an embodiment of the present invention provides a facial expression synthesis system based on an AU controlled diffusion model, comprising: The acquisition module is configured to: acquire an original image, wherein the original image includes any expression; The initial processing module is configured to: analyze the original image using a pre-trained AU estimator, extract a current AU vector as an initial AU vector, and set a target AU vector; The optimization and adjustment module is configured to: optimize and adjust the pre-trained diffusion model through an AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features; A forward diffusion process inversion module is configured to: invert the diffusion process of the original image using an optimization-based inversion technique to obtain its latent representation, and generate specific textual cues based on the content of the input image and the target AU vector; The update module is configured to: replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU; The generation module is configured to extract samples from the standard Gaussian distribution as starting points, start the reverse diffusion process, gradually remove noise, and finally generate a high-quality target image that not only meets the target AU vector but also maintains the spatial structure of the original image.

[0015] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a processor, a communication interface, a memory and a bus, wherein the processor, the communication interface and the memory communicate with each other through the bus, and the processor can call logic instructions in the memory to execute the steps of the method provided in the first aspect.

[0016] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the facial expression synthesis method based on the AU controlled diffusion model as described in the first aspect.

[0017] Compared with the prior art, the beneficial effects of the facial expression synthesis system, electronic device and storage medium based on the AU controlled diffusion model of the present invention are the same as the facial expression synthesis method based on the AU controlled diffusion model described in the first aspect, so they will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 is a flow chart of a facial expression synthesis method based on an AU controlled diffusion model in an embodiment of the present invention; Figure 2 Schematic diagram of a facial expression synthesis method based on an AU controlled diffusion model in an embodiment of the present invention; Figure 3A model diagram of a method for synthesizing facial expressions from a single image based on an AU controlled diffusion model in an embodiment of the present invention; Figure 4 is a structural block diagram of a facial expression synthesis system based on an AU controlled diffusion model in an embodiment of the present invention; Figure 5 1 is a structural block diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0021] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0022] In an embodiment of the present invention, a facial expression synthesis method based on an AU controlled diffusion model is provided. Figure 1 is a flowchart of the facial expression synthesis method based on the AU controlled diffusion model of the present invention, combined with Figure 2 and Figure 3 As shown, the process includes the following steps: S100, obtaining an original image, wherein the original image includes any expression; In this embodiment, the RGB image As input, the image can contain any expression. Each expression is represented by a set of N action units (AUs) Indicates that each It is a normalized value ranging from 0 to 1 and is used to adjust the This continuous representation allows smooth interpolation of expressions between different states, thus generating diverse, realistic, and natural facial expressions.

[0023] The object of the present invention is to learn a conversion function , the input image Convert to target expression image , and its corresponding target action unit is That is, determine the following conversion relationship: .

[0024] S200, using a pre-trained AU estimator to analyze the original image, extracting a current AU vector as an initial AU vector, and setting a target AU vector; It should be noted that the AU estimator can receive a face image as input and output an action unit intensity value compatible with the FACS system.

[0025] Specifically, the original image is analyzed using a pre-trained AU estimator, a current AU vector is extracted as an initial AU vector, and a target AU vector is set, including: Preprocess the original image; if necessary, the input image can be preprocessed first, such as face detection, alignment, and size adjustment, to ensure that the image is suitable for input into the AU estimator.

[0026] Input the processed original image into the pre-trained AU estimator; The AU estimator outputs a series of AU intensity values ​​according to the input image, indicating the activation degree of each AU; according to the result of the AU estimator, an initial AU vector is constructed, and a target vector is set.

[0027] The AU estimator outputs a series of action unit intensity values, which are usually standardized between 0 and 1, indicating the activation level of each AU. From the results obtained by the AU estimator, an initial AU vector is constructed. This vector contains the intensity values ​​of all relevant AUs (action units) and is used to describe the expression state in the current image.

[0028] The target AU vector can be specified by the user or automatically selected by an algorithm. For example, the user can select expressions such as "happy" and "surprised", and the system will convert them into corresponding AU combinations. It should be noted that for the selected target expression, the corresponding AU configuration is found. This may involve consulting the FACS manual or using a predefined expression-AU mapping table to find the AUs and their intensities required to activate a specific expression. Based on the above mapping results, a new AU vector is created, namely the target AU vector. This vector is also a standardized vector in which each element corresponds to the expected intensity value of an AU. This is a routine step of the FACS system, so it will not be repeated here.

[0029] S300, optimize and adjust the pre-trained diffusion model through AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features; It should be noted that the present invention adopts a two-stage method, including a modulation stage and an editing stage.

[0030] For the modulation stage, a pre-trained latent diffusion model (LDM) is used. Unlike traditional diffusion models, LDM operates in the latent space of the autoencoder, significantly reducing the computational complexity while maintaining high-quality generation performance.

[0031] Specifically, the pre-trained diffusion model is optimized and adjusted through the AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features, including: Obtaining a pre-trained diffusion model, wherein the pre-trained diffusion model is operable in the latent space of the autoencoder; Input an image through the autoencoder Convert to latent representation , in the latent space, gradually move to the latent representation Add Gaussian noise to simulate the gradual degradation process and the gradually degraded potential representation Used to simulate the distribution change from clear images to completely random noise; Through iterative optimization, the diffusion model gradually learns to separate noise from meaningful features in the latent space to capture the underlying structure of the data.

[0032] Specifically, during the training process of the modulation phase, each noise sample is estimated by a neural network. The model gradually learns to separate noise from meaningful features in the latent space through iterative optimization, thereby effectively capturing the underlying structure of the data.

[0033] In addition, the pre-trained diffusion model is optimized and adjusted through the AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features, including: Construct the corresponding text prompt based on the content of the input image and the target AU vector ; Text prompt Combined with image features, the attention mechanism is used to make the target AU vector consistent with the text prompt Alignment; For the case without explicit text guidance, a fixed empty text embedding is used for unconditional generation; Specific image-cue pairs are designed to introduce AU-specific cues so that the diffusion model learns to associate the visual features of the image with the corresponding AU attributes.

[0034] Specifically, this method also introduces a text-guided generation mechanism. Text prompts are encoded through pre-trained language models (BERT, CLIP) is a series of word embeddings, and uses the attention mechanism to align the generated content with the text description. For scenarios without text guidance, the diffusion model uses a fixed empty text embedding On this basis, in order to better capture the relationship between facial images and action units (AUs), for each pair of facial images and its AU vector , design a specific image-prompt pair. For example, the prompt is defined as 'photo with AU'. By introducing AU-specific cues, the model is able to learn to associate the visual features of an image with the corresponding AU attributes, thereby achieving precise expression editing.

[0035] S400, inverting the diffusion process of the original image using an optimization-based inversion technique to obtain its latent representation, and generating a specific text prompt according to the content of the input image and the target AU vector; Specifically, an optimization-based inversion technique is used to invert the diffusion process of the original image to obtain its latent representation, and specific textual cues are generated according to the content of the input image and the target AU vector, including: By adjusting the null text embedding at each step, we ensure that the input image is correctly reconstructed during the forward diffusion process; Generate specific textual cues based on the content of the input image and the target AU vector.

[0036] In the editing stage of the present invention, the original image is processed by using an optimization-based inversion technique. The diffusion process of By adjusting the empty text embedding at each step , we hope To ensure that the input image is correctly reconstructed during the forward diffusion process. Subsequently, the input image is analyzed using the pre-trained AU estimator , get its AU vector , and generates hints that match the content of the input image .

[0037] S500, replacing the initial AU vector with the target AU vector, combining the original cross-attention map and updating its value according to the new text prompt to ensure that the generated content meets the target AU; Specifically, the initial AU vector is replaced with the target AU vector, the original cross-attention map is combined and its value is updated according to the new text prompt to ensure that the generated content meets the target AU, including: Replace the initial AU vector with the target AU vector, instructing the model to generate images with new expressions; Combining the original criss-cross attention maps and updating their values ​​according to the new cues ensures that the generated image both conforms to the target AU vector and maintains the spatial structure of the original image.

[0038] On this basis, continue to edit the expression of the picture. Replace with the target AU vector During the sampling process, the original cross-attention map is combined , and according to the new prompt Update its value to ensure that the generated image not only conforms to the target AU vector but also maintains the spatial structure of the original image. By modifying the expression-related terms in the prompt, the method focuses on expression-related pixels and uses cross-attention mapping Pinpoint areas of expression in an image for precise, controlled expression editing.

[0039] S600, extracting samples from the standard Gaussian distribution as starting points, starting the reverse diffusion process, gradually removing noise, and finally generating a high-quality target image that not only meets the target AU vector but also maintains the original image spatial structure.

[0040] Samples are drawn from the standard Gaussian distribution as the starting point to start the back diffusion process, and finally a high-quality target image is generated that not only conforms to the target AU vector but also maintains the spatial structure of the original image, including: A random noise vector is drawn from a standard Gaussian distribution as a starting point; this vector is located in the latent space and serves as the starting point for image generation.

[0041] Start the back-diffusion process to gradually remove noise, including predicting noise, updating latent representation, and iterative optimization; The decoder converts the latent representation back to the original image space to generate the target image. Once the latent representation becomes clear enough after multiple denoising, the decoder part of the autoencoder is used to convert the latent representation back to the original image space to generate a high-quality target image.

[0042] Specifically, the diffusion model gradually learns to separate noise from meaningful features in the latent space through iterative optimization, thereby effectively capturing the underlying structure of the data. In the reverse diffusion process, samples are drawn from a standard Gaussian distribution, gradually denoised, and finally a high-quality image is generated. The formal representation is as follows: ; in, and is random Gaussian noise, and They are respectively from The corresponding noise latent code is obtained. In order to retain the rich image priors learned by the diffusion model, it should be noted that in this embodiment, the number of fine-tuning steps is limited to a small value, usually about 150 steps.

[0043] It can be seen from the technical solution provided by the present invention that through the above steps, the target image finally generated is It not only retains the overall structure of the input image, but also presents the New expressions defined. This method can achieve high-fidelity, high-precision and natural expression editing. It not only supports continuous interpolation and precise control of expressions, but also ensures high fidelity and detail retention of generated images. Accurate quantification and editing control of expressions are achieved through AU vectors, which is suitable for a variety of human-computer interaction scenarios.

[0044] The following experiments were conducted on the public datasets EmotionNet and MEAD datasets, and compared with many advanced methods. The recognition results are shown in Table 1:

[0045] Table 1 It can be seen from Table 1 that the method proposed in the present invention performs well in terms of prediction accuracy, image generation quality and user satisfaction compared with other methods.

[0046] The embodiment of the present invention also provides a facial expression synthesis system based on the AU controlled diffusion model, which is used to implement the above method embodiment, which has been described and will not be repeated. The terms "module", "unit", "subunit", etc. used below can implement a combination of software and / or hardware for predetermined functions. Although the system described in the following embodiments is preferably implemented in software, the implementation of hardware or a combination of software and hardware is also possible and conceivable.

[0047] like Figure 4 As shown, Figure 4 : is a structural block diagram of a facial expression synthesis system based on an AU controlled diffusion model in the present invention, the system comprising: The acquisition module 101 is configured to: acquire an original image, wherein the original image includes any expression; The initial processing module 102 is configured to: analyze the original image using a pre-trained AU estimator, extract a current AU vector as an initial AU vector, and set a target AU vector; The optimization and adjustment module 103 is configured to: optimize and adjust the pre-trained diffusion model through an AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features; The forward diffusion process inversion module 104 is configured to: invert the diffusion process of the original image using an optimization-based inversion technique to obtain its latent representation, and generate a specific text prompt according to the content of the input image and the target AU vector; An updating module 105 is configured to: replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU; The generation module 106 is configured to extract samples from a standard Gaussian distribution as a starting point, start a reverse diffusion process, gradually remove noise, and finally generate a high-quality target image that not only meets the target AU vector but also maintains the original image spatial structure.

[0048] Figure 5 A structural block diagram of an electronic device provided by an embodiment of the present invention, such as Figure 5 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the following method: Acquire an original image, wherein the original image includes any expression; Analyze the original image using a pre-trained AU estimator, extract a current AU vector as an initial AU vector, and set a target AU vector; The pre-trained diffusion model is optimized and adjusted through AU-aware fine-tuning strategies to improve the model's ability to distinguish expression-related AUs from irrelevant features. An optimization-based inversion technique is used to invert the diffusion process of the original image to obtain its latent representation and generate specific textual cues based on the content of the input image and the target AU vector; Replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU; Samples are drawn from the standard Gaussian distribution as the starting point to start the back diffusion process, gradually remove the noise, and finally generate a high-quality target image that not only conforms to the target AU vector but also maintains the spatial structure of the original image.

[0049] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0050] An embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including: Acquire an original image, wherein the original image includes any expression; Analyze the original image using a pre-trained AU estimator, extract a current AU vector as an initial AU vector, and set a target AU vector; The pre-trained diffusion model is optimized and adjusted through AU-aware fine-tuning strategies to improve the model's ability to distinguish expression-related AUs from irrelevant features. An optimization-based inversion technique is used to invert the diffusion process of the original image to obtain its latent representation and generate specific textual cues based on the content of the input image and the target AU vector; Replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU; Samples are drawn from the standard Gaussian distribution as the starting point to start the back diffusion process, gradually remove the noise, and finally generate a high-quality target image that not only conforms to the target AU vector but also maintains the spatial structure of the original image.

[0051] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.

[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A facial expression synthesis method based on AU controlled diffusion model, characterized in that: The facial expression synthesis method based on the AU controlled diffusion model includes: Acquire an original image, wherein the original image includes any expression; Analyze the original image using a pre-trained AU estimator, extract a current AU vector as an initial AU vector, and set a target AU vector; The pre-trained diffusion model is optimized and adjusted through AU-aware fine-tuning strategies to improve the model's ability to distinguish expression-related AUs from irrelevant features. An optimization-based inversion technique is used to invert the diffusion process of the original image to obtain its latent representation and generate specific textual cues based on the content of the input image and the target AU vector; Replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU; Samples are drawn from the standard Gaussian distribution as the starting point to start the back diffusion process, gradually remove the noise, and finally generate a high-quality target image that not only conforms to the target AU vector but also maintains the spatial structure of the original image.

2. The facial expression synthesis method based on the AU controlled diffusion model according to claim 1, characterized in that: The original image is analyzed using a pre-trained AU estimator, a current AU vector is extracted as an initial AU vector, and a target AU vector is set, including: Preprocess the original image; Input the processed original image into the pre-trained AU estimator; The AU estimator outputs a series of AU intensity values ​​according to the input image, indicating the activation degree of each AU; according to the result of the AU estimator, an initial AU vector is constructed, and a target vector is set.

3. The facial expression synthesis method based on the AU controlled diffusion model according to claim 1, characterized in that: The pre-trained diffusion model is optimized and adjusted through the AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features, including: Obtaining a pre-trained diffusion model, wherein the pre-trained diffusion model is operable in the latent space of the autoencoder; The input image is converted into a latent representation through an autoencoder. In the latent space, Gaussian noise is gradually added to the latent representation to simulate the gradual degradation process. The gradually degraded latent representation is used to simulate the distribution change from a clear image to completely random noise. Through iterative optimization, the diffusion model gradually learns to separate noise from meaningful features in the latent space to capture the underlying structure of the data.

4. The facial expression synthesis method based on the AU controlled diffusion model according to claim 1, characterized in that: The pre-trained diffusion model is optimized and adjusted through the AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features, including: Construct the corresponding text prompt according to the content of the input image and the target AU vector; Combine text cues with image features and use the attention mechanism to align the target AU vector with the text cues. For situations without explicit text guidance, use a fixed empty text embedding for unconditional generation. Specific image-cue pairs are designed to introduce AU-specific cues so that the diffusion model learns to associate the visual features of the image with the corresponding AU attributes.

5. The facial expression synthesis method based on the AU controlled diffusion model according to claim 4, characterized in that: The diffusion process of the original image is inverted using an optimization-based inversion technique to obtain its latent representation and generate specific textual cues based on the content of the input image and the target AU vector, including: By adjusting the null text embedding at each step, we ensure that the input image is correctly reconstructed during the forward diffusion process; Generate specific textual cues based on the content of the input image and the target AU vector.

6. The facial expression synthesis method based on the AU controlled diffusion model according to claim 5, characterized in that: Replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU, including: Replace the initial AU vector with the target AU vector, instructing the model to generate images with new expressions; Combining the original criss-cross attention maps and updating their values ​​according to the new cues ensures that the generated image both conforms to the target AU vector and maintains the spatial structure of the original image.

7. The facial expression synthesis method based on the AU controlled diffusion model according to claim 6, characterized in that: Samples are drawn from the standard Gaussian distribution as the starting point to start the back diffusion process, and finally a high-quality target image is generated that not only conforms to the target AU vector but also maintains the spatial structure of the original image, including: Draw a random noise vector from a standard Gaussian distribution as a starting point; Start the back-diffusion process to gradually remove noise, including predicting noise, updating latent representation, and iterative optimization; The decoder transforms the latent representation back to the original image space to generate the target image.

8. A facial expression synthesis system based on AU controlled diffusion model, characterized in that: include: The acquisition module is configured to: acquire an original image, wherein the original image includes any expression; The initial processing module is configured to: analyze the original image using a pre-trained AU estimator, extract a current AU vector as an initial AU vector, and set a target AU vector; The optimization and adjustment module is configured to: optimize and adjust the pre-trained diffusion model through an AU-aware fine-tuning strategy to improve the model's ability to distinguish expression-related AUs from irrelevant features; A forward diffusion process inversion module is configured to: invert the diffusion process of the original image using an optimization-based inversion technique to obtain its latent representation, and generate specific textual cues based on the content of the input image and the target AU vector; The update module is configured to: replace the initial AU vector with the target AU vector, combine the original cross-attention map and update its value according to the new text prompt to ensure that the generated content meets the target AU; The generation module is configured to extract samples from the standard Gaussian distribution as starting points, start the reverse diffusion process, gradually remove noise, and finally generate a high-quality target image that not only meets the target AU vector but also maintains the spatial structure of the original image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the facial expression synthesis method based on the AU controlled diffusion model as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the facial expression synthesis method based on the AU controlled diffusion model as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Facial expression generation method, device and equipment based on expression parameter control

    CN118644589A