A preference optimization enhanced three-dimensional face generation method and system across modal learning
By using a preference optimization enhancement method based on cross-modal learning, the problems of insufficient cross-modal modeling capability and misalignment of user preferences in the generation of 3D speaker facial animation are solved, achieving high-precision lip-sync and natural expression of multiple emotions, which is suitable for virtual human interaction and immersive human-computer interface.
Patent Information
- Application Number
- CN202511006660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing methods for generating 3D speaker facial animations lack cross-modal modeling capabilities, have limited model generalization ability, and lack user preference alignment mechanisms, making it difficult to achieve high-precision lip-syncing, natural expression of multiple emotions, and personalized adaptation.
We employ a cross-modal learning-based preference optimization enhancement method. By constructing a generative architecture with modality fusion capabilities, designing a multi-stage curriculum-based training mechanism, and introducing a temporal-aware hybrid expert mechanism and an emotion control mechanism, combined with a closed-loop preference optimization mechanism, we achieve high-precision lip-phonetic synchronization and natural expression of multiple emotions.
It significantly improves the model's ability to generalize to speaker identity, emotion category, and language content. The generated results are highly consistent with user preferences, and it has high lip-movement synchronization accuracy and natural, rich, and controllable emotional expression, making it suitable for virtual human interaction and immersive human-computer interfaces.
Smart Images

Figure CN120510260B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D speaker face generation, and more particularly to a method and system for enhancing 3D face generation through cross-modal learning and preference optimization. Background Technology
[0002] With the rapid development of virtual humans, digital humans, virtual reality (VR), and immersive interactive systems, audio-driven 3D facial animation generation has become one of the key technologies for achieving natural human-computer interaction. This task requires the model to generate a corresponding 3D facial animation sequence based on the input audio content. It must not only ensure high-precision lip-sync between the lips and the audio, but also possess the ability to express emotional states and the ability to personalize the generation behavior, thereby more realistically simulating human facial movements.
[0003] The existing technologies have the following shortcomings: 1) Insufficient cross-modal modeling capability: Current mainstream methods mostly use convolutional neural networks (CNN) or variational autoencoders (VAE) as the backbone of the generative model. Their modeling capabilities are mainly limited to short-term dependencies or static frame reconstruction, lacking the ability to express long-term facial motion and cross-temporal emotional dynamics, and making it difficult to accurately capture complex and varied facial behaviors driven by audio; 2) Limited model generalization ability: High-quality emotionally labeled 3D facial data acquisition is costly. Due to insufficient data, the model is difficult to achieve good generalization and emotional expression capabilities; 3) Lack of user preference alignment mechanism: Traditional training objectives rely on likelihood estimation or auxiliary loss functions, which cannot accurately reflect users' subjective evaluation of the naturalness of facial expressions and the accuracy of synchronization, resulting in an inconsistency between the optimization direction and the actual perceived effect.
[0004] Currently, there is a lack of a unified architecture that can simultaneously achieve a balance between lip-syncing, emotional control, diverse expression, and alignment with user subjective preferences. Summary of the Invention
[0005] To overcome the shortcomings of existing 3D speaker facial animation generation methods, such as limited model structure expression, limited emotion control, and misalignment of user preferences, this invention proposes a cross-modal learning-based preference optimization and enhancement 3D facial generation method and system. By constructing a generation architecture with modal fusion capabilities, designing a multi-stage course-based training mechanism, and introducing preference optimization and enhancement strategies, it achieves high-precision lip-phonetic synchronization, natural expression of multiple emotions, and personalized adaptation capabilities under cross-modal input conditions.
[0006] The present invention adopts the following specific technical solution:
[0007] In a first aspect, the present invention provides a cross-modal learning-based preference optimization-enhanced 3D face generation method, comprising:
[0008] (1) Obtain a first dataset containing emotion labels and a second dataset not containing emotion labels. Both datasets contain facial image sequences and their corresponding driving audio. The emotion labels are parameter groups that reflect the expression state of each frame of the image.
[0009] (2) A face generation diffusion model with temporal awareness hybrid expert mechanism and emotion control mechanism is trained by course learning. During the course learning process, the model is first pre-trained with audio-lip movement synchronization using the second dataset. During the pre-training, the embedding of the person's identity feature and the driving audio embedding are used as input conditions. Then, the model is fine-tuned with emotion driving using the first dataset. During the fine-tuning training, the embedding of the person's identity feature and the driving audio embedding are used as input conditions, and the emotion label embedding is injected based on the emotion control mechanism.
[0010] (3) Introduce a closed-loop preference optimization mechanism, predict the emotional naturalness score of the facial animation generated by the facial generation diffusion model through the reward estimation model, generate winning and losing sample pairs by combining the lip-sync evaluation tool, and update the facial generation diffusion model by combining the preference optimization loss function and the flow matching loss, iterating multiple times to achieve preference optimization enhancement;
[0011] (4) Use the optimized facial generation diffusion model to generate facial animations that meet the preference conditions.
[0012] Furthermore, the facial frame images in both datasets are three-dimensional scan images of the face, and the parameter set reflecting the expression state of each frame image includes 50-dimensional facial expression parameters and 3-dimensional jaw pose parameters from the three-dimensional scan images.
[0013] Furthermore, the face generation diffusion model is implemented based on the flow matching DiT model, and a time-aware hybrid expert mechanism and an emotion control mechanism are introduced on the basis of the DiT model. The time-aware hybrid expert mechanism uses a hybrid expert system to replace the feedforward network of the DiT model, and the emotion control mechanism refers to using an emotion adaptive normalization module to replace the standard normalization layer of the DiT model.
[0014] Furthermore, the hybrid expert system includes multiple expert subnetworks and a dynamic router. The dynamic router activates several expert subnetworks and assigns weights at each inference time step, and outputs the weighted result of the expert subnetworks.
[0015] Furthermore, the calculation formula for the emotion adaptive normalization module is as follows:
[0016] ;
[0017] in, , It is a linear projection layer, where x and e are the input features and sentiment label embeddings of the sentiment adaptive normalization module, respectively. Let L2 norm be the input feature dimension. This represents the emotion adaptive normalization module.
[0018] Further, step (3) includes:
[0019] (3.1) Construct a preference dataset containing generated animations, real animations and their respective preference ratings, and train a reward estimation model using the preference dataset. The reward estimation model takes the animation as input and predicts the preference rating.
[0020] (3.2) Use the facial generation diffusion model to generate facial animations in batches. Use the lip-phonetic synchronization scorer and the reward estimation model to estimate the lip-phonetic synchronization score and the emotional naturalness score, respectively. Select the win-loss sample pairs with significant score differences from multiple facial animations generated by the same driving audio, and calculate the preference optimization loss function.
[0021] (3.3) Combine the preference optimization loss function and the flow matching loss to update the face generation diffusion model;
[0022] (3.4) Iterate through multiple rounds, and in each round, reconstruct the preference dataset and update the reward estimation model using the updated face generation diffusion model.
[0023] Furthermore, the preference optimization loss function is as follows:
[0024]
[0025]
[0026]
[0027] in, Indicates preference optimization loss; Indicates the parameters of the current facial generation diffusion model. The difference in flow matching loss between the winning and losing sample pairs. This represents the parameters of the previous round of facial generation diffusion model. The difference in flow matching loss between the winning and losing sample pairs. This indicates that the time step t follows a continuous uniform distribution on the interval [0,1]. , Indicates win / loss sample pairs, This represents the scaling factor, used to control the degree to which the differences in loss values are amplified. This represents the sigmoid function, where t represents the time step. and They represent respectively to , The result after adding noise at time step t This represents the velocity vector predicted by the facial generation diffusion model. and These represent the true velocity vectors of the winning and losing sample pairs, This represents the square of the L2 norm.
[0028] Furthermore, the process of constructing the preference dataset is as follows:
[0029] Given audio and emotion tags for a batch of real animations, the face generation diffusion model after two courses is run twice to obtain two generated animations for each real animation;
[0030] The real animation and its two corresponding generated animations are combined into a triplet. The annotator assigns a preference score to each animation in different triplets, resulting in a preference dataset consisting of the animations and preference scores.
[0031] Secondly, this invention proposes a cross-modal learning preference optimization enhancement 3D face generation system to implement the aforementioned cross-modal learning preference optimization enhancement 3D face generation method.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] This invention proposes a facial generation diffusion model that incorporates a time-aware hybrid expert mechanism and an emotion control mechanism, breaking through the modeling capacity bottleneck of traditional convolutional neural networks or U-Net. Simultaneously, it employs a two-stage learning strategy, effectively combining large-scale data with high-quality, refined data, significantly improving the model's generalization ability to conditions such as speaker identity, emotion category, and language content. A closed-loop preference optimization mechanism is introduced, ensuring that the model's optimization metrics are highly consistent with actual user preferences. While maintaining high lip-movement synchronization accuracy, it achieves natural, rich, and controllable emotional expression. This method is designed for multimodal 3D facial animation modeling tasks, taking audio, character identity encoding, and emotion tag control commands as inputs, and outputting a 3D facial deformation parameter sequence that is precisely synchronized with the audio semantic content and can express the target's emotional state. The generated results simultaneously possess structural naturalness, expressive diversity, and consistency with user subjective preferences, making it suitable for applications such as virtual human interaction, digital human generation, and immersive human-computer interfaces. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating a cross-modal learning-based preference optimization enhancement method for 3D face generation according to the present invention.
[0035] Figure 2 This is a schematic diagram of the facial generation diffusion model of the present invention.
[0036] Figure 3 This is a diagram illustrating the first stage of training in the course.
[0037] Figure 4 This is a schematic diagram of the second stage of course learning and closed-loop preference optimization training.
[0038] Figure 5 This is a schematic diagram of the three-dimensional face generation process. Detailed Implementation
[0039] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.
[0040] The accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0041] The flowchart shown in the attached diagram is merely an illustrative example and does not necessarily include all steps. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0042] like Figure 1 As shown, the present invention proposes a cross-modal learning-based preference optimization enhancement method and system for 3D face generation, which mainly includes the following steps:
[0043] S1, obtain a first dataset containing emotion labels and a second dataset not containing emotion labels. Both datasets contain facial image sequences and their corresponding driving audio. The emotion labels are parameter groups that reflect the expression state of each frame of the image.
[0044] S2 trains a facial generation diffusion model that incorporates a temporally aware hybrid expert mechanism and an emotion control mechanism through a course-based learning approach. During the course-based learning process, the model is first pre-trained on the second dataset using audio-lip movement synchronization, with the person's identity feature embedding and the driving audio embedding as input conditions. Then, the model is fine-tuned on the first dataset using emotion-driven training, with the person's identity feature embedding and the driving audio embedding as input conditions, and emotion label embedding is injected based on the emotion control mechanism.
[0045] S3 introduces a closed-loop preference optimization mechanism, which predicts the emotional naturalness score of the facial animation generated by the facial generation and diffusion model through a reward estimation model, generates winning and losing sample pairs by combining a lip-sync evaluation tool, and updates the facial generation and diffusion model by combining the preference optimization loss function and the flow matching loss, iterating multiple times to achieve preference optimization enhancement.
[0046] S4 uses the optimized facial generation diffusion model to generate facial animations that meet the preference conditions.
[0047] Step S1 above is the step of obtaining the training dataset. One possible method is as follows:
[0048] The training datasets in this embodiment come from large-scale audio and video datasets, such as LRS3-TED and HDTF, which contain a large number of video clips of real people speaking, used to establish a robust basic mapping relationship from audio to lip movements; 3D scan data VOCASET; and high-quality emotional facial animation datasets, such as MEAD and RAVDESS, which have clear emotional labels and high-quality video quality.
[0049] The datasets described above can be processed to obtain a first dataset containing emotion labels and a second dataset without emotion labels. Both datasets contain facial image sequences and their corresponding driving audio. The facial frame images are three-dimensional scan images of the face, and the emotion labels are a set of parameters reflecting the expression state of each frame image.
[0050] When acquiring emotion tags, advanced 3D reconstruction tools (such as EMOCA) can be used to reconstruct the 3D facial mesh of each frame from the video frames and convert it into the FLAME model parameter format. The FLAME model represents the face as identity shape parameters, facial pose parameters, and facial expression parameters. To improve computational efficiency and modeling accuracy, only 53-dimensional parameters (50-dimensional facial expression parameters + 3-dimensional jaw pose parameters) of each frame are retained as the final emotion tag, which is a set of 53-dimensional parameter groups; the identity of each person is encoded as a fixed value.
[0051] This invention employs the FLAME parameter extraction tool EMOCA to extract 53-dimensional facial expression and jaw pose parameters from videos, which are then used as real data for model learning. Replacing the 5023*3-dimensional vertex coordinates with these 53-dimensional facial expression and jaw pose parameters as the prediction target significantly improves generation efficiency and stability, and can be utilized by mainstream animation engines.
[0052] In step S2 above, a face generation diffusion model that incorporates a time-aware hybrid expert mechanism and an emotion control mechanism is trained using a course-based learning approach. One possible method is as follows:
[0053] In this embodiment, the face generation diffusion model is implemented based on the flow matching DiT model. A time-aware hybrid expert mechanism and an emotion control mechanism are introduced on the basis of the DiT model. The time-aware hybrid expert mechanism uses a hybrid expert system to replace the feedforward network of the DiT model, and the emotion control mechanism uses an emotion adaptive normalization module to replace the standard normalization layer of the DiT model.
[0054] like Figure 2 As shown, the facial generation diffusion model consists of N DiT blocks. Each DiT contains, in sequence, an emotion adaptive normalization module, a self-attention module, another emotion adaptive normalization module, and a hybrid expert system. The emotion adaptive normalization module can inject facial expression control. Its core includes:
[0055] The time-aware hybrid expert mechanism replaces the feedforward network of each Transformer layer with multiple expert subnetworks. A dynamic router, based on the current inference time step t, selects the active set of experts to achieve division of labor at different stages. Its output is: ,in It is the set of k selected experts. For the gating weights of each expert, For expert sub-networks, These are the input features for the expert subnetwork.
[0056] Emotion control mechanism: The emotion tag embedded in 'e' is injected into the normalization layer of each layer. Specifically, the standard normalization layer is replaced with an emotion-adaptive normalization module to achieve implicit control and driving of emotions.
[0057]
[0058] in, , It is a linear projection layer, where x and e are the input features and sentiment label embeddings of the sentiment adaptive normalization module, respectively. Let L2 norm be the number of input features, and D be the dimension of the input features. This represents the emotion adaptive normalization module.
[0059] Generation strategy: The model adopts a multimodal collaborative generation mechanism, inputting noisy states at each time step. Time step coding Audio features Given a condition vector, output the flow field velocity vector at the current time step. , used to advance to the target distribution.
[0060] The backbone network is a Transformer architecture that integrates a hybrid expert mechanism and a time-aware module, called the MoE-DiT model, which enables cross-modal semantic parsing and generation prediction. Compared with traditional convolutional networks or U-Net structures, this architecture has advantages in modeling long-term temporal dynamics and handling high-dimensional facial expression parameters. It can capture fine-grained changes in facial expressions as they evolve over time, generating more realistic and natural animation effects.
[0061] Specifically, the hybrid expert module uses a time-aware expert router to dynamically activate specific expert subnetworks at different inference stages, enabling the model to flexibly switch between coarse prediction and fine-tuning, thereby improving the adaptability of the generated sequence at different time scales.
[0062] In terms of emotion control, an adaptive root mean square normalization module is used, combined with the injected emotion condition vector, to modulate the normalization process of each layer, achieving control over different emotion categories. Simultaneously, a classifier-free guidance mechanism is employed, discarding emotion regulation vectors with a certain probability during training, and adjusting the guidance strength during the inference phase to achieve tunability of emotion control, enhancing the model's flexibility in expressing different emotional styles.
[0063] Considering the scarcity of sentiment data, this invention designs a two-stage "synchronization-sensitivity" training mechanism. This mechanism first pre-trains on a large-scale audio-video dataset to establish a robust audio-lip movement correspondence, ensuring high-quality audio-lip synchronization in the generated results. Subsequently, it performs sentiment-supervised fine-tuning on a small-scale, high-quality sentiment dataset, enabling the expression and control of multiple sentiment categories, forming a learning path from unimodal driving to multimodal fusion.
[0064] When training the above model, the person identification encoding, driving audio, and emotion label need to be processed separately to obtain person identification feature embedding, driving audio embedding, and emotion label embedding. For example, using the self-supervised pre-trained HuBERT (Hidden-Unit BERT) model as the audio encoder, HuBERT can extract content, prosody, and semantic features from the audio. The input is an aligned video audio segment A, and the output is the audio embedding for each frame. Since the output frame rate of HuberT may not match the target video frame rate, a linear interpolation method is used to align the audio feature time axis to the target video frame rate T. Cross-modal time synchronization preprocessing is then completed, forming a one-to-one cross-modal sequence of "audio frames – animation frames." This processed audio feature sequence will be used as conditional input to the subsequent generative model to drive the synthesis of 3D facial animation.
[0065] To alleviate the problems of scarce emotional data and unstable training, this invention introduces a course-based learning mechanism, which is divided into two stages:
[0066] Phase 1: Synchronous pre-training. For example... Figure 3 As shown, a second dataset without emotion labels is used, with person identity feature embedding and driving audio embedding as input conditions. This minimizes the flow matching loss between the audio-driven generation results and the pseudo-ground values, enabling audio-lip movement synchronization pre-training of the model. This stage strengthens the model's lip movement synchronization capability, serving as the foundation for downstream tasks.
[0067] Phase Two: Emotional Fine-tuning. For example... Figure 4 As shown, using a first dataset containing emotion labels, fine-tuning training is performed with character identity feature embeddings and driving audio embeddings as input conditions, and emotion label embeddings are injected based on an emotion control mechanism. The training objective is also to optimize the stream matching loss, enabling the model to possess emotion recognition, control, and expression capabilities.
[0068] In flow matching, data generation is modeled as a continuous trajectory from a noise distribution (t=1) to the target data (t=0). A face generation diffusion model predicts the velocity vector, and the trained model's predicted velocity vector approximates the true velocity vector. The ODE solver then performs a numerical solution based on the velocity vector, progressively transforming the noise into the target data. The flow matching loss is formulated as follows:
[0069]
[0070] In each training batch, multiple t values are randomly sampled, and noisy samples are constructed for each t. Calculation model prediction speed With real speed The deviation is averaged, and the error across all sampling points is used as the loss. The role of the course is reflected in its progression from general to specific, from broad to specialized, learning synchronous data before emotion data, ensuring the stability and convergence quality of model training.
[0071] In step S3 above, a closed-loop preference optimization mechanism is introduced. One possible approach is as follows:
[0072] To make the model output more closely match the user's subjective perception, this invention introduces a preference optimization enhancement mechanism into the audio-driven 3D face generation task for the first time. First, a preference dataset is constructed, containing generated animations, realistic animations, and preference scores for each animation. A reward estimation model is trained using the preference dataset, which predicts preference scores based on the animations as input. Then, a face generation diffusion model is used to generate face animations in batches. The lip-phone synchronicity scorer and the reward estimation model are used to estimate the lip-phone synchronicity score and the emotional naturalness score, respectively. Win-loss sample pairs with significant score differences are selected from multiple face animations generated from the same driving audio, and the preference optimization loss function is calculated. Finally, the preference optimization loss function and the flow matching loss are combined to update the face generation diffusion model, iterating multiple times.
[0073] In one specific embodiment of the present invention, when constructing the preference dataset, multiple samples are randomly selected from MEAD and RAVDESS; two different 3D facial animations are generated for each audio segment (due to the uncertainty of stream matching, the inference results are different each time); sample triples are constructed by combining real sequences; multiple human annotators are invited to score the "naturalness of expression"; each video is scored by no fewer than 3 people, and the voting results are used as the final label; samples without consensus are excluded; after constructing the dataset, it is used to train a reward estimation model. This replaces manual scoring in subsequent stages. In this embodiment, the reward estimation model... It is implemented using a convolutional neural network.
[0074] The common scoring criteria used by the annotators were:
[0075] 1 point: Poor video quality, extremely uncoordinated lip movements or never opening the lips, obviously wrong facial expressions, and obvious shaking or numerous artifacts in the character.
[0076] 2 points: The video quality is average. There are some deviations in lip movements or they are occasionally unnatural. Expressions are visible but not clear or coherent. Slight shaking or artifacts can be detected but are not significant.
[0077] 3 points: The video quality is average. The lip movements are basically accurate, the facial expressions are fairly natural, and there are occasional jitters or artifacts.
[0078] 4 points: The video quality is good, the lip movements are natural with only minor flaws, the expressions are clear and expressive, and the emotions are accurately conveyed; there are very few distracting jitters or artifacts.
[0079] 5 stars: Excellent video quality, perfect lip synchronization, clear and natural expressions, no jitter or artifacts.
[0080] After constructing the triples, perform the following steps:
[0081] Sample generation: Generate multiple facial animation sequences in batches using the current model;
[0082] Scoring estimation: Two evaluators were used for evaluation. The lip-phonetic synchronization score used the SyncNet model as an automated synchronization detector; the emotional naturalness score used a trained reward model. Make predictions;
[0083] Sample pairing and loss optimization: From multiple results generated by the same driving audio, select "win-loss" sample pairs with significant performance differences, and calculate the preference-optimized loss function as follows:
[0084]
[0085]
[0086]
[0087] The loss function introduces subjective preference guidance signals into the model generation path, enabling the model to evolve towards higher user ratings while maintaining data consistency, thus adapting to the optimization requirements of multimodal feedback signals.
[0088] To prevent the model's output from deviating from the correct samples during preference optimization, a flow matching loss term is added to the loss function:
[0089]
[0090]
[0091] The flow matching loss term is:
[0092]
[0093] in, Indicate the loss of preference optimization; Indicates the parameters of the current facial generation diffusion model. The difference in flow matching loss between the winning and losing sample pairs. This represents the parameters of the previous round of facial generation diffusion model. The difference in flow matching loss between the winning and losing sample pairs. This indicates that the time step t follows a continuous uniform distribution on the interval [0,1]. , Indicates win / loss sample pairs, This represents the scaling factor, used to control the degree to which the differences in loss values are amplified. This represents the sigmoid function, i.e. ; t represents the time step, and They represent respectively to , The result after adding noise at time step t This represents the velocity vector predicted by the facial generation diffusion model. and These represent the true velocity vectors of the winning and losing sample pairs, This represents the square of the L2 norm.
[0094] An optimization objective is constructed by pairing user preference samples (win-loss pairs), guiding the model to update in the streaming matching space in a direction that better matches user preferences.
[0095] Multiple rounds of updates: training the model for After that, the results generated by the new model are scored by the user, the reward model r is updated, and the next round of training begins, repeating for 3 rounds.
[0096] The model training objectives encompass minimizing lip movement synchronization error and maximizing the naturalness of emotional expression. It also constructs win-loss pairs for facial expression naturalness and win-loss pairs for lip-phoneme alignment, employing multi-task loss joint optimization to enhance the model's balance between expressiveness and controllability. This mechanism ensures that each round of model updates evolves in a direction more favored by the user, maximizing subjective satisfaction.
[0097] In step S4 above, the optimized facial generation diffusion model is used to generate facial animations that meet the preference conditions.
[0098] like Figure 5 As shown, character identity feature embedding and driving audio embedding are used as input conditions, and emotion label embedding is injected based on emotion control mechanism. The facial generation diffusion model is iteratively optimized to gradually denoise and generate a speaker sequence with expressions that meets the input conditions.
[0099] The above method will be applied to the following embodiments to demonstrate the technical effects of the present invention. The specific steps in the embodiments will not be repeated.
[0100] This invention was tested on three datasets: HDTF, MEAD, and RAVDESS. HDTF is a high-resolution audio-video dataset, while MEAD and RAVDESS are emotional audio-video datasets used to test the quality of generated emotionally charged speakers. To objectively evaluate the performance of this invention, LVE (Lip Corner Error) and LSE-D / LSE-C (Lip-Speech Alignment Distance, Confidence) were used to evaluate lip alignment on the selected test sets, and VE-FID and the reward model score EmoScore were used to evaluate the naturalness of facial expressions. A comparison was made with the following existing models:
[0101] Comparison Method 1: 3D speaker generation models without facial expression control, including FaceFormer, CodeTalker, FaceDiffuser, and UniTalker. These generative models primarily aim to achieve lip alignment, but their relatively simple structure results in poor performance when trained on large-scale datasets.
[0102] Comparison Method 2: 3D speaker generation models with facial expression control, including EMOTE and ProbTalk3D. These models achieve lip-syncing and controllable facial expression generation, similar to the application scenario of this invention. However, they are not optimized for user preferences, and the generated quality often does not meet audience preferences in actual applications.
[0103] The experimental results obtained by following the steps described in the specific implementation method are shown in Tables 1 to 3:
[0104] Table 1: Test results of lip alignment for the HDTF dataset in this invention
[0105]
[0106] Table 2: Test results of lip alignment and facial expression naturalness for the RAVDESS dataset in this invention.
[0107]
[0108] Table 3: Test results of lip alignment and facial expression naturalness for the MEAD dataset in this invention
[0109]
[0110] As can be seen from Tables 1, 2, and 3, the three-dimensional speaker sequences generated by this invention outperform previous state-of-the-art methods in terms of lip-phonetic synchronization. The LVE index improves from 3.65 to 3.21, 5.23 to 5.05, and 4.41 to 4.31, respectively, while the LSE-D index decreases from 11.324 to 11.209, 9.598 to 9.512, and 10.754 to 10.631, respectively. Benefiting from the two-stage learning process and improved generation framework introduced in this invention, more consistent audio-lip alignment is achieved.
[0111] As shown in Tables 2 and 3, this invention achieves superior performance in terms of the naturalness of generated expressions and user preferences. On the RAVDESS dataset, the VE-FID index decreased from 29.75 to 27.67, while the reward model score (EmoScore) increased from 3.72 to 4.12 and from 3.62 to 4.07, respectively. Benefiting from the closed-loop preference optimization mechanism proposed in this invention, the model's generated results are more aligned with user preferences, significantly reducing the probability of low-quality generated results.
[0112] It is worth noting that all experimental results in Tables 1 to 3 are based on the same pre-trained model of this invention, which demonstrates the excellent generative capability of the model of this invention.
[0113] Based on the same inventive concept, this embodiment also proposes a cross-modal learning preference optimization enhancement 3D face generation system, including:
[0114] The training dataset acquisition module is used to acquire a first dataset containing emotion labels and a second dataset not containing emotion labels. Both datasets contain facial image sequences and their corresponding driving audio. The emotion labels are parameter sets that reflect the expression state of each frame of the image.
[0115] The course learning module is used to train a face generation diffusion model that incorporates a time-aware hybrid expert mechanism and an emotion control mechanism through a course learning approach. During the course learning process, the model is first pre-trained with audio-lip movement synchronization using the second dataset, with the person's identity feature embedding and the driving audio embedding as input conditions. Then, the model is fine-tuned with emotion driving using the first dataset, with the person's identity feature embedding and the driving audio embedding as input conditions, and emotion label embedding is injected based on the emotion control mechanism.
[0116] The preference optimization enhancement module is used to introduce a closed-loop preference optimization mechanism. It predicts the emotional naturalness score of the facial animation generated by the facial generation diffusion model through the reward estimation model, generates winning and losing sample pairs by combining the lip-sync evaluation tool, and updates the facial generation diffusion model by combining the preference optimization loss function and the flow matching loss. It iterates for multiple rounds to achieve preference optimization enhancement.
[0117] The 3D face generation module is used to generate facial animations that meet preference conditions using an optimized face generation diffusion model.
[0118] For the system embodiments, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments; the implementation methods of the remaining modules will not be repeated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0119] The system embodiments of the present invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The system embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution.
[0120] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A cross-modal learning-based preference optimization-enhanced 3D face generation method, characterized in that, include: (1) Obtain a first dataset containing emotion labels and a second dataset not containing emotion labels. Both datasets contain facial image sequences and their corresponding driving audio. The emotion labels are parameter groups that reflect the expression state of each frame of the image. (2) A face generation diffusion model with temporal awareness hybrid expert mechanism and emotion control mechanism is trained by course learning. During the course learning process, the model is first pre-trained with audio-lip movement synchronization using the second dataset. During the pre-training, the embedding of character identity features and the driving audio embedding are used as input conditions. The model is then fine-tuned using the first dataset. During fine-tuning, character identity feature embeddings and driving audio embeddings are used as input conditions, and emotion label embeddings are injected based on an emotion control mechanism. The emotion control mechanism refers to replacing the standard normalization layer of the DiT model with an emotion adaptive normalization module. The calculation formula for the emotion adaptive normalization module is as follows: ; in, , It is a linear projection layer, where x and e are the input features and sentiment label embeddings of the sentiment adaptive normalization module, respectively. Let L2 norm be the number of input features, and D be the dimension of the input features. This indicates the emotion adaptive normalization module; (3) Introduce a closed-loop preference optimization mechanism, predict the emotional naturalness score of the facial animation generated by the facial generation diffusion model through the reward estimation model, generate winning and losing sample pairs by combining the lip-sync evaluation tool, and update the facial generation diffusion model by combining the preference optimization loss function and the flow matching loss, iterating multiple times to achieve preference optimization enhancement; The preference optimization loss function is as follows: ; in, Indicate the loss of preference optimization; Indicates the parameters of the current facial generation diffusion model. The difference in flow matching loss between the winning and losing sample pairs. This represents the parameters of the previous round of facial generation diffusion model. The difference in flow matching loss between the winning and losing sample pairs. This indicates that the time step t follows a continuous uniform distribution on the interval [0,1]. , Indicates win / loss sample pairs, Indicates the scaling factor. Let E represent the sigmoid function and E represent the expectation. (4) Use the optimized facial generation diffusion model to generate facial animations that meet the preference conditions.
2. The cross-modal learning-based preference optimization enhancement method for 3D face generation according to claim 1, characterized in that, The facial frame images in both datasets are three-dimensional scans of the face. The parameter set reflecting the expression state of each frame image includes 50-dimensional facial expression parameters and 3-dimensional jaw pose parameters from the three-dimensional scan images.
3. The cross-modal learning-based preference optimization enhancement method for 3D face generation according to claim 1, characterized in that, The facial generation diffusion model is implemented based on the flow matching DiT model. A time-aware hybrid expert mechanism and an emotion control mechanism are introduced on the basis of the DiT model. The time-aware hybrid expert mechanism uses a hybrid expert system to replace the feedforward network of the DiT model.
4. The cross-modal learning-based preference optimization-enhanced 3D face generation method according to claim 3, characterized in that, The hybrid expert system includes multiple expert subnetworks and a dynamic router. The dynamic router activates several expert subnetworks and assigns weights at each inference time step, and outputs the weighted result of the expert subnetworks.
5. The cross-modal learning-based preference optimization enhancement method for 3D face generation according to claim 1, characterized in that, Step (3) includes: (3.1) Construct a preference dataset containing generated animations, real animations and their respective preference ratings, and train a reward estimation model using the preference dataset. The reward estimation model takes the animation as input and predicts the preference rating. (3.2) Use the facial generation diffusion model to generate facial animations in batches. Use the lip-phonetic synchronization scorer and the reward estimation model to estimate the lip-phonetic synchronization score and the emotional naturalness score, respectively. Select the win-loss sample pairs with significant score differences from multiple facial animations generated by the same driving audio, and calculate the preference optimization loss function. (3.3) Combine the preference optimization loss function and the flow matching loss to update the face generation diffusion model.
6. The cross-modal learning-based preference optimization enhancement method for 3D face generation according to claim 5, characterized in that, The process of constructing the preference dataset is as follows: Given audio and emotion tags for a batch of real animations, the face generation diffusion model after two courses is run twice to obtain two generated animations for each real animation; The real animation and its two corresponding generated animations are combined into a triplet. The annotator assigns a preference score to each animation in different triplets, resulting in a preference dataset consisting of the animations and preference scores.
7. A cross-modal learning preference optimization enhancement 3D face generation system, used to implement the cross-modal learning preference optimization enhancement 3D face generation method of claim 1, characterized in that, The system includes: The training dataset acquisition module is used to acquire a first dataset containing emotion labels and a second dataset not containing emotion labels. Both datasets contain facial image sequences and their corresponding driving audio. The emotion labels are parameter sets that reflect the expression state of each frame of the image. The course learning module is used to train a face generation diffusion model that incorporates a time-aware hybrid expert mechanism and an emotion control mechanism through a course learning approach. During the course learning process, the model is first pre-trained with audio-lip movement synchronization using the second dataset, with the person's identity feature embedding and the driving audio embedding as input conditions. Then, the model is fine-tuned with emotion driving using the first dataset, with the person's identity feature embedding and the driving audio embedding as input conditions, and emotion label embedding is injected based on the emotion control mechanism. The preference optimization enhancement module is used to introduce a closed-loop preference optimization mechanism. It predicts the emotional naturalness score of the facial animation generated by the facial generation diffusion model through the reward estimation model, generates winning and losing sample pairs by combining the lip-sync evaluation tool, and updates the facial generation diffusion model by combining the preference optimization loss function and the flow matching loss. It iterates for multiple rounds to achieve preference optimization enhancement. The 3D face generation module is used to generate facial animations that meet preference conditions using an optimized face generation diffusion model.
8. The cross-modal learning preference optimization enhancement 3D face generation system according to claim 7, characterized in that, The facial generation diffusion model is implemented based on the flow matching DiT model. A time-aware hybrid expert mechanism and an emotion control mechanism are introduced on the basis of the DiT model. The time-aware hybrid expert mechanism uses a hybrid expert system to replace the feedforward network of the DiT model, and the emotion control mechanism uses an emotion adaptive normalization module to replace the standard normalization layer of the DiT model.