Method for generating mouth shape expression through audio of large model based on emotion reference and discrimination
By combining multimodal feature extraction from speech, text, and emotional reference videos, the generative model can more accurately express complex emotional states, solving the problem that a single emotion category label cannot express complex emotions and improving the accuracy of expression generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING JULI DIMENSION TECH CO LTD
- Filing Date
- 2025-12-15
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, relying on a single emotion category label cannot effectively express complex emotional states, resulting in low accuracy in generating expressions in complex emotional scenarios.
By acquiring audio data, text data, and emotional reference videos, and utilizing large-scale speech models, large-scale language models, and facial parameter generation models, multimodal features are extracted and facial parameter predictions are performed to generate a composite emotional feature space, enabling emotion discrimination and content discrimination, and generating corresponding facial lip movements.
It improves the accuracy of facial expression generation in complex emotional scenarios, enabling more accurate expression of complex emotional states.
Smart Images

Figure CN121962371A_ABST
Abstract
Description
A method for generating lip-sync expressions from audio using a large model based on emotion reference and discrimination. Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method for generating lip-sync expressions from audio based on a large model of emotion reference and discrimination. Background Technology
[0002] In recent years, speech-driven lip-sync generation has received increasing research and widespread attention, especially in the field of 3D digital humans. Facial expressions and lip-sync based on speech and text are popular in animation, games, virtual reality, and short video production and service industries. As a result, there are increasingly higher requirements for the richness, realism, subtlety, and emotional tension of virtual humans.
[0003] Research shows that human emotions can be categorized into several major types, such as happiness, sadness, anger, fear, disgust, surprise, and anticipation. However, in reality, human emotions are often not singular, and relying solely on a single emotional label cannot express complex emotional states. Furthermore, the accuracy of facial expression generation in complex emotional scenarios is low. Summary of the Invention
[0004] This application provides a method for generating lip-sync expressions from audio based on a large model of emotion reference and discrimination, in order to alleviate or solve one or more technical problems existing in the prior art.
[0005] In a first aspect, embodiments of this application provide an audio-based lip-syncing expression generation method based on a large-scale model of emotion reference and discrimination, comprising: acquiring audio data and corresponding text data, emotion category labels, and emotion reference videos; extracting features from the audio data using a large-scale speech model to obtain audio features, and extracting features from the text data using a large-scale language model to obtain text features; extracting features from the emotion category labels and the emotion reference videos to obtain emotion features and emotion reference features; using a facial parameter generation model to predict facial parameters based on the audio features, text features, emotion features, and emotion reference features to obtain predicted facial parameters; the facial parameter generation model is a large-scale model obtained by performing emotion discrimination and content discrimination based on historically predicted facial parameters and training the model according to the corresponding discrimination results; and generating facial lip-syncing expressions based on the predicted facial parameters.
[0006] Secondly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the methods of embodiments of this application when executing the computer program.
[0007] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of any one of the embodiments of this application.
[0008] Fourthly, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, implements any of the methods described in the embodiments of this application.
[0009] According to the method of this application embodiment, audio features are extracted from audio data using a large speech model, text features are extracted from the corresponding text data using a large language model, emotional features are extracted from the emotional category labels corresponding to the audio data, and emotional reference features are extracted from emotional reference features. A facial parameter generation model then predicts facial parameters based on the audio features, text features, emotional features, and emotional reference features to obtain predicted facial parameters. Facial lip-sync expressions are then generated based on the predicted facial parameters. According to this method, the facial parameter generation model performs emotional and content discrimination based on historically predicted facial parameters, and a large model is obtained by training the model based on the corresponding discrimination results. Utilizing the facial reference generation model, facial parameter prediction is performed based on speech, semantic, emotional features, and emotional reference features. This facilitates the construction of a composite emotional feature space by combining multimodal features. Based on this composite emotional feature space, it is beneficial to predict facial parameters expressing complex emotional states and generate corresponding facial lip-sync expressions, effectively solving the technical problem that traditional single emotional category labels cannot express complex emotional states, thereby improving the accuracy of expression generation in complex emotional scenarios.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0011] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this application and should not be construed as limiting the scope of this application.
[0012] Figure 1 shows a flowchart of the audio lip-syncing generation method based on a large model of emotion reference and discrimination according to an embodiment of this application; Figure 2 shows a detailed flowchart of the audio lip-syncing generation method based on a large model of emotion reference and discrimination according to an exemplary embodiment of this application; Figure 3 shows a schematic diagram of the training process of the emotion reference feature extraction network according to an exemplary embodiment of this application; Figure 4 shows a schematic diagram of the module structure of the emotion reference feature extraction network according to an exemplary embodiment of this application; Figure 5 shows a flowchart of the overall network module structure of an exemplary embodiment of this application; Figure 6 shows a schematic diagram of the processing method for multiple emotion reference videos according to an exemplary embodiment of this application; Figure 7 shows a schematic diagram of the structure of the audio lip-syncing generation device based on a large model of emotion reference and discrimination according to an embodiment of this application; Figure 8 shows a block diagram of the electronic device provided in an embodiment of this application. Detailed Implementation
[0013] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the concept or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0014] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.
[0015] In real-world scenarios, human emotions generally fall into categories such as: happiness, sadness, anger, fear, disgust, surprise, and anticipation. However, human emotions are rarely singular; rather, they manifest in varying degrees, such as happiness, relief, ecstasy, or schadenfreude; or they can be complex emotions derived from a combination of different underlying emotions, such as a mixture of joy and sorrow, or a tangled web of love and hate. Relying solely on a single emotional label is insufficient to fully express the complete content of an emotion.
[0016] It should be noted that the application scenarios or examples provided in the embodiments of this application are for ease of understanding, and the embodiments of this application do not specifically limit the application of the technical solutions. In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0017] The technical solution of this application and how it solves the aforementioned technical problems are described in detail below with specific embodiments. The listed specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0018] Figure 1 shows a flowchart of an audio lip-syncing method based on a large model of emotion reference and discrimination according to an embodiment of this application. As shown in Figure 1, the method may include steps S101 to S105.
[0019] S101, acquire audio data and corresponding text data, sentiment category labels and sentiment reference videos.
[0020] S102: Extract features from audio data using a large speech model to obtain audio features; extract features from text data using a large language model to obtain text features.
[0021] S103, extract features from emotion category labels and emotion reference videos to obtain emotion features and emotion reference features.
[0022] S104 utilizes a facial parameter generation model to predict facial parameters based on audio features, text features, sentiment features, and sentiment reference features. The facial parameter generation model performs sentiment and content discrimination based on historically predicted facial parameters, and trains the model based on the corresponding discrimination results to obtain a large model.
[0023] S105 generates facial lip movements based on predicted facial parameters.
[0024] According to the method of this application embodiment, audio features are extracted from audio data using a large speech model, text features are extracted from the corresponding text data using a large language model, emotional features are extracted from the emotional category labels corresponding to the audio data, and emotional reference features are extracted from emotional reference features. A facial parameter generation model then predicts facial parameters based on these audio features, text features, emotional features, and emotional reference features to obtain predicted facial parameters. Facial lip movements are then generated based on these predicted facial parameters. According to this method, the facial parameter generation model performs emotional and content discrimination based on historically predicted facial parameters, and a large model is obtained by training the model based on the corresponding discrimination results. Utilizing the facial reference generation model, which predicts facial parameters based on speech, semantics, emotional features, and emotional reference features, facilitates the construction of a composite emotional feature space by combining multimodal features. Based on this composite emotional feature space, it is beneficial to predict facial parameters expressing complex emotional states and generate corresponding facial lip movements. This effectively solves the technical problem that traditional single emotional category labels cannot express complex emotional states, thereby improving the accuracy of expression generation in complex emotional scenarios.
[0025] In step S101, the audio data and text data can be speech and annotations or subtitles from the same video content, or audio data and text data from the same speech segment. The emotion category label is used to mark the emotion-related classifications made by humans after understanding the video content or speech segment.
[0026] As an example, based on training needs, emotion category labels are divided into multiple major categories, each containing multiple subcategories. For instance, emotion labels comprise 7 major categories and 33 subcategories. The 7 major categories are: happiness, sadness, anger, fear, disgust, surprise, and anticipation. Based on the aforementioned emotion category labels, the 7 major categories are further subdivided as follows.
[0027] For example, happiness is subdivided into: comfort, joy, pleasure, satisfaction, happiness, excitement, ecstasy, and frenzy; sadness is subdivided into: melancholy, sorrow, loss, grief, and despair; anger is subdivided into: irritability, hostility, indignation, rage, and fury; fear is subdivided into: unease, timidity, anxiety, and dread; disgust is subdivided into: aversion, disgust, aversion, and contempt; surprise is subdivided into: astonishment, bewilderment, surprise, and shock; and expectation is subdivided into: hope, longing, and eager anticipation.
[0028] As an example, the 7 major categories and 33 subcategories of sentiment labels can be numerically encoded, such as one-hot encoding or label encoding. For example, the 7 major categories and 33 subcategories of sentiment can be treated as two independent namespaces, each encoded using one-hot encoding, and then concatenated.
[0029] In this embodiment, each major category can be referred to as a first-level emotion category, and the corresponding emotion tag is a first-level emotion tag, that is, a first-level labeling of the emotion. The subcategories under each major category can be referred to as second-level emotion categories, and the corresponding emotion tags are second-level emotion tags, that is, second-level labeling of the emotion.
[0030] Based on the data types of the aforementioned emotion categories, emotion labeling is a classification made by humans after understanding the video. For example, the broad category might be sadness, and the subcategory might be loss. For data where it is difficult to clearly distinguish between secondary emotion categories, only the primary emotion category needs to be labeled.
[0031] Emotional reference videos are video data containing emotional expressions used to assist model training and feature extraction. Their role is to provide a multimodal reference standard for emotion recognition. For example, emotional reference videos typically contain real-life facial expressions or scene clips with known emotional labels (such as happiness, ecstasy, anger, rage, disgust, etc.), providing learning samples for the model. Sources of emotional reference videos include, but are not limited to, at least one of the following: movies, TV series, speech videos, and social media videos, as long as they contain emotional expressions from individuals.
[0032] As an example, an emotional reference video can be a pre-processed segment of the original reference video that expresses the emotion. Pre-processing includes cleaning and cropping. Cleaning includes: removing segments irrelevant to the emotional expression, such as black screens, stutters, repetitions, or blurry images; adjusting the brightness and / or contrast of the video footage, etc. Cropping is to ensure the integrity of each emotional segment, for example, retaining a few seconds (1-2 seconds) of video content before the emotion occurs, to ensure the continuity of the emotional expression and avoid misinterpretation.
[0033] In some embodiments, in step S102 above, features can be extracted from text data using a Large Language Model (LLM) to obtain text features. As examples, besides the Large Language Model, text features can also be extracted using word embeddings (such as Word2Vec and GloVe), pre-trained models such as Bidirectional Encoder Representations from Transformers (BERT), and Generative Pre-trained Transformers (GPT). The core architecture of the GPT model is the Transformer. In the field of text processing, applying the GPT model for text feature extraction allows for the consideration of information from other words when processing each word, obtaining the dependencies between contexts and providing high-quality features for subsequent processing.
[0034] In step S102, audio features can be extracted using a Hidden-Unit BERT (HuBERT) speech model. As an example, besides the Hidden-Unit BERT, audio features can also be extracted using a self-supervised learning-based audio pre-trained model (wav2vec). Audio features include both lexical and non-lexical features. Lexical features include acoustic parameters directly related to specific pronunciation content, such as phonemes, syllables, or words. Non-lexical features include speech attributes such as intonation, speech rate, and intensity. These non-lexical attributes can be used to more accurately predict sentiment.
[0035] As an example, the input sentiment category label can be encoded into the latent space using one-hot encoding and an embedding layer, resulting in a low-dimensional continuous vector called the sentiment feature.
[0036] In some embodiments, the step of extracting features from the emotion reference video in step S103 above may specifically include: for any emotion reference video among the emotion reference videos, extracting features from the emotion reference video to obtain the emotion reference features of the emotion reference video; performing weighted fusion on the emotion reference features of each emotion reference video, and using the weighted fused features as the emotion reference features.
[0037] For example, the emotion feature extraction network supports inputting one or more emotion reference videos. Each emotion reference video is input into the emotion feature extraction network, which extracts features from each video to obtain an emotion reference feature code (feature vector). If multiple emotion reference videos exist, multiple emotion reference feature codes are obtained. These codes are then weighted and summed according to a certain ratio to obtain a weighted fused emotion reference feature code, which serves as the extracted emotion reference feature. For instance, the emotion feature extraction network supports inputting two emotion reference videos. After passing through the network, two feature vectors are obtained, which are then superimposed in a certain ratio to obtain a composite emotion, such as "tears of joy" or "a mixture of laughter and tears."
[0038] It should be understood that the weights of multiple sentiment reference feature encodings can be customized according to actual conditions, and the embodiments of this application do not impose specific limitations.
[0039] In this embodiment, features of each emotion reference video are extracted one by one, and then the features of multiple videos are weighted and fused. This is beneficial to combine the features of multiple emotion reference videos and improve the robustness and generalization ability of subsequent facial parameter prediction.
[0040] In some embodiments, the step of extracting features from the emotional reference video to obtain the emotional reference features of the emotional reference video can specifically be as follows: using the visual branch network and the audio branch network of the emotional feature extraction network to extract features from the emotional reference video respectively, to obtain the extracted visual emotional features and acoustic emotional features; and performing multimodal fusion of the visual emotional features and the acoustic emotional features to obtain the emotional reference features.
[0041] For example, the emotion feature extraction network is a dual-branch network: a visual branch and an audio branch. The visual branch is used to capture emotion-related information based on the content of consecutive frames, and the audio branch is used to obtain speech emotion based on audio features and speech content.
[0042] For example, when fusing visual and audio dual-branch features using multimodal fusion, a cross-attention mechanism can be used to perform cross-attention calculations on the extracted visual and acoustic emotion features. For instance, a cross-attention model can be used; specifically, this model adopts a query-key-value (QKV) pattern, where visual emotion features are used as the query vector, and acoustic emotion features are used as the key and value vectors. Multi-head attention is used to achieve modal interaction, and finally, a fully connected layer is used to reduce the dimensionality to 256 dimensions, providing an efficient and compact feature representation for subsequent tasks and helping to control the computational load and memory consumption of the model.
[0043] In this embodiment, cross-attention calculation is performed on visual and acoustic emotion features through multi-head cross-attention, which can establish data association between the two and establish a dependency relationship between them. This is beneficial for accurately obtaining the similarity between the two, so that the relationship between different modalities of image frames and spectrograms can promote and complement each other, thereby better performing feature fusion between multiple modalities, effectively capturing cross-modal emotional correlation, and improving the accuracy of the extracted emotional reference features.
[0044] In some embodiments, the steps of using the visual branch network and audio branch network of the emotion feature extraction network to extract features from the emotion reference video to obtain extracted visual emotion features and acoustic emotion features may specifically include: in the visual branch network, extracting spatiotemporal features from the emotion reference video, using a fast branch to extract dynamic detail features from the spatiotemporal features, using a slow branch to extract semantic features from the spatiotemporal features, and performing feature fusion on the dynamic detail features and semantic features to obtain visual emotion features; in the audio branch network, extracting a spectrogram from the emotion reference video, extracting local features from the spectrogram through convolutional layers and pooling layers, and using an encoder to extract long temporal dependencies from the local features to obtain acoustic emotion features based on long temporal dependencies.
[0045] As an example, the visual branch network comprises a backbone network and a feature fusion network. The backbone network uses a video action recognition model (SlowFast ResNet-50-3DCNN) to extract spatiotemporal features. This model employs a dual-path architecture: the slow pathway uses a low sampling rate and a high number of channels to capture semantic information at a low frame rate, such as 4-8 FPS; the fast pathway uses a high sampling rate and a low number of channels to capture dynamic details at a high frame rate, such as 16-32 FPS.
[0046] As an example, the input to the audio branch network can be the Log-Mel spectrogram of an emotion reference video. The audio branch network can include multiple convolutional layers and one max-pooling layer to extract local features, and finally, a Transformer encoder is used to model long-term temporal dependencies. For example, multiple convolutional layers could be four layers, used to extract local features of the audio signal contained in the emotion reference video layer by layer. The convolutional layers can efficiently capture local patterns and structural information in the audio data by using sliding convolution kernels. The max-pooling layer is used to reduce the dimensionality of the feature map, reducing computation while retaining the most important feature information. As an example, the Transformer encoder can effectively model long-term temporal dependencies in the audio signal.
[0047] As an example, visual and acoustic emotion features can be fused using average pooling. Specifically, the feature maps of the two branches are reduced in dimensionality using global average pooling to obtain their respective fixed-length feature vectors, which are then fused through weighted summation or concatenation.
[0048] In this embodiment, the visual branch extracts dynamic details and semantic features through fast and slow branches respectively, and then fuses them to obtain visual sentiment features; the audio branch extracts local features and long-term temporal dependencies from the spectrogram to obtain acoustic sentiment features. This provides a data foundation for subsequent multimodal feature fusion.
[0049] In some embodiments, the steps described above for predicting facial parameters based on audio features, text features, emotion features, and emotion reference features using a facial parameter generation model to obtain predicted facial parameters may specifically include: performing feature fusion processing using the feature extraction network of the facial parameter generation model. The feature fusion processing includes the following steps: performing spatial attention calculation on facial parameters at each historical time point to obtain spatial attention features of the facial parameters; performing causal multi-head self-attention calculation on the spatial attention features of the facial parameters to obtain spatiotemporal fusion attention features; fusing audio features, text features, and emotion features to obtain multimodal fusion features; performing cross-multi-head attention calculation on the spatiotemporal fusion attention features and the multimodal fusion features to obtain cross-attention features of the multimodal fusion; performing cross-multi-head attention calculation on the cross-attention features of the multimodal fusion and the emotion reference features to obtain emotion-enhanced multimodal fusion features; and decoding the multimodal fusion features using the decoder of the facial parameter generation model to obtain the predicted facial parameters.
[0050] As an example, the input to the facial parameter generation model includes facial parameters up to historical time points t1:t2-1, and the output is the facial parameters at the current time t2. t1:t2-1 refers to all historical time points from time t1 to the end of time t2-1. t2-1 is the time point preceding t2 (also called a time point). Accordingly, the output of the facial parameter generation model is the facial parameters at time t2. t1:t2-1 can be understood as all historical time points before the current time, and t2 is the current time.
[0051] As an example, the facial parameter generation model, as the generator body, can be an autoregressive generation model. The autoregressive model models the time series by representing the value at the current time point as a linear combination of the values at several past time points. The input is the facial parameters at each time point from t1 to t2-1 in history, and the output is the facial parameters at the current time t2.
[0052] For example, facial parameters are four-dimensional facial parameters that include facial features and motion parameters of other facial regions. As an example, facial expressions can include various facial movements and states, which can be simply referred to as facial expressions in this embodiment. Examples include raising eyebrows, frowning, widening or squinting eyes, and raising or lowering the corners of the mouth. Different regions of facial expressions (such as eyes, mouth, eyebrows, etc.) have different importance in expressing emotions. For example, the degree of eye opening and closing and the degree of mouth raising are very important in expressing happiness. To enable the facial parameter generation model to adaptively focus on key spatial regions in the input parameter sequence, a spatial attention mechanism can be introduced to assign different weights to different regions. For example, a weight matrix can be generated, where each element represents the importance weight of the corresponding facial region. The facial parameter generation model extracts feature maps from the input facial parameter sequence, which contain feature information of each facial region. Multiplying the weight matrix element-wise with the original feature map can add feature responses of key regions, enhancing the ability to process important facial expressions.
[0053] As an example, spatial attention is calculated on facial parameters at each historical moment to obtain spatial attention features of the facial parameters. Spatial attention calculation is used to obtain important spatial information in the facial parameters at each moment, such as key expression regions. Causal multi-head self-attention calculation is then performed on the spatial attention features of the facial parameters to obtain spatiotemporally fused attention features. Causal multi-head self-attention calculation is used to capture the dynamic changes of facial lip movements and expressions over time while maintaining causal relationships, that is, the current moment depends only on information from previous moments.
[0054] In this embodiment, to strictly adhere to causal constraints, a causal multi-head self-attention mechanism is employed to perform causal multi-head self-attention computation on the spatial attention features of facial parameters. Its core function is to capture the temporal dependencies in the facial parameter sequence data; that is, the features at the current moment can only depend on the facial parameter features at historical moments, and cannot utilize information from future moments. Causality is ensured through causal masking. Furthermore, multi-head parallel computation can learn diverse causal association patterns from different subspaces, improving the model's hierarchical modeling of facial expression features. For example, one subspace focuses on periorbital muscle movement, while another focuses on lip movements; one subspace focuses on global motion, while another focuses on local motion; one subspace analyzes dynamic changes in facial expressions, while another focuses on the correlation between semantics, etc.
[0055] As an example, feature fusion of audio features, text features, and sentiment features can be achieved through fully connected layers or by concatenation. Integrating information from different modalities provides a foundation for subsequent cross-attention calculations.
[0056] As an example, when performing cross-multi-head attention computation on spatiotemporal fusion attention features and multimodal fusion features, the spatiotemporal fusion attention features can be used as the query, and the multimodal fusion features can be used as the key and value. The cross-multi-head attention mechanism can be used to process them, allowing facial lip-shape expression features to interact with multimodal features. The resulting multimodal fusion cross-attention features can further enhance the expressive power of the features.
[0057] As an example, when performing cross-head attention computation on the cross-attention features and sentiment reference features of multimodal fusion, the cross-attention features of multimodal fusion can be used as the query, and the sentiment reference features as the key and value. The cross-head attention mechanism can be used to further utilize the sentiment reference features and enhance the sentiment-related information. The resulting sentiment-enhanced multimodal fusion features can make the final features closer to the target sentiment.
[0058] In this embodiment, spatial attention and causal multi-head self-attention calculations are performed on facial parameters from historical moments to extract spatiotemporally fused attention features. These features are then fused with audio features, text features, emotional features, and positive emotional reference features through multiple rounds of cross-multi-head attention, ultimately yielding emotionally enhanced multimodal fusion features for predicting facial parameters at the current moment. This process fully utilizes multimodal and time-series information, which is beneficial for more accurately predicting facial lip movements that match emotional expression.
[0059] In some embodiments, the training process of the sentiment feature extraction network includes: performing dimensionality reduction on the sentiment reference features to obtain dimensionality-reduced sentiment reference features; predicting the first-level sentiment category and the second-level sentiment category of the dimensionality-reduced sentiment reference features respectively to obtain a first-level predicted category and a second-level predicted category, wherein the second-level sentiment category is a subcategory of the first-level sentiment category; constructing a first loss function based on the first-level predicted category and the first-level label category; constructing a second loss function based on the second-level predicted category and the second-level label category; constructing a joint loss function according to the first loss function and the second loss function; and updating the model parameters of the sentiment feature extraction network using the joint loss function to train the model of the sentiment feature extraction network to obtain the trained sentiment feature extraction network.
[0060] As an example, sentiment reference features can be dimensionality reduced using fully connected layers. Dimensionality reduction reduces computational complexity, making features more compact and efficient, thus facilitating efficient execution of subsequent classification tasks.
[0061] As an example, in the training method of the sentiment feature extraction network in this application embodiment, the loss function can be cross-entropy loss, which includes two loss functions, one for each of the two levels of classification labels. The cross-entropy loss corresponding to the first-level label category can be applied to all samples, while the cross-entropy loss corresponding to the second-level label category is applied to samples containing the second-level label.
[0062] As an example, the joint loss function can be expressed using the following formula: (1) Among them, It is the first loss function, also known as the first-level label cross-entropy loss (i.e., first-level loss) for all samples. It is the second loss function, also known as the second-level label cross-entropy loss (i.e., second-level loss) applied to samples with second-level labels.
[0063] As an example, the loss weights in formula (1) above can be expressed using adaptive weights as follows: (2) Among them This represents the number of samples including secondary labels. This represents the total number of samples.
[0064] During the training phase, each sample can have multiple levels of label categories. For example, the primary label can be any of the seven major emotion categories such as "happiness" or "sadness," while the secondary label could be more granular emotion categories under "happiness," such as "comfort," "joy," or "pleasure." In actual data annotation, some samples may only be labeled with primary labels without secondary labels, resulting in missing secondary labels.
[0065] As an example, dynamic masking is a mechanism that dynamically generates masks during training to label which samples have secondary labels and which do not. When calculating the second loss function, the mask value of each sample is multiplied by the corresponding loss value. Thus, the loss value for samples without secondary labels (mask value of 0) is set to 0. Therefore, through the dynamic masking mechanism, the model can automatically ignore samples lacking secondary labels during training, ensuring that the calculation of the second loss function is based only on samples with complete labels. This achieves the effect of automatically filtering out unlabeled samples and improves the model's prediction accuracy for secondary label categories.
[0066] In this embodiment, after dimensionality reduction of the emotion reference features, predictions are made for the first-level label category and the second-level label category, respectively. A two-level loss function is constructed and jointly optimized, enabling the emotion feature extraction network to learn both coarse-grained and fine-grained emotion categories simultaneously, thereby improving the accuracy and robustness of emotion recognition.
[0067] In some embodiments, after the step of obtaining the predicted facial parameters described above, the method further includes: using the predicted facial parameters to determine emotion and content, obtaining emotion determination results and content determination results; determining emotion determination loss based on the emotion determination results and predetermined emotion determination labels; determining content determination loss based on the content determination results and predetermined content determination labels; generating a comprehensive loss based on the emotion determination loss and content determination loss; and updating the model parameters of the facial parameter generation model based on the comprehensive loss to train the facial parameter generation model, thereby obtaining a trained facial parameter generation model.
[0068] As an example, the sentiment discrimination result refers to the emotion category expressed by the facial lip-sync expression predicted by the model. The predefined sentiment discrimination label refers to the true emotion category of the sample. The sentiment discrimination loss is used to characterize whether the emotion expressed by the facial lip-sync expression generated based on the predicted facial parameters is consistent with the target emotion, in order to measure the accuracy of sentiment prediction. For example, if the predefined sentiment discrimination label indicates that the emotion expressed by the audio is "excitement," but the sentiment discrimination result is that the emotion category expressed by the facial lip-sync expression predicted by the model is "sadness," then the sentiment loss will be high because the emotion expression is inconsistent.
[0069] As an example, the content discrimination result refers to the detail of the content expressed by the predicted facial lip-sync expression. The predefined content discrimination label is used to characterize: the true content detail of the sample. Content details include, but are not limited to, lip shape, micro-expressions, etc. The content discrimination loss is used to characterize: whether the generated facial lip-sync expression matches the target content in specific details. For example: if the predefined content discrimination label indicates that the spoken content is "ooh," but the content discrimination result is: the facial parameters predicted by the model make the mouth look like it's saying "ah," this indicates that the lip shape in the content detail is completely incorrect, so the content discrimination loss will be high because the specific content detail (lip shape) does not match.
[0070] In this embodiment, emotion discrimination loss and content discrimination loss are calculated separately through dual discrimination of emotion and content, and then combined into a comprehensive loss to update the parameters of the facial parameter generation model. This method helps to simultaneously optimize the model's performance in both emotional expression and content details, making the generated facial lip movements more realistic and consistent with the target emotion.
[0071] In some embodiments, in step S104 above, the facial parameter generation model is a large model obtained by performing emotion and content discrimination based on historically predicted facial parameters, and training the model according to the corresponding discrimination results. Here, the historically predicted facial parameters refer to the facial parameters obtained by executing the audio-to-lip-gesture expression method based on the large model of emotion reference and discrimination described in the above embodiments on historically acquired audio data, corresponding text data, emotion category labels, and emotion reference videos.
[0072] For example, emotion and content are discriminated using historically predicted facial parameters. The resulting emotion and content discrimination results are referred to as historical sentiment discrimination results and historical content discrimination results. Based on the historical emotion discrimination results and corresponding emotion discrimination labels, a historical emotion discrimination loss is determined. Based on the historical content discrimination results and corresponding content discrimination labels, a historical content discrimination loss is determined. Based on the historical emotion discrimination loss and historical content discrimination loss, a historical comprehensive loss is generated. The model parameters of the facial parameter generation model to be trained are updated based on the historical comprehensive loss to train the facial parameter generation model, resulting in a trained facial parameter generation model.
[0073] In this embodiment, historically predicted facial parameters are used for emotion and content discrimination. A comprehensive loss is generated by combining historical emotion discrimination loss and content discrimination loss, and the model parameters are updated accordingly. This helps to improve the training accuracy of the facial parameter generation model and the accuracy of emotion and content discrimination.
[0074] According to the method of this application embodiment, by utilizing a facial reference generation model, facial parameter prediction is performed based on speech, semantics, emotion features and emotion reference features. This is beneficial for constructing a composite emotion feature space by combining multimodal features. Based on this composite emotion feature space, it is beneficial for predicting facial parameters that express complex emotional states and generating corresponding facial lip expressions. This effectively solves the technical problem that traditional single emotion category labels cannot express complex emotional states, thereby improving the accuracy of expression generation in complex emotional scenarios.
[0075] Figure 2 shows a detailed flowchart of an audio lip-syncing method based on a large model of emotion reference and discrimination, according to an exemplary embodiment of this application. As shown in Figure 2, in some embodiments, the method includes the following steps.
[0076] S201, retrieve voice, text, and emotion category tags, and emotion reference videos.
[0077] Specifically, the audio data consists of speech segments, the text consists of annotation text corresponding to the speech segments, and the sentiment category label is used to mark the sentiment category of the audio data. Specifically, it can include a first-level label category, or it can include a first-level label category and a second-level label category.
[0078] For emotion reference videos, primary and secondary emotion category labels need to be added. The processed dataset contains two types of data: labeled supervised data and unlabeled self-supervised data. Each piece of supervised data contains {audio, audio text annotation, audio emotion category label, and four-dimensional facial motion parameters}, while each piece of self-supervised data contains {audio, emotion category label, and text annotation}. Each piece of emotion reference video contains {video, primary emotion category label, and secondary emotion category label [optional]}. Four-dimensional facial motion parameters refer to the dynamic parameters of facial feature points changing over time in three-dimensional space. These parameters include the geometric features of the face (such as the position and shape of feature points like eyes, mouth, and eyebrows) and the motion changes of these feature points in the time dimension, such as displacement, velocity, and acceleration.
[0079] S202 utilizes a large language model to extract text features.
[0080] Specifically, text features are obtained by extracting features from the text using a large language model.
[0081] S203 utilizes a large speech model to extract audio features.
[0082] Specifically, large speech models are used to extract audio features from speech.
[0083] S204, Encode the sentiment category label to obtain sentiment features.
[0084] Specifically, the input sentiment category labels are encoded into the latent space through one-hot encoding and an embedding layer to obtain sentiment features.
[0085] S205, Encode the emotional reference video to obtain emotional reference features.
[0086] Specifically, the emotion reference video is input into the emotion feature extraction network to obtain emotion reference feature codes. If there are multiple emotion reference videos, multiple emotion reference feature codes are obtained, and the multiple codes are weighted and summed according to a certain ratio.
[0087] S206 fuses the various features, passes them through a network decoder, and outputs the predicted facial parameters.
[0088] Specifically, the above features are fused and decoded, and the predicted four-dimensional facial parameters are obtained through the output layer. These parameters include, for example, motion parameters of the facial features and other facial areas. Motion parameters of the facial features include, but are not limited to, at least one of the following: eye opening and closing degree, horizontal and vertical eye movement, pupil size change, gaze direction, eyebrow height, eyebrow corner position, nasal flare degree, lip opening and closing degree, corner of mouth position, and tooth exposure degree. Motion parameters of other facial areas include, but are not limited to, at least one of the following: contraction and relaxation state of cheek muscles, cheek rosiness, appearance and disappearance of forehead wrinkles, and chin movement. Different combinations of parameters can convey different emotional states.
[0089] S207: After the input and output are passed through the sentiment discriminator, sentiment consistency and content consistency are determined, the loss is calculated, and the network is guided to update the gradient.
[0090] Specifically, the predicted four-dimensional facial parameters can be used to calculate the emotion category and emotion intensity through an emotion discriminator. The calculation results can be compared and analyzed with the input emotion category and emotion intensity to obtain an emotion consistency judgment.
[0091] As an example, the sentiment discriminator is used to evaluate whether the predicted sentiment category and intensity corresponding to the four-dimensional facial parameters are consistent with the input sentiment category and intensity. Specifically, by analyzing facial parameters (such as the movements of eyebrows, eyes, and mouth), the sentiment discriminator can predict the sentiment category (such as happiness, sadness, anger, etc.) expressed by the facial parameters, as well as the intensity (such as mild happiness, extreme anger, etc.) of the expression expressed by the facial parameters.
[0092] As an example, the input emotion category and emotion intensity can be predefined target emotion labels. The prediction results of the emotion discriminator (predicted emotion category and predicted emotion intensity) are compared with the target emotion labels (labeled emotion category and labeled emotion intensity). If they match, the generated facial expression parameters are accurate. If they do not match, the model parameters of the facial parameter generation model need to be adjusted. For example, based on the emotion consistency discrimination results, the parameters of the facial parameter generation model are updated through backpropagation, making the generated facial lip movements more accurately reflect the input emotion. In virtual scenes, consistency discrimination significantly improves the expressiveness of emotional facial lip movements in virtual characters.
[0093] In real-world scenarios, emotion reference feature extraction plays a crucial role in generating emotionally rich lip-syncing expressions. While research and application of voice-driven virtual animation characters are widespread, the lack of rich and nuanced emotions and expressive facial lip-syncing significantly diminishes the user experience. This emotional deficiency directly impacts visual perception and experience, hindering the transmission of emotion. When the generated facial lip-syncing lacks subtlety, it can create a mechanical feeling for the user.
[0094] In this embodiment, the emotion reference feature extraction network is a pre-trained network module. After being trained with large datasets, this module can accurately extract the emotions and facial movement habits of people speaking in videos. When a certain emotion cannot be represented by discrete emotion labels, the emotion features extracted by the emotion reference feature network, i.e., the emotion reference features, have continuity in the emotion space and can express rich emotions more delicately and accurately.
[0095] During the facial expression generation process, specifically the model inference stage, the parameters of this network module are fixed and not modified; it is directly used as an emotion feature extraction network module. Furthermore, when the user provides two emotion reference videos, the emotion reference feature extraction network can extract features from both videos separately. The extracted features are then weighted and fused to obtain a composite emotion feature, such as a mixture of laughter and tears, or intense joy. This composite emotion feature is then used as the extracted emotion reference feature.
[0096] In some scenarios, each individual's facial lip-syncing habits and intensity differ. Using an emotion feature extraction network, these facial movement habits can be extracted as abstract, high-dimensional features, thereby guiding speech generation to produce facial lip-syncing movements with different performance habits. It should be noted that when training the emotion feature extraction network, it is trained using an emotion category task, providing two levels of emotion labels: primary label categories and secondary label categories. Primary label categories are used to label basic human emotions, such as the seven major emotion categories described in the above embodiment. Secondary label categories are used to further subdivide these seven major emotion categories in terms of type and intensity, such as the 33 subcategories of emotion categories described in the above embodiment. When labeling emotion, if it is not possible to clearly distinguish the secondary label categories, only the primary label categories can be provided.
[0097] Figure 3 illustrates a schematic diagram of the training process of the sentiment reference feature extraction network according to an exemplary embodiment of this application. As shown in Figure 3, the training process includes the following steps.
[0098] Step S301: Obtain the emotional reference video, and synchronize the visual and audio information in the emotional reference video in time and label the emotional categories.
[0099] Specifically, time synchronization refers to synchronizing visual information (such as facial expressions and lip movements) and audio information (such as speech) in the emotional reference video in time, ensuring that they are aligned on the timeline. Emotional annotation is performed on each emotional reference video. Emotional annotation includes setting primary and secondary emotion tags. For videos where secondary emotion tags cannot be clearly distinguished, only primary emotion tags can be used.
[0100] As an example, the data sources for emotional reference videos include, but are not limited to, at least one of the following: open-source datasets, internet datasets, and data collected using proprietary devices. During data collection, different age groups and genders can be covered, and as many emotional types as possible can be included to ensure a uniform distribution of the nuances of different emotions.
[0101] As an example, based on training requirements, the emotion reference video is a video sequence of frames (RGB format). These video frames can be uniformly sampled to a fixed length with a spatial resolution of 256x256. The audio contained within the video sequence can be processed to a uniform sampling rate, matching the visual frame timestamps. In other words, the synchronization of audio and video within the video sequence requires a timestamp mechanism, ensuring synchronized playback by resampling the audio to a uniform sampling rate and aligning it with the video frame timestamps.
[0102] Step S302: Train the emotion feature extraction network using emotion reference videos.
[0103] Specifically, emotion reference videos can be used as training samples to train the emotion feature extraction network.
[0104] Step S303: Calculate the loss function for the sentiment category.
[0105] Specifically, the primary and secondary sentiment categories of the reduced-dimensional sentiment reference features of the training samples are predicted to obtain primary and secondary predicted categories, with the secondary sentiment category being a subcategory of the primary sentiment category.
[0106] The loss function for sentiment classification includes a first loss function and a second loss function. The first loss function is constructed based on the first-level predicted category and the corresponding first-level label category of the training samples. The second loss function is constructed based on the second-level predicted category and the corresponding second-level label category of the training samples. For sample data without second-level label categories, only the first loss function is calculated to determine the loss for the first-level sentiment category.
[0107] Step S304: Update network parameters based on backpropagation of loss.
[0108] Through the above steps S301-S304, an emotion feature extraction network can be trained using emotion reference videos as training data.
[0109] Figure 4 illustrates a schematic diagram of the module structure of the sentiment reference feature extraction network according to an exemplary embodiment of this application. In Figure 4, the processing flow based on this module structure includes the following steps.
[0110] S401, Input an emotional reference video.
[0111] S402, Feature extraction from visual branch networks.
[0112] Specifically, in the visual branch network, spatiotemporal features of the emotional reference video are extracted, dynamic detail features are extracted from the spatiotemporal features using the fast branch, and semantic features are extracted from the spatiotemporal features using the slow branch. The dynamic detail features and semantic features are fused to obtain the visual emotional features.
[0113] S403, feature extraction of the audio branch network.
[0114] Specifically, in the audio branch network, spectrograms are extracted from the emotional reference video, local features are extracted from the spectrograms through convolutional and pooling layers, and long temporal dependencies are extracted from the local features using an encoder to obtain acoustic emotional features based on long temporal dependencies.
[0115] S404, feature fusion based on cross-multi-head attention.
[0116] Specifically, in the fusion layer, cross-head attention computation is performed on visual emotional features and acoustic emotional features to fuse the two features and obtain emotional reference features.
[0117] S405, dimensionality reduction processing of fully connected layers.
[0118] Specifically, the sentiment reference features are dimensionality-reduced by a fully connected layer (FC1) to obtain the dimensionality-reduced sentiment reference features. For example, the dimensionality is reduced to 256 degrees.
[0119] S406, Level 1 Sentiment Category Prediction.
[0120] Specifically, the output layer of the sentiment feature extraction network adopts a dual-output design, with one branch used to predict the first-level sentiment category.
[0121] For example, a fully connected layer (FC2-1) with an output dimension of 7 and a normalization layer (softmax-1) can be used, where each output value represents the probability of the corresponding class.
[0122] S407, Secondary Sentiment Category Prediction.
[0123] Specifically, another branch of the output layer of the sentiment feature extraction network is used to predict the secondary sentiment category. It uses a fully connected layer (FC2-2) with an output dimension of 33 and a normalization layer (softmax-2). This branch can be activated when the sample data contains secondary label categories.
[0124] Through steps S401-S407 above, the first-level predicted category and the second-level predicted category are obtained. Subsequently, a first loss function can be constructed based on the first-level predicted category and the first-level label category; a second loss function can be constructed based on the second-level predicted category and the second-level label category; a joint loss function can be constructed based on the first and second loss functions; and the model parameters of the sentiment feature extraction network can be updated using the joint loss function to train the sentiment feature extraction network, resulting in the trained sentiment feature extraction network.
[0125] In this embodiment, the sentiment feature extraction network is trained using a classification task. After training, the model parameters are fixed, and the dimensionality-reduced output of the sentiment reference features obtained from the fusion layer (corresponding to step S404) is used as the output of the sentiment feature layer. For example, after the sentiment feature extraction network model is trained, the dimensionality-reduced output (e.g., 256 dimensions) of the fusion layer after removing the classification layer (corresponding to steps S406 and S407) is used as the sentiment reference feature.
[0126] The overall processing flow of this application embodiment is described below with reference to Figure 5. Figure 5 shows a flowchart of the overall network module structure of an exemplary embodiment of this application. Figure 6 shows a schematic diagram of the processing method for multiple emotion reference videos in an exemplary embodiment of this application. As shown in Figure 5, the processing flow includes the following steps.
[0127] S501, Feature Extraction.
[0128] Specifically, a large-scale speech model is used to extract audio features from audio data; corresponding text is obtained from the audio data, and features are extracted from the text using a large-scale language model to obtain text features. The input sentiment category label is encoded into the latent space through a one-hot encoding and an embedding layer to obtain sentiment features. Sentiment reference videos are input into the sentiment feature extraction network to obtain sentiment reference feature encodings. If there are multiple sentiment reference videos, multiple sentiment reference feature encodings are obtained, and these encodings are weighted and summed according to a certain ratio.
[0129] As shown in Figure 6, feature extraction is performed on the two input emotion reference videos, resulting in emotion reference feature 1 and emotion reference feature 2 as shown in Figure 6. The weight of emotion reference feature 1 is... The weight of sentiment reference feature 2 is Emotional reference feature 1 and emotional reference feature 2 are weighted and fused to obtain emotional reference feature 3. This weighted and fused emotional reference feature 3 is then used as the obtained emotional reference feature. In other words, at least one (e.g., one or two) emotional reference videos are input, and after passing through the emotional feature extraction network, various emotional feature vectors are obtained. These emotional feature vectors are then weighted and fused to serve as emotional reference features, guiding the generation of facial parameters.
[0130] S502, Feature Fusion.
[0131] Specifically, facial parameters at each historical moment are input, and spatial attention is calculated on the facial parameters at each historical moment to obtain the spatial attention features of the facial parameters; causal multi-head self-attention is calculated on the spatial attention features of the facial parameters to obtain the spatiotemporal fusion attention features; feature fusion is performed on audio features, text features, and emotion features to obtain multimodal fusion features; cross-multi-head attention is calculated on the spatiotemporal fusion attention features and multimodal fusion features to obtain the cross-attention features of multimodal fusion; cross-multi-head attention is calculated on the cross-attention features of multimodal fusion and emotion reference features to obtain the emotion-enhanced multimodal fusion features.
[0132] When performing feature fusion of audio features, text features, and emotional features (also known as emotion coding features), cross-multi-head attention is used to establish connections with facial parameters at each historical moment, guiding the generation of facial parameters (lip shape, expression, etc.) at the current moment.
[0133] When establishing a connection between cross-attention and facial parameters at historical moments, emotional reference features and facial parameter features at each historical moment are projected into a shared attention space to form a query key-value triple. The similarity between historical features and emotional reference features is calculated through dot product operation, and then normalized through softmax to generate attention weights. For example, when generating an "angry" expression, the system will assign higher weights to high-energy frequency bands in the speech, emotional features, and frowning movements, thereby achieving efficient fusion of emotional reference features and facial lip movements.
[0134] By calculating attention weights, the model can focus on key parts of multimodal information that are related to historical facial states, improving prediction accuracy and the alignment between speech and historical states, and providing key feature vectors for subsequent feature processing and the decoder.
[0135] S503, decoder.
[0136] Specifically, the aforementioned features are predicted by the decoder to output the facial parameters at the current moment.
[0137] S504, sentiment discrimination and content discrimination.
[0138] S505 calculates emotional and content loss.
[0139] Specifically, the predicted facial parameters are processed by emotion discrimination and content discrimination, and the emotion loss and content loss are calculated respectively. The network is then backpropagated using the emotion loss and content loss until the loss converges, at which point the model training ends.
[0140] In this embodiment, a two-stage training approach can be adopted. The first stage trains the emotion reference feature extraction network. In the second stage, the model parameters of the emotion reference network are fixed, and the subject generation network (the generator subject) is trained, i.e., the facial parameter generation model is trained. During model training, any of the following optimizers can be used to update the model parameters, such as Adam, stochastic gradient descent (SGD), AdaGrad, AdamW, etc. The specific optimizer can be selected according to actual needs, and this embodiment does not impose any specific limitations.
[0141] This application provides an autoregressive large model technology for generating emotionally rich lip-sync expressions from text-to-speech based on emotion discrimination and emotion reference. By extracting emotion reference features from provided emotion reference videos, it guides the accurate expression of facial lip-sync expressions, enhances the richness of emotions, improves the expressive tension, and enhances the user experience.
[0142] Corresponding to the application scenarios and methods provided in the embodiments of this application, the embodiments of this application also provide an apparatus for generating lip-sync expressions from audio based on a large model of emotion reference and discrimination.
[0143] Figure 7 shows a schematic diagram of the structure of an audio lip-syncing device based on a large model of emotion reference and discrimination according to an embodiment of this application. The device is used to perform the method provided in any of the above embodiments. As shown in Figure 7, the device includes the following modules.
[0144] The data acquisition module 710 is used to acquire audio data and corresponding text data, sentiment category tags, and sentiment reference videos.
[0145] The feature extraction module 720 is used to extract features from audio data using a large speech model to obtain audio features, and to extract features from text data using a large language model to obtain text features; it also extracts features from emotion category labels and emotion reference videos to obtain emotion features and emotion reference features; the facial parameter generation model is a large model obtained by performing emotion discrimination and content discrimination based on historically predicted facial parameters and training the model according to the corresponding discrimination results.
[0146] The parameter prediction module 730 is used to predict facial parameters based on audio features, text features, emotion features and emotion reference features using a facial parameter generation model, and obtain the predicted facial parameters.
[0147] The expression generation module 740 is used to generate facial lip movements based on predicted facial parameters.
[0148] In some embodiments, the feature extraction module 720, when performing feature extraction on the emotional reference video, specifically performs the following: for any emotional reference video among the various emotional reference videos, it extracts features from the emotional reference video to obtain the emotional reference features of the emotional reference video; it performs weighted fusion on the emotional reference features of each emotional reference video, and uses the weighted fused features as the emotional reference features.
[0149] In some embodiments, when the feature extraction module 720 is used to extract features from the emotional reference video to obtain the emotional reference features of the emotional reference video, it is specifically used to: use the visual branch network and the audio branch network of the emotional feature extraction network to extract features from the emotional reference video respectively to obtain the extracted visual emotional features and acoustic emotional features; and perform multimodal fusion on the visual emotional features and the acoustic emotional features to obtain the emotional reference features.
[0150] In some embodiments, when the feature extraction module 720 extracts features from the emotional reference video using the visual branch network and the audio branch network of the emotional feature extraction network to obtain extracted visual emotional features and acoustic emotional features, it specifically performs the following steps: In the visual branch network, it extracts the spatiotemporal features of the emotional reference video, extracts dynamic detail features from the spatiotemporal features using a fast branch, extracts semantic features from the spatiotemporal features using a slow branch, and fuses the dynamic detail features and semantic features to obtain visual emotional features; In the audio branch network, it extracts a spectrogram from the emotional reference video, extracts local features from the spectrogram through convolutional layers and pooling layers, and extracts long-term dependencies from the local features using an encoder to obtain acoustic emotional features based on long-term dependencies.
[0151] In some embodiments, the parameter prediction module 730 is specifically used for: performing feature fusion processing using the feature extraction network of the facial parameter generation model. The feature fusion processing includes the following steps: performing spatial attention calculation on facial parameters at each historical time to obtain spatial attention features of facial parameters; performing causal multi-head self-attention calculation on the spatial attention features of facial parameters to obtain spatiotemporal fusion attention features; performing feature fusion on audio features, text features, and emotion features to obtain multimodal fusion features; performing cross-multi-head attention calculation on the spatiotemporal fusion attention features and multimodal fusion features to obtain cross-attention features of multimodal fusion; performing cross-multi-head attention calculation on the cross-attention features of multimodal fusion and emotion reference features to obtain emotion-enhanced multimodal fusion features; and decoding the multimodal fusion features using the decoder of the facial parameter generation model to obtain predicted facial parameters.
[0152] In some embodiments, the apparatus further includes: an emotion feature extraction network training module for training an emotion feature extraction network. The training process of the emotion feature extraction network includes: performing dimensionality reduction on the emotion reference features to obtain dimensionality-reduced emotion reference features; predicting the first-level emotion category and the second-level emotion category of the dimensionality-reduced emotion reference features to obtain a first-level predicted category and a second-level predicted category, wherein the second-level emotion category is a subcategory of the first-level emotion category; constructing a first loss function based on the first-level predicted category and the first-level label category; constructing a second loss function based on the second-level predicted category and the second-level label category; constructing a joint loss function based on the first loss function and the second loss function; and updating the model parameters of the emotion feature extraction network using the joint loss function to train the model of the emotion feature extraction network to obtain the trained emotion feature extraction network.
[0153] In some embodiments, the apparatus further includes: a facial parameter generation model training module, configured to, after obtaining predicted facial parameters, use the predicted facial parameters to perform emotion and content discrimination to obtain emotion discrimination results and content discrimination results; determine emotion discrimination loss based on the emotion discrimination results and predetermined emotion discrimination labels; determine content discrimination loss based on the content discrimination results and predetermined content discrimination labels; generate a comprehensive loss based on the emotion discrimination loss and content loss; and update the model parameters of the facial parameter generation model based on the comprehensive loss to train the facial parameter generation model and obtain a trained facial parameter generation model.
[0154] The functions of each module in each device in the embodiments of this application can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.
[0155] Figure 8 is a block diagram of an electronic device used to implement the embodiments of this application. As shown in Figure 8, the electronic device includes a memory 801 and a processor 802. The memory 801 stores a computer program that can run on the processor 802. When the processor 802 executes the computer program, it implements the methods in the above embodiments. The number of memories 801 and processors 802 can be one or more. In a specific implementation, the electronic device may also include a communication interface 803 for communicating with external devices and performing data exchange and transmission.
[0156] In practical implementation, if the memory 801, processor 802, and communication interface 803 are implemented independently, they can be interconnected via a bus to complete communication. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in Figure 8, but this does not indicate that there is only one bus or one type of bus.
[0157] Optionally, in a specific implementation, if the memory 801, the processor 802, and the communication interface 803 are integrated on a single chip, then the memory 801, the processor 802, and the communication interface 803 can communicate with each other through an internal interface.
[0158] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in this application.
[0159] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method provided in this application.
[0160] This application also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device with the chip installed to perform the method provided in this application.
[0161] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.
[0162] It should be understood that the aforementioned processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0163] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0164] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0165] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0166] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0167] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.
[0168] The logic and / or steps described in the flowchart or otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0169] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.
[0170] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0171] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for generating lip-sync expressions from audio based on a large model of emotion reference and discrimination, characterized in that, The method includes: acquiring audio data and corresponding text data, emotion category labels, and emotion reference videos; extracting features from the audio data using a large speech model to obtain audio features; extracting features from the text data using a large language model to obtain text features; extracting features from the emotion category labels and the emotion reference videos to obtain emotion features and emotion reference features; using a facial parameter generation model, predicting facial parameters based on the audio features, text features, emotion features, and emotion reference features to obtain predicted facial parameters; the facial parameter generation model is a large model obtained by performing emotion discrimination and content discrimination based on historically predicted facial parameters and training the model according to the corresponding discrimination results; and generating facial lip movements based on the predicted facial parameters.
2. The method according to claim 1, characterized in that, The feature extraction of the emotional reference videos includes: for any emotional reference video in each emotional reference video, performing feature extraction on the emotional reference video to obtain the emotional reference features of the emotional reference video; performing weighted fusion on the emotional reference features of each emotional reference video, and using the weighted fused features as the emotional reference features.
3. The method according to claim 2, characterized in that, The step of extracting features from the emotional reference video to obtain the emotional reference features of the emotional reference video includes: using the visual branch network and the audio branch network of the emotional feature extraction network to extract features from the emotional reference video respectively, to obtain the extracted visual emotional features and acoustic emotional features; and performing multimodal fusion on the visual emotional features and the acoustic emotional features to obtain the emotional reference features.
4. The method according to claim 3, characterized in that, The visual and audio branch networks of the emotion feature extraction network are used to extract features from the emotion reference video to obtain extracted visual and acoustic emotion features. This includes: in the visual branch network, extracting spatiotemporal features from the emotion reference video; using a fast branch to extract dynamic detail features from the spatiotemporal features; using a slow branch to extract semantic features from the spatiotemporal features; and fusing the dynamic detail features and semantic features to obtain the visual emotion features. In the audio branch network, extracting a spectrogram from the emotion reference video; extracting local features from the spectrogram using convolutional and pooling layers; and using an encoder to extract long-term temporal dependencies from the local features to obtain acoustic emotion features based on the long-term temporal dependencies.
5. The method according to claim 3, characterized in that, The method of using a facial parameter generation model to predict facial parameters based on the audio features, text features, emotion features, and emotion reference features to obtain predicted facial parameters includes: using the feature extraction network of the facial parameter generation model for feature fusion processing, the feature fusion processing including the following steps: performing spatial attention calculation on facial parameters at each historical time to obtain spatial attention features of the facial parameters; performing causal multi-head self-attention calculation on the spatial attention features of the facial parameters to obtain spatiotemporal fusion attention features; performing feature fusion on the audio features, text features, and emotion features to obtain multimodal fusion features; performing cross-multi-head attention calculation on the spatiotemporal fusion attention features and the multimodal fusion features to obtain cross-multimodal fusion attention features; performing cross-multi-head attention calculation on the cross-multimodal fusion attention features and the emotion reference features to obtain emotion-enhanced multimodal fusion features; and decoding the multimodal fusion features using the decoder of the facial parameter generation model to obtain the predicted facial parameters.
6. The method according to claim 5, characterized in that, The training process of the sentiment feature extraction network includes: performing dimensionality reduction on the sentiment reference features to obtain dimensionality-reduced sentiment reference features; predicting the first-level sentiment category and the second-level sentiment category of the dimensionality-reduced sentiment reference features to obtain a first-level predicted category and a second-level predicted category, wherein the second-level sentiment category is a subcategory of the first-level sentiment category; constructing a first loss function based on the first-level predicted category and the first-level label category; constructing a second loss function based on the second-level predicted category and the second-level label category; constructing a joint loss function according to the first loss function and the second loss function; and updating the model parameters of the sentiment feature extraction network using the joint loss function to train the model of the sentiment feature extraction network, thereby obtaining the trained sentiment feature extraction network.
7. The method according to claim 1, characterized in that, After obtaining the predicted facial parameters, the method further includes: using the predicted facial parameters to determine emotion and content, obtaining emotion determination results and content determination results; determining emotion determination loss based on the emotion determination results and predetermined emotion determination labels; determining content determination loss based on the content determination results and predetermined content determination labels; generating a comprehensive loss based on the emotion determination loss and the content determination loss; and updating the model parameters of the facial parameter generation model based on the comprehensive loss to train the facial parameter generation model, thereby obtaining a trained facial parameter generation model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.
10. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.