Plug-and-play semantic feature decoupling method
By integrating the Wav2Sem module in the self-supervised audio encoder, decoupling the feature space of homophones, the inaccuracy and unnatural problems of facial animation in the prior art are solved, and a more natural and accurate speech-driven facial animation generation is achieved.
Patent Information
- Application Number
- CN202510484084.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, homophones lead to feature spatial coupling when generating facial animations, resulting in inaccurate and unnatural lip generation, and the self-supervised audio model ignores the global semantic information of the audio signal, affecting the accuracy and naturalness of facial expressions.
By training the Wav2Sem module, audio semantic features are learned using large-scale audio-text alignment data and integrated into a self-supervised pre-trained audio encoder. Time series features are generated to drive virtual human face animation by adding and fusing semantics and phoneme features through corresponding channels.
Effectively decouple semantic confusion of proximity syllables, improve the naturalness and accuracy of facial animation, enhance the expressiveness of voice-driven facial animation, has plug-and-play and high robustness, and is suitable for fields such as speech recognition, speech synthesis and virtual human interaction.
Smart Images

Figure CN120339467A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of audio processing and computer graphics, and particularly to a plug-and-play semantic feature decoupling method. Background Art
[0002] With the rapid development of virtual reality, augmented reality, and digital human technologies, 3D facial animation generation driven by speech has become a research hotspot. Existing facial animation generation methods usually drive the changes of facial expressions through speech signals, especially using self-supervised audio models in deep learning technology for audio feature extraction. However, when dealing with syllables with similar pronunciations but different mouth shapes, existing technologies often have the problem of feature space coupling, resulting in an averaging effect of near-sound syllables in the generated facial expression animations.
[0003] Specifically, audio signals contain rich acoustic features. Different syllables may have similar pronunciations, but due to their corresponding different mouth shapes, traditional audio feature modeling methods cannot effectively distinguish these near-sound syllables, resulting in excessive averaging in the feature space when generating facial animations, which affects the accuracy and naturalness of facial animations. In addition, existing self-supervised audio models mostly focus on phoneme-level feature modeling and ignore the global semantic information in audio signals, making them lack a full understanding of the overall semantics and context when generating facial expressions.
[0004] Therefore, a new method is needed that can effectively decouple the influence of near-sound syllables in audio features and can improve the expressiveness of audio-driven facial animations by introducing global semantic features, thereby improving the accuracy and naturalness of facial expression generation. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to achieve decoupling of homophones in the feature space. Aiming at the above-mentioned defects of the existing technology, a plug-and-play audio semantic decoupling method is provided, aiming to solve the audio feature space coupling problem in the existing 3D speech-driven facial animation technology, especially the averaging phenomenon when dealing with homophones. By introducing global semantic features, the present invention can effectively decouple the influence of near-sound syllables in audio features, thereby improving the accuracy and naturalness of facial animations.
[0006] The technical solution adopted by the present invention to solve the problem is as follows:
[0007] A plug-and-play semantic feature decoupling method, the method comprising:
[0008] Step S100: Training of the Wav2Sem module: Using a large audio-text corpus database containing large-scale audio-text alignment data, train the Wav2Sem module using a supervised learning strategy so that it can learn audio semantic features and decouple the semantic information between near-homophones;
[0009] Step S200: Integration of the Wav2Sem module: Insert the Wav2Sem module into an existing self-supervised pre-trained audio encoder;
[0010] Step S300: Construction of a new audio encoder: Replace the audio encoder in the existing voice drive control framework with the new audio encoder and train the voice drive control framework;
[0011] Step S400: Speech-driven facial animation generation: Use the trained new audio encoder to extract features from the input speech signal, generate time series features, and map them to the parameter space of the three-dimensional facial model to finally drive the virtual human face animation.
[0012] Advantages of the present invention:
[0013] The Wav2Sem module proposed by the present invention enhances the expressive ability of audio features and improves the problem of precise averaging of lip shape generation caused by feature coupling due to homophones by extracting semantic features from the complete audio sequence and using the extracted semantic information to decouple the feature space of audio coding. The feature decoupling of homophones in the self-supervised audio encoder is achieved through the plug-and-play Wav2Sem module. By tightly integrating self-supervised pre-trained audio encoders (such as Wav2Vec 2.0, HuBERT, DeepSpeech) with the Wav2Sem model, users can easily apply these modules directly to different speech-driven tasks without additional configuration or complex debugging. This plug-and-play feature makes the technology highly flexible and scalable in various practical applications. Figure 4 It shows the improvement of the Wav2Sem module on the currently commonly used speech-driven facial generation method, which can be quickly deployed in fields such as speech recognition, speech synthesis, virtual assistants, etc., greatly reducing the complexity of system integration and update.
[0014] The present invention is compatible with a variety of speech-driven models, effectively reducing the averaging effect of near-syllables and improving the naturalness and realism of facial animations. It solves the problems of inaccurate and unnatural lip shape generation caused by speech feature coupling in the prior art. Description of the Drawings
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a flowchart of the plug-and-play semantic feature decoupling method provided by the present invention.
[0017] Figure 2 It is a working principle diagram of the plug-and-play module Wav2Sem provided by the embodiments of the present invention.
[0018] Figure 3 It is an overall architecture diagram of the plug-and-play module Wav2Sem provided by the embodiments of the present invention.
[0019] Figure 4 It is a schematic diagram comparing the transitional generation effects of facial actions when different methods are used with and without Wav2Sem in the near-homophone syllable decoupling task on the VOCA-Test (left) and BIWI-Test-B (right) datasets provided by the embodiments of the present invention. Detailed implementation manners
[0020] The present invention discloses a plug-and-play semantic feature decoupling method. To make the purpose, technical solutions and effects of the present invention clearer and more definite, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0021] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0022] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.
[0023] In view of the above-mentioned defects of the prior art, the present invention provides a plug-and-play semantic feature decoupling method. The method first performs data cleaning, alignment, and feature extraction on a large-scale audio-text alignment dataset covering multiple languages, pronunciation styles, and background noises. Secondly, the Wav2Sem module is trained to be able to extract semantic information in audio and effectively decouple the features of near-homophones, reducing semantic confusion. Then, the trained Wav2Sem module is embedded in a self-supervised pre-trained audio encoder (such as Wav2Vec 2.0, HuBERT, DeepSpeech), and the semantic and phoneme features are fused by adding the corresponding channels. The enhanced audio encoder is used to optimize the speech-driven facial animation task, improving the accuracy and generation effect of semantic alignment, and is applicable to scenarios such as speech recognition, speech synthesis, and virtual human interaction. This method has the characteristics of plug-and-play, high robustness, and wide applicability, providing more accurate semantic control capabilities for speech understanding and generation.
[0024] Figure 1 is a flowchart of the plug-and-play semantic feature decoupling method provided by the present invention. As Figure 1 shown, the method includes:
[0025] Step S100: Training of the Wav2Sem module: Based on a large audio-text corpus database containing large-scale audio-text alignment data, a supervised learning method is used to train a speech feature decoupling module (Wav2Sem module) so that it can learn audio semantic features and decouple the semantic information between near-homophones;
[0026] Step S200: Integration of the Wav2Sem module: Insert the Wav2Sem module into an existing self-supervised pre-trained audio encoder to obtain a new audio encoder, so as to enhance its ability to decouple semantic features in speech, and reduce the confusion of near-homophones during the encoding process;
[0027] Step S300: Construction of the new audio encoder: Replace the audio encoder in the existing speech drive control framework with the new audio encoder, and train the speech drive control framework to optimize the performance of the speech-driven facial animation task;
[0028] Step S400: Speech-driven Facial Animation Generation: Use the trained new audio encoder to extract features from the input speech signal, generate time-series features, and map them to the parameter space of the 3D facial model, ultimately driving the facial animation of the virtual human to exhibit natural and realistic speech-synchronized facial movements.
[0029] Specifically, the present invention proposes a plug-and-play speech feature decoupling module Wav2Sem. By integrating the Wav2Sem module into self-supervised pre-trained audio encoders (such as Wav2Vec 2.0, HuBERT, DeepSpeech), the model's understanding of speech semantics is improved. Wav2Sem is trained with large-scale audio-text alignment data to align the semantic information in the audio and effectively decouple the features of near-homophones, reducing semantic confusion. The integration method uses channel-wise addition to enhance the fusion of phoneme and semantic features and optimize the naturalness of speech-driven facial animation. This method is both plug-and-play, highly robust, and widely applicable, and can be used in multiple fields such as speech recognition, speech synthesis, and virtual human interaction.
[0030] In one implementation, the large audio-text corpus database containing large-scale audio-text alignment data is Librispeech. The database contains audio data in multiple languages and their corresponding transcribed texts, and covers various pronunciation styles, speaker identities, emotional states, and different background noise conditions to improve the generalization ability of the Wav2Sem model.
[0031] In one implementation, as Figure 3 shown, the Wav2Sem module includes a Temporal Convolutional Network (TCN) module for obtaining the input raw audio signal, processing the raw audio signal using the Temporal Convolutional Network (TCN) structure, and modeling the short-term features of the audio features; a context feature extraction module for capturing the context information between the short-term features modeled by the TCN module of the temporal convolutional network to enhance the separability of semantic features; and a semantic feature extraction module responsible for aggregating the context information captured by the context feature extraction module to obtain the global semantic information of the entire audio segment and generate the final speech semantic representation.
[0032] The TCN module can be composed of 7 convolutional blocks. Each layer uses the ReLU activation function and is equipped with residual connections to maintain the stability of the information flow. The entire Temporal Convolutional Network (TCN) module can be expressed as , where A is the input audio sequence, is the local audio feature information, N is the token output length, and C is the feature dimension of 768.
[0033] Each TCN convolutional block can have 768 channels, and the time step of the output features is 49Hz, that is, each time step corresponds to an audio segment of about 20ms to ensure sufficient time resolution.
[0034] The context feature extraction module can adopt 12 Transformer blocks. Each Transformer block uses the multi-head self-attention mechanism to extract the global dependencies of speech features. The multi-head self-attention (MHSA) mechanism contains 8 attention heads, which are used to focus on different context information in different semantic dimensions. The number of channels of the feed-forward neural network (FFN) in the Transformer module is set to 3072, and the GeLU activation function is used to improve the non-linear modeling ability and enhance the modeling ability of complex semantic relationships. Layer normalization (LayerNorm) and residual connections are used to stabilize the training of the deep network and improve the convergence speed of the model. The Transformer block can be expressed as:
[0035] ,
[0036] where MHSA refers to the multi-head self-attention mechanism, and LN is the layer normalization in the Transformer block. represents the output of the (l-1)-th layer Transformer block, is the intermediate result, is the feed-forward neural network in the Transformer block.
[0037] The semantic feature extraction module can globally aggregate the temporal features extracted by the Transformer block through global average pooling to obtain the global semantic information of the entire audio segment, which can be expressed as:
[0038] ,
[0039] where, is the semantic feature extracted by the Wav2Sem module, is the context feature calculated by the Transformer block, N is the output token length, and i corresponds to each token.
[0040] By using the global average pooling method to globally aggregate the temporal features, a fixed-length semantic feature vector is obtained, ensuring that the finally output semantic information has higher stability and generalization ability. The generated semantic features can be used for different downstream tasks.
[0041] In one implementation, such as Figure 3As shown, the labels in the Wav2Sem training process are the text encoding results of the BERT language model, including:
[0042] CLS token scheme: Extract the first [CLS] token of the sequence output by the BERT language model as the label in the Wav2Sem training process. The first [CLS] token is the semantic feature vector of the entire sentence, and the entire sentence refers to all the text input into the BERT model; this scheme utilizes the encoding ability of the BERT model to compress the overall semantics of the sentence into a fixed-length vector representation, which helps to achieve precise alignment between speech and text semantics.
[0043] Average pooling scheme: Perform mean pooling on all tokens output by the BERT language model, and use the obtained average value as the label in the Wav2Sem training process to generate a stable text semantic representation. By calculating the mean value of all tokens, the global semantic information in the sentence can be retained, further improving the alignment effect between speech and text features.
[0044] In the CLS token scheme, the training label of Wav2Sem is provided by the [CLS] token, and in the average pooling scheme, the training label of Wav2Sem is provided by the average value encoded by each input text token.
[0045] In one implementation, the BERT model is the pre-trained BERT-base or BERT-large to enhance the ability to extract text semantic information and improve the accuracy of speech-semantic alignment, and its output feature dimension is 768.
[0046] In one implementation, the loss function in the Wav2Sem training process is the L1 loss function, which can be expressed as: where represents the semantic feature generated by Wav2Sem, is the first [CLS] token feature extracted from the text sequence output by the BERT model, represents the 1-norm. By optimizing with the L1 loss function, the error between speech and text can be effectively reduced, and the accuracy of speech-semantic alignment can be improved.
[0047] In one implementation, the optimizer in the Wav2Sem training process is Adam.
[0048] In one implementation, the number of training epochs of Wav2Sem is 200 to fully learn the mapping relationship between audio and text semantics.
[0049] In one implementation, the learning rate of Wav2Sem is 1 10-4 。
[0050] In one implementation, the batch number of Wav2Sem is 1.
[0051] Step S200: Plug-and-play: Insert the Wav2Sem module into an existing self-supervised pre-trained audio encoder to obtain a new audio encoder, where the existing self-supervised pre-trained audio encoder is one of Wav2Vec 2.0, HuBERT, and DeepSpeech. These self-supervised learning models are trained with a large amount of unlabeled audio data and can extract rich phoneme features from audio.
[0052] Figure 2 is the working principle diagram of the plug-and-play module Wav2Sem, showing the feature decoupling at the word level. At the same time, as Figure 3 shown, it shows the feature decoupling of Wav2Sem at the phoneme level. The insertion method of inserting the Wav2Sem module into an existing self-supervised pre-trained audio encoder uses the method of adding corresponding channels, that is, adding the semantic features generated by Wav2Sem and the phoneme features generated by the self-supervised pre-trained audio encoder corresponding channels by channel, thereby reducing the confusion of near-homophonic words and improving semantic discrimination. The insertion method can be described as , where represents the semantic features generated by Wav2Sem, is the feature of the self-supervised pre-trained audio encoder, is the new audio representation feature, represents the feature mapping layer. The feature mapping layer is a layer of parameters trained when wav2sem generates new audio features during plug-and-play.
[0053] The text semantic feature vector generated by Wav2Sem and the phoneme feature vector generated by the self-supervised pre-trained audio encoder are added on the corresponding channels. This method can closely combine the semantic information of the text with the phoneme feature information of the audio, improving the performance of the model in speech-driven tasks. Through the addition operation, the semantic information of speech and text can be fused in the same feature space, promoting the semantic alignment of speech and text.
[0054] Step S300: Model training: Replace the audio encoder in the existing speech driving framework with the new audio encoder and train it, where the speech driving framework can be CodeTalker, and the audio encoder in the existing speech driving framework can be a self-supervised pre-trained encoder.
[0055] In one implementation, the VOCASET dataset is used to train the speech-driven facial framework, that is, the speech driving framework.
[0056] In one implementation, the optimizer for the CodeTalker training process is Adam.
[0057] In one implementation, the number of training rounds for the CodeTalker codebook is 400.
[0058] In one implementation, the learning rate for the CodeTalker codebook training is 1 10 -4 。
[0059] In one implementation, the batch size for the CodeTalker codebook training is 1.
[0060] In one implementation, the number of training rounds in the CodeTalker backbone model is 200.
[0061] In one implementation, the learning rate for the CodeTalker backbone model training is 1 10 -4 。
[0062] In one implementation, the batch size for the CodeTalker backbone model training is 1.
[0063] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can implement the processes of the embodiments of the above-described methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0064] In summary, the present invention provides a plug-and-play semantic feature decoupling method, aiming to improve the accuracy of speech-semantic alignment, reduce semantic confusion between near-homophones, and optimize the generation effect of speech-driven facial animation. The core steps of this method include: data preparation and preprocessing, collecting and cleaning a large-scale audio-text alignment dataset covering multiple languages, pronunciation styles, and background noises to ensure data quality and alignment accuracy. Training of the Wav2Sem module, training the Wav2Sem module through supervised learning so that it can extract semantic information in the audio and decouple the features of near-homophones. This module includes TCN, Transformer, and a semantic feature extraction module to enhance the semantic separability of the audio. Plug-and-play integration, embedding the trained Wav2Sem module into a self-supervised pre-trained audio encoder (such as Wav2Vec 2.0, HuBERT, DeepSpeech), and fusing the speech features and semantic information of the audio by way of channel addition to improve the alignment accuracy of speech semantics. Model training, replacing the audio encoder in the existing framework with the enhanced audio encoder and training it to optimize the performance of tasks such as speech-driven facial animation and speech recognition. Inference generation, extracting features from the given audio through the trained enhanced audio encoder to drive facial animation or complete other speech-driven tasks. As Figure 2 , Figure 3 shown, this method has the characteristics of plug-and-play, can complete the decoupling of homophones and corresponding syllables in the feature space, and has strong adaptability, and can be widely applied to fields such as speech recognition, speech synthesis, and virtual human interaction to improve the accuracy of speech understanding and generation.
[0065] Although the present invention has been described based on a limited number of embodiments, those skilled in the art in this technical field will understand, based on the above description, that other embodiments can be envisioned within the scope of the present invention thus described. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention.
Claims
1. A plug-and-play semantic feature decoupling method, characterized in that The method includes: Step S100: Training of the Wav2Sem module: Using a large audio-text corpus database containing large-scale audio-text alignment data, training the Wav2Sem module using a supervised learning strategy so that it can learn audio semantic features and decouple the semantic information between near-homophones; Step S200: Integration of the Wav2Sem module: Inserting the Wav2Sem module into an existing self-supervised pre-trained audio encoder; Step S300: Construction of a new audio encoder: Replacing the audio encoder in the existing voice drive control framework with the new audio encoder and training the voice drive control framework; Step S400: Speech-driven facial animation generation: Using the trained new audio encoder to extract features from the input speech signal, generating time series features, and mapping them to the parameter space of the 3D facial model to finally drive the virtual human face animation.
2. The plug-and-play semantic feature decoupling method according to claim 1, wherein The large audio-text corpus database contains audio data in multiple languages and their corresponding transcribed texts, and covers multiple pronunciation styles, speaker identities, emotional states, and different background noise conditions.
3. The plug-and-play semantic feature decoupling method according to claim 1, characterized in that, The Wav2Sem module includes: Temporal Convolutional Network (TCN) module: Used to obtain the input raw audio signal, processing the raw audio signal using a temporal convolutional network structure and modeling the short-term features of audio features; Context feature extraction module: Used to capture the context information between the short-term features modeled by the Temporal Convolutional Network (TCN) module to enhance the separability of semantic features; Semantic feature extraction module: Responsible for aggregating the context information captured by the context feature extraction module to obtain the global semantic information of the entire audio segment and generating the final speech semantic representation.
4. The plug-and-play semantic feature decoupling method according to claim 3, characterized in that The TCN module consists of 7 convolutional blocks, each layer using the ReLU activation function and equipped with residual connections to maintain the stability of the information flow.
5. The plug-and-play semantic feature decoupling method according to claim 4, characterized in that Each TCN convolutional block has 768 channels, and the time step of the output feature is 49Hz.
6. The plug-and-play semantic feature decoupling method according to claim 3, characterized in that The context feature extraction module consists of 12 Transformer blocks, and each Transformer block uses the multi-head self-attention mechanism.
7. The plug-and-play semantic feature decoupling method according to claim 6, characterized in that: The multi-head self-attention mechanism contains 8 attention heads, used to focus on different context information in different semantic dimensions; the number of channels of the feed-forward neural network of the Transformer block is 3072, and GeLU is used as the activation function.
8. The plug-and-play semantic feature decoupling method according to claim 6, characterized in that The semantic feature extraction module aggregates the temporal features extracted by the Transformer module through global average pooling to obtain the global semantic information of the entire audio segment.
9. The plug-and-play semantic feature decoupling method according to claim 1, characterized in that, The labels in the Wav2Sem training process come from the text encoding results of the BERT language model, including: CLS token scheme: Extracting the first [CLS] token of the output sequence of the BERT language model as the label in the Wav2Sem training process; Average pooling scheme: Perform mean pooling on all tokens output by the BERT language model, and use the obtained average value as the label during the training process of Wav2Sem.
10. The plug-and-play semantic feature decoupling method according to claim 1, characterized in that The training process of Wav2Sem is optimized using the L1 loss.
11. The plug-and-play semantic feature decoupling method according to claim 9, wherein The BERT language model is the pre-trained BERT-base or BERT-large.
12. The plug-and-play semantic feature decoupling method according to claim 1, wherein The self-supervised pre-trained audio encoder is Wav2Vec 2.0, HuBERT, or DeepSpeech.
13. The plug-and-play semantic feature decoupling method according to claim 1, characterized in that Step S200 includes: The method of inserting the Wav2Sem module into the existing self-supervised pre-trained audio encoder is by adding corresponding channels, that is, adding the semantic features generated by the Wav2Sem module and the phoneme features generated by the self-supervised pre-trained audio encoder in corresponding channels.
14. The plug-and-play semantic feature decoupling method according to claim 1, characterized in that In step S300, the audio encoder in the existing voice control framework uses a self-supervised pre-trained audio encoder.