Virtual human mouth shape driving method, device, equipment and medium based on dynamic and static feature fusion

Through the combination of dynamic and static feature fusion and noise reduction module, the problem of insufficient accuracy and efficiency of mouth animation in the prior art is solved, and a more natural and smooth virtual mouth animation generation is achieved.

CN119724193BActive Publication Date: 2025-05-16TIANDU (XIAMEN) TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510246463.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-16
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing speech-driven three-dimensional virtual human mouth shape change technology still lacks accuracy and efficiency, especially in the generation of mouth animations and inaccurate mouth closure during speech silence.

Method used

The method of fusion of dynamic and static characteristics is adopted, and the speech characteristics are obtained and voice data are extracted, and the global information and local characteristics are combined to generate enhanced speech dynamic and static characteristics. Then, these features are input to the noise-reducing module and processed through the Transformer layer of the encoding block, the intermediate block and the decoding block to predict the hybrid deformation parameter sequence to drive the virtual human mouth animation.

Benefits of technology

It improves the accuracy of mouth animation, solves the problem of inaccurate mouth closure when voice silence, and reduces abnormal mouth shaking, and the generated mouth animation is more natural and smooth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724193B_ABST
    Figure CN119724193B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, equipment and medium for driving a virtual human mouth shape by fusing dynamic and static features, and relates to the technical field of audio processing and image processing. The method for driving a virtual human mouth shape comprises: S1, obtaining voice data and extracting voice features. S2, extracting global information from voice features to obtain enhanced dynamic global features. S3, extracting static local features from voice features to obtain static local features. S4, fusing voice features, enhanced dynamic global features and static local features to obtain enhanced voice dynamic and static features. S5, obtaining a noisy mixed deformation parameter sequence based on facial mesh deformation parameters. S6, inputting the enhanced voice dynamic and static features and the mixed deformation parameter sequence with noise into a denoising module to obtain a predicted mixed deformation parameter sequence. The denoising module comprises a coding block, an intermediate block and a decoding block. The mixed deformation parameter sequence with noise is input into the convolution layer of the coding block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing and image processing, and in particular to a method, device, equipment and medium for driving a virtual human mouth shape by fusing dynamic and static features. Background Art

[0002] With the continuous development of virtual reality, augmented reality and digital content creation technology, 3D virtual human technology has gradually become a hot research field. Among them, voice-driven 3D virtual human mouth shape change technology is particularly important. It can automatically generate corresponding 3D virtual human facial mouth shape animation according to any voice signal during the interaction between users and virtual humans, thereby enhancing the user's sense of immersion. This technology has broad application prospects in scenes such as film production and virtual character interaction.

[0003] In real life, the changes in people's mouth shapes when speaking are complex and diverse, involving the coordinated movement of multiple parts such as lips, tongue and chin. The changes in mouth shape are not only determined by pronunciation, but also affected by emotions and intonation. Therefore, in the process of 3D mouth animation synthesis, how to accurately map audio features to mouth shape movements is a technical problem that needs to be solved urgently.

[0004] Early lip-driven technology was mainly based on visemes (the visual embodiment of phonemes). By formulating a one-to-one mapping rule between phonemes and lip movements, lip movements were integrated into the binding animation of the lips and chin. This method uses prior speech knowledge to extract viseme parameters from speech videos, which has strong comprehensibility and operability. However, the lip-shaped effects generated by this method are unstable and usually require manual adjustments by animators in the later stage, which is inefficient.

[0005] In recent years, with the development of deep learning technology, deep learning-based lip-driven methods have made significant progress. These methods can be roughly divided into two categories: methods based on facial mesh vertex data and methods based on facial mesh deformation parameters. Methods based on facial mesh vertex data can generate detailed facial mesh vertex data, but they are usually based on standard models and difficult to extend to other models. Methods based on facial mesh deformation parameters solve the model scalability problem of methods based on mesh vertex data, but they mainly rely on local facial information and audio information of adjacent frames, lack of attention to global information, resulting in the accuracy of the generated lip animation still needs to be improved.

[0006] The above problems limit the effect and efficiency of voice-driven three-dimensional virtual human mouth shape change technology in practical applications, and therefore need further improvement and optimization. Summary of the invention

[0007] The present invention provides a method, device, equipment and medium for driving a virtual human mouth shape by fusing dynamic and static features, so as to improve at least one of the above-mentioned technical problems.

[0008] In a first aspect, the present invention provides a method for driving a virtual human mouth shape by fusing dynamic and static features, which comprises steps S1 to S6.

[0009] S1. Acquire speech data and extract speech features.

[0010] S2. Extract global information from the speech features to obtain enhanced dynamic global features.

[0011] S3. Extract static local features from the speech features to obtain static local features.

[0012] S4. Fusing the speech feature, the enhanced dynamic global feature and the static local feature to obtain enhanced speech dynamic and static features.

[0013] S5. Obtain a mixed deformation parameter sequence with noise based on the facial mesh deformation parameters.

[0014] S6. Input the enhanced speech dynamic and static features and the mixed deformation parameter sequence with noise into a denoising module to obtain a predicted mixed deformation parameter sequence. The denoising module includes a coding block, an intermediate block and a decoding block. The mixed deformation parameter sequence with noise is input into the convolution layer of the coding block. The enhanced speech dynamic and static features are input into the Transformer layers of the coding block, the intermediate block and the decoding block.

[0015] In an optional implementation, step S2 specifically includes step S21 to step S22.

[0016] S21, the speech feature is first compressed by a linear layer to the number of speech channels, and then the standard The activation function activates the deep features of the speech, and finally the number of speech channels is expanded back to the original number of channels through the linear layer to obtain dynamic global features. The dynamic global feature extraction model is: , where For dynamic global features, For voice features, For the linear layer, for Activation function.

[0017] S22, the dynamic global feature and the speech feature are combined Product, to obtain enhanced dynamic global features. The product model is: , where To enhance dynamic global features, For dynamic global features, for product, For voice features.

[0018] In an optional implementation, step S3 specifically includes: first performing a convolution operation on the speech feature to capture the local features of the audio, and then using Normalize and use it last The activation function extracts high-dimensional data and obtains static local features. , where For voice features, is a one-dimensional convolutional layer, is the normalization layer, for Activation function, It is a static local feature.

[0019] In an optional implementation, step S4 specifically includes step S41 to step S42.

[0020] S41, adding the enhanced dynamic global feature and the static local feature, and performing Operation to obtain the fused dynamic and static features. , where To integrate the dynamic and static features, is the normalized exponential function, To enhance dynamic global features, It is a static local feature.

[0021] S42, the fused dynamic and static features and the voice features are combined The product is used to obtain the enhanced speech dynamic and static features. , where To enhance the dynamic and static characteristics of speech, To integrate the dynamic and static features, for product, For the linear layer, For voice features.

[0022] In an optional implementation, step S5 specifically includes steps S51 to S54.

[0023] S51, obtaining a mouth shape animation data set based on mesh vertex data as a target model.

[0024] S52, obtaining a three-dimensional virtual human face model with mixed deformation parameters of a blended deformation deformer representing speech mouth shape drive as a source model.

[0025] S53, using the virtual human facial animation data mapping model, adjusting the shape of the source model to the shape of the target model, and obtaining a final hybrid deformation parameter sequence.

[0026] S54: adding noise to the final mixed deformation parameter sequence to obtain a mixed deformation parameter sequence with noise.

[0027] In an optional implementation, the virtual human facial animation data mapping model is:

[0028] .

[0029] In the formula, is the sequence of mixed deformation parameters after mapping, is the total number of frames of facial animation data, is the frame number, For the Frame facial mesh vertex data, is the standard face mesh vertex vector, is the total number of blend deformation parameters, is the sequence number of the blend deformation parameter, For the Frame No. The value of the blend shape parameter, For the The vector of facial mesh vertices corresponding to the blend shape parameters.

[0030] In an optional implementation, the noise addition model is:

[0031] .

[0032] In the formula, is a mixed deformation parameter sequence with noise, represents the noise adding module, is the sequence of mixed deformation parameters after mapping, is the noise intensity, is standard Gaussian noise.

[0033] In an optional implementation, the encoding block, the intermediate block and the decoding block are connected in sequence.

[0034] The encoding block includes a convolutional layer, a residual layer and a Transformer layer connected in sequence.

[0035] The intermediate block includes a residual layer and a Transformer layer connected in sequence.

[0036] The decoding block includes a residual layer, a Transformer layer and a normalization layer connected in sequence.

[0037] The coded block, the intermediate block and the decoded residual layer are all embedded with time steps.

[0038] In an optional implementation, the model of the noise reduction module is:

[0039] .

[0040] In the formula, is the predicted hybrid deformation parameter sequence, Indicates noise reduction, is a hyperparameter, is a mixed deformation parameter sequence with noise, To enhance the dynamic and static characteristics of speech, is the total time step, represents the weight of the linear combination, Indicates that the condition is empty.

[0041] In an optional implementation, the loss function of the virtual human mouth shape driving method is:

[0042] .

[0043] .

[0044] .

[0045] In the formula, is the deep smoothing feature loss function, is the sample loss function, is the speed matching loss function, Express expectations, Represents the true sequence of blended deformation parameters for a single frame, represents the predicted hybrid deformation parameter sequence obtained conditionally by the parameters in brackets, is a mixed deformation parameter sequence with noise, To enhance the dynamic and static characteristics of speech, is the local time step, is the total number of frames of facial animation data, is the frame number, Indicates Frame noise mixed deformation parameter sequence, Indicates Frame noise mixed deformation parameter sequence, Represents the predicted hybrid deformation parameter sequence obtained according to the parameters in brackets.

[0046] In the second aspect, the present invention provides a virtual human mouth shape driving method device with dynamic and static feature fusion, which includes a speech feature extraction module, a global feature extraction module, a local feature extraction module, a feature fusion module, a mapping module and a noise reduction module.

[0047] The speech feature extraction module is used to obtain speech data and extract speech features.

[0048] The global feature extraction module is used to extract global information from the speech features to obtain enhanced dynamic global features.

[0049] The local feature extraction module is used to extract static local features from the speech features to obtain static local features.

[0050] The feature fusion module is used to fuse the speech feature, the enhanced dynamic global feature and the static local feature to obtain the enhanced speech dynamic and static feature.

[0051] A mapping module is used to obtain a noisy hybrid deformation parameter sequence based on the facial mesh deformation parameters.

[0052] The denoising module is used to input the enhanced speech dynamic and static features and the mixed deformation parameter sequence with noise into the denoising module to obtain a predicted mixed deformation parameter sequence. The denoising module includes a coding block, an intermediate block and a decoding block. The mixed deformation parameter sequence with noise is input into the convolution layer of the coding block. The enhanced speech dynamic and static features are input into the Transformer layers of the coding block, the intermediate block and the decoding block.

[0053] In a third aspect, the present invention provides a method and device for driving a virtual human mouth shape by fusing static and dynamic features, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a method for driving a virtual human mouth shape by fusing static and dynamic features as described in any paragraph of the first aspect.

[0054] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a virtual human mouth shape driving method with dynamic and static feature fusion as described in any paragraph of the first aspect.

[0055] By adopting the above technical solution, the present invention can achieve the following technical effects:

[0056] The virtual human mouth shape driving method of the present invention solves the problem of inaccurate mouth shape closure when the speech is silent by realizing a balanced transition between the global and local features of the speech.

[0057] In addition, a deep feature smoothing loss function is added to reduce abnormal jitter of the mouth shape, and ultimately generate a more accurate mouth animation, which can be applied to voice-driven 3D virtual human mouth animation related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the specific implementation methods of the present invention. It should be understood that the following drawings only show certain specific implementation methods of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0059] Figure 1 It is a flowchart of a virtual human mouth shape driving method that integrates static and dynamic features.

[0060] Figure 2 It is a schematic diagram of the neural network structure of the virtual human mouth shape driving method that integrates dynamic and static features. DETAILED DESCRIPTION

[0061] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] Example 1, please refer to Figure 1 to Figure 2 The first embodiment of the present invention provides a method for driving a virtual human mouth shape by integrating dynamic and static features, which can be performed by a virtual human mouth shape driving device by integrating dynamic and static features (hereinafter referred to as: virtual human mouth shape driving device). In particular, it is performed by one or more processors in the virtual human mouth shape driving device to implement steps S1 to S6.

[0063] S1. Acquire speech data and extract speech features.

[0064] Specifically, step S1 uses the pre-trained wav2vec2.0 model to extract features from the speech data to obtain input speech features.

[0065] It is understandable that the virtual human mouth shape driving device can be an electronic device with computing performance, such as a portable notebook computer, a desktop computer, a server, a smart phone or a tablet computer.

[0066] For voice features In this embodiment, a collaborative attention module that integrates dynamic global features and static local features is designed. The collaborative attention module that integrates dynamic global features and static local features mainly includes a dynamic global feature extraction module and a static local feature extraction module.

[0067] S2, extracting global information from the speech features to obtain enhanced dynamic global features. Figure 2 As shown, step S2 specifically includes step S21 to step S22.

[0068] S21, the speech feature is first compressed by a linear layer to the number of speech channels, and then the standard The activation function activates the deep features of the speech, and finally the number of speech channels is expanded back to the original number of channels through the linear layer to obtain dynamic global features. The dynamic global feature extraction model is: , where For dynamic global features, For voice features, For the linear layer, for Activation function.

[0069] Specifically, for the global information extraction of the input speech features, the number of speech channels is first compressed through a linear layer, and the standard The activation function activates the deep features of the speech, and then expands the number of speech channels back to the original number of channels through the linear layer to obtain the dynamic global features of the input speech features.

[0070] The first linear layer compresses the number of channels of the input speech features to 1 / 8 of the original. As an activation function, it can make negative inputs have certain outputs, avoid the problem of dead neurons, and provide better gradient information during training. The number of channels is expanded back to the original number of channels for subsequent fusion of dynamic global features and static local features.

[0071] The dynamic global feature extraction module, through the combination of linear layers and nonlinear activation functions, can effectively compress and expand the speech feature channel and achieve accurate extraction of global features. The activation function captures deep semantic features, improves the model's adaptability to diverse speech inputs, and ensures that the output features are consistent with the original dimensions, making it easier to connect with subsequent modules. This module enhances the dynamics and robustness of feature extraction while maintaining computational efficiency, and can handle the dynamic global features of speech well.

[0072] S22, the dynamic global feature and the speech feature are combined Product, to obtain enhanced dynamic global features. The product model is: , where To enhance dynamic global features, For dynamic global features, for product, For voice features.

[0073] Considering that the original speech features still have a certain influence on the global information, this embodiment compares the obtained dynamic global features with the original speech features. Product, get enhanced dynamic global features .

[0074] S3, extracting static local features from the speech features to obtain static local features. Figure 2 As shown, step S3 specifically includes: first performing a convolution operation on the speech feature to capture the local features of the audio, and then using Normalize and use it last The activation function extracts high-dimensional data and obtains static local features. , where For voice features, is a one-dimensional convolutional layer, is the normalization layer, for Activation function, It is a static local feature.

[0075] For the static local features of the input speech features, the input speech features are first convolved to capture the local features of the audio, and then used Normalize and use it last The activation function extracts high-dimensional data and obtains static local features.

[0076] It is a one-dimensional convolution layer. The number of input and output channels is the same as the number of channels of the original input speech features, and the convolution kernel size is 7. The layer normalization layer normalizes all dimensions of the input speech features. is the activation function, which activates the deep features of speech features.

[0077] The static local feature extraction module extracts local features of the input speech features in a single frame. It first uses a one-dimensional convolution layer with a convolution kernel size of 7 to increase the receptive field of the model. Then, through layer normalization, all dimensions of the input speech data are normalized to avoid gradient vanishing or exploding problems during training. Finally, by using The activation function captures the deep features of speech features and has corresponding outputs for negative values, providing a more stable gradient flow for subsequent deep networks.

[0078] S4, fusing the speech feature, the enhanced dynamic global feature and the static local feature to obtain enhanced speech dynamic and static features. Figure 2 As shown, step S4 specifically includes step S41 to step S42.

[0079] S41, adding the enhanced dynamic global feature and the static local feature, and performing Operation to obtain the fused dynamic and static features. , where To integrate the dynamic and static features, is the normalized exponential function, To enhance dynamic global features, It is a static local feature.

[0080] After obtaining the enhanced dynamic global features and static local features, in order to simultaneously represent the dynamic global features and the static local features, this embodiment uses an addition operation and performs Operation to obtain the fused dynamic and static features.

[0081] S42, the fused dynamic and static features and the voice features are combined The product is used to obtain the enhanced speech dynamic and static features. , where To enhance the dynamic and static characteristics of speech, To integrate the dynamic and static features, for product, For the linear layer, For voice features.

[0082] Considering that the original speech features will still have an impact on subsequent training, this embodiment uses a linear layer to operate on the input speech features, and the number of input and output channels is consistent with the number of channels of the original speech features. Further, an operation similar to the self-attention mechanism is used to perform the fused dynamic and static features on them. The product is obtained by enhancing the dynamic and static features of speech, so as to further enhance the representation ability of the input speech.

[0083] S5, obtaining a mixed deformation parameter sequence with noise based on the facial mesh deformation parameters. Figure 2 As shown, step S5 specifically includes steps S51 to S54.

[0084] This embodiment uses a virtual human facial animation data mapping module to improve the mouth shape animation data set VOCASET based on mesh vertex data to generate a three-dimensional virtual human mouth shape animation based on mesh deformation parameters.

[0085] S51, obtaining a mouth shape animation data set based on mesh vertex data as a target model.

[0086] Specifically, before using the virtual human facial animation data mapping module, the mouth animation dataset VOCASET based on mesh vertex data is preprocessed to eliminate regional elements such as eyeballs, ears, back of the head and neck in VOCASET, thereby enhancing the stability of the three-dimensional virtual human face model generation process.

[0087] S52, obtaining a three-dimensional virtual human face model with mixed deformation parameters of a blended deformation deformer representing speech mouth shape drive as a source model.

[0088] Specifically, a 3D virtual human face model with 52 blendshape parameters is used as the source model, and 32 blendshape parameters are selected to represent key facial features in speech lip shape driving, such as the mouth, chin, cheeks, nose, etc.

[0089] S53, using the virtual human facial animation data mapping model, adjusting the shape of the source model to the shape of the target model, and obtaining a final hybrid deformation parameter sequence.

[0090] The dataset VOCASET is used as the target model, and the 3D virtual human face model with 32 hybrid deformation parameters is used as the source model. By using the virtual human facial animation data mapping module, the shape of the source model is adjusted to the shape of the target model to obtain the final hybrid deformation parameter sequence.

[0091] Assumptions For the corresponding The goal of the virtual human facial animation data mapping module is to generate an approximate hybrid deformation parameter sequence corresponding to the source model. , Indicates Mesh vertex data for frame facial animation, Indicates The blend shape parameter sequence of the frame facial animation. The length of the blend shape parameter sequence is 32.

[0092] This embodiment designs this process as a quadratic programming problem (i.e., a virtual human facial animation data mapping model), and realizes the smoothness of the mixed deformation parameter sequence by limiting the maximum transformation between two adjacent frames. By solving the quadratic programming problem of the virtual human facial animation data mapping model, the mapped mixed deformation parameter sequence is obtained.

[0093] Preferably, the virtual human facial animation data mapping model is:

[0094] .

[0095] In the formula, is the sequence of mixed deformation parameters after mapping, is the total number of frames of facial animation data, is the frame number, For the Frame facial mesh vertex data, is the standard face mesh vertex vector, is the total number of blend deformation parameters, is the sequence number of the blend deformation parameter, For the Frame No. The value of the blend shape parameter, For the The vector of facial mesh vertices corresponding to the blend shape parameters.

[0096] S54, adding noise to the final mixed deformation parameter sequence to obtain a mixed deformation parameter sequence with noise. Preferably, the noise adding model is:

[0097] .

[0098] In the formula, is a mixed deformation parameter sequence with noise, represents the noise adding module, is the sequence of mixed deformation parameters after mapping, is the noise intensity, is standard Gaussian noise.

[0099] S6. Input the enhanced speech dynamic and static features and the mixed deformation parameter sequence with noise into a noise reduction module to obtain a predicted mixed deformation parameter sequence.

[0100] Specifically, for data distribution, this embodiment uses an improved UNet network structure for training. Compared with the original UNet structure, this embodiment modifies the input from two-dimensional input to one-dimensional input according to the form of data. And deletes the downsampling and upsampling modules to reduce the model size. The specific structure is as follows Figure 2As shown, the denoising module includes a coding block, an intermediate block and a decoding block. The coding block, the intermediate block and the decoding block are connected in sequence. The coding block includes a convolutional layer, a residual layer and a Transformer layer connected in sequence. The intermediate block includes a residual layer and a Transformer layer connected in sequence. The decoding block includes a residual layer, a Transformer layer and a normalization layer connected in sequence.

[0101] The mixed deformation parameter sequence with noise is input into the convolution layer of the coding block. The coding block, the intermediate block and the decoded residual layer are all embedded with time steps. The enhanced speech dynamic and static features are input into the Transformer layers of the coding block, the intermediate block and the decoding block.

[0102] Furthermore, the mixed deformation parameter sequence with noise obtained by the above operation is input into the encoding block, the enhanced speech dynamic and static features are input into the Transformer layer of each module, and the time step is embedded into the residual layer of each module, and the conditional and unconditional linear combination sampling method is used for noise reduction processing to obtain the predicted mixed deformation parameter sequence. That is, the mouth shape animation controlled by the predicted mixed deformation parameter sequence is obtained through the noise reduction module.

[0103] The model of the noise reduction module is:

[0104] .

[0105] In the formula, is the predicted hybrid deformation parameter sequence, Indicates noise reduction, is a hyperparameter, is a mixed deformation parameter sequence with noise, To enhance the dynamic and static characteristics of speech, is the total time step, represents the weight of the linear combination, Indicates that the condition is empty.

[0106] The virtual human mouth shape driving method of the present embodiment that integrates dynamic and static features solves the problem of inaccurate mouth shape closure when the speech is silent by achieving a balanced transition between the global and local features of the speech.

[0107] In order to further improve the accuracy of mouth shape animation and reduce the abnormal jitter of mouth shape during large-scale animation switching, this embodiment introduces a deep smooth feature loss function. By adding a deep feature smooth loss function, the abnormal jitter of mouth shape is reduced, and a more accurate mouth shape animation is finally generated, which can be applied to voice-driven 3D virtual human mouth shape animation related technologies.

[0108] Based on the above embodiment, in an optional embodiment of the present invention, the loss function of the virtual human mouth shape driving method is:

[0109] .

[0110] .

[0111] .

[0112] In the formula, is the deep smoothing feature loss function, is the sample loss function, is the speed matching loss function, Express expectations, Represents the true sequence of blended deformation parameters for a single frame, represents the predicted hybrid deformation parameter sequence obtained conditionally by the parameters in brackets, is a mixed deformation parameter sequence with noise, To enhance the dynamic and static characteristics of speech, is the local time step, is the total number of frames of facial animation data, is the frame number, Indicates Frame noise mixed deformation parameter sequence, Indicates Frame noise mixed deformation parameter sequence, Represents the predicted hybrid deformation parameter sequence obtained according to the parameters in brackets.

[0113] During the training process, the sample loss function is used to select the absolute error between the minimum noise and the predicted noise as the optimization target for each time step, and the perceived distance between the predicted data and the real data distribution is continuously reduced. In this way, the accuracy of the lip animation is further improved during the continuous training process.

[0114] This embodiment also introduces a speed matching loss function to solve the problem of abnormal jitter in the generated mouth shape animation. This loss function smoothes the dynamic changes in the animation by minimizing the time difference distance between the noisy hybrid deformation parameter sequence and the noise reduction prediction value. Specifically, the time difference represents the rate of change of the model prediction value in continuous time steps, while the noise reduction prediction value represents the expected change trend in the real animation. By optimizing the gap between the two, the jitter phenomenon in the mouth shape animation can be significantly reduced, making the animation more coherent and smooth.

[0115] The deep smoothing feature loss function can significantly reduce jitter when generating lip animations. This function successfully reduces unnecessary dynamic noise by deeply smoothing the time series features of the generated animation, making the lip animation smoother and more natural. In addition, it also uses the perceptual distance as an indicator to measure the difference between the generated samples and the real data distribution. By reducing this distance, it effectively enhances the realism and stability of the generated lip animation.

[0116] The virtual human mouth shape driving method based on the fusion of dynamic and static features provided in this embodiment can well generate accurate three-dimensional virtual human mouth shape animation. The collaborative attention module of this method that integrates dynamic global features and static local features can effectively mine the global features and local features of speech, and effectively fuse them, thereby providing assistance for subsequent mouth shape animation generation.

[0117] In addition, this embodiment also designs a loss function of deep smoothness features, which improves the accuracy and richness of the generated mouth animation through the sample loss function, and reduces the abnormal jitter phenomenon of the mouth animation through the speed matching loss function, thereby generating natural and realistic three-dimensional virtual human mouth animation.

[0118] This embodiment uses an improved conditional denoising UNet model, modifies its main architecture, changes the input dimension from two dimensions to one dimension, and deletes the downsampling and upsampling modules in UNet to reduce the model size. Through these improvements, the scale of the model is greatly reduced, and the performance requirements for the running device are reduced.

[0119] This embodiment can well mine the deep features of speech, and generate 3D virtual human mouth shape animation in the form of facial mesh deformation parameters through the improved conditional denoising UNet model, and has strong scalability and can be expanded on any 3D virtual human. At the same time, this embodiment can generate 3D virtual human mouth shape animation in a relatively short time, improve the reaction speed of the user's interaction with the virtual human, and improve the user's experience.

[0120] Embodiment 2: The present invention provides a method and device for driving a virtual human mouth shape by fusing dynamic and static features, which comprises a speech feature extraction module, a global feature extraction module, a local feature extraction module, a feature fusion module, a mapping module and a noise reduction module.

[0121] The speech feature extraction module is used to obtain speech data and extract speech features.

[0122] The global feature extraction module is used to extract global information from the speech features to obtain enhanced dynamic global features.

[0123] The local feature extraction module is used to extract static local features from the speech features to obtain static local features.

[0124] The feature fusion module is used to fuse the speech feature, the enhanced dynamic global feature and the static local feature to obtain the enhanced speech dynamic and static feature.

[0125] A mapping module is used to obtain a noisy hybrid deformation parameter sequence based on the facial mesh deformation parameters.

[0126] The denoising module is used to input the enhanced speech dynamic and static features and the mixed deformation parameter sequence with noise into the denoising module to obtain a predicted mixed deformation parameter sequence. The denoising module includes a coding block, an intermediate block and a decoding block. The mixed deformation parameter sequence with noise is input into the convolution layer of the coding block. The enhanced speech dynamic and static features are input into the Transformer layers of the coding block, the intermediate block and the decoding block.

[0127] Embodiment 3, the present invention provides a method and device for driving a virtual human mouth shape by fusing static and dynamic features, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a method for driving a virtual human mouth shape by fusing static and dynamic features as described in any paragraph of Embodiment 1.

[0128] Embodiment 4: The present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a virtual human mouth shape driving method with dynamic and static feature fusion as described in any paragraph of Embodiment 1.

[0129] In several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the apparatus, method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0130] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0131] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code. It should be noted that in this article, the term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.

[0132] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0133] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0134] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0135] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0136] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A virtual human mouth shape driving method based on the fusion of dynamic and static features, characterized in that: Include: Acquire voice data and extract voice features; Extracting global information from the speech features to obtain enhanced dynamic global features; Extracting static local features from the speech features to obtain static local features; fusing the speech feature, the enhanced dynamic global feature and the static local feature to obtain enhanced speech dynamic and static features; Obtaining a mixed deformation parameter sequence with noise based on the facial mesh deformation parameters; The enhanced speech dynamic and static features and the mixed deformation parameter sequence with noise are input into a denoising module to obtain a predicted mixed deformation parameter sequence; wherein the denoising module includes a coding block, an intermediate block and a decoding block; the mixed deformation parameter sequence with noise is input into the convolution layer of the coding block; the enhanced speech dynamic and static features are input into the Transformer layers of the coding block, the intermediate block and the decoding block; Obtain a mixed deformation parameter sequence with noise based on the facial mesh deformation parameters, including: Obtain a mouth animation dataset based on mesh vertex data as a target model; A three-dimensional virtual human face model with mixed deformation parameters of a blended deformation deformer representing speech mouth shape drive is obtained as a source model; Using the virtual human facial animation data mapping model, the shape of the source model is adjusted to the shape of the target model to obtain the final hybrid deformation parameter sequence; Adding noise to the final mixed deformation parameter sequence to obtain a mixed deformation parameter sequence with noise; The virtual human facial animation data mapping model is: ; In the formula, is the sequence of mixed deformation parameters after mapping, is the total number of frames of facial animation data, is the frame number, For the Frame facial mesh vertex data, is the standard face mesh vertex vector, is the total number of blend deformation parameters, is the sequence number of the blend deformation parameter, For the Frame No. The value of the blend shape parameter, For the The facial mesh vertex vector corresponding to the blend deformation parameters; The noise addition model is: ; In the formula, is a mixed deformation parameter sequence with noise, represents the noise adding module, is the sequence of mixed deformation parameters after mapping, is the noise intensity, is standard Gaussian noise.

2. The method for driving a virtual human mouth shape by fusing dynamic and static features according to claim 1, characterized in that: Performing global information extraction on the speech features to obtain enhanced dynamic global features specifically includes: The speech features are first compressed by a linear layer to reduce the number of speech channels, and then the standard The activation function activates the deep features of the speech, and finally the number of speech channels is expanded back to the original number of channels through the linear layer to obtain dynamic global features. The dynamic global feature extraction model is: , where For dynamic global features, For voice features, For the linear layer, for Activation function; The dynamic global feature and the speech feature are combined Product, to obtain enhanced dynamic global features; where the product model is: , where To enhance dynamic global features, For dynamic global features, for product, For voice features.

3. The method for driving a virtual human mouth shape by fusing dynamic and static features according to claim 1, characterized in that: Extracting static local features from the speech features to obtain static local features specifically includes: The speech features are first convolved to capture the local features of the audio, and then used Normalize and use it last The activation function extracts high-dimensional data and obtains static local features; among them, , where For voice features, is a one-dimensional convolutional layer, is the normalization layer, for Activation function, It is a static local feature.

4. The method for driving a virtual human mouth shape by fusing dynamic and static features according to claim 1, characterized in that: The speech feature, the enhanced dynamic global feature and the static local feature are integrated to obtain the enhanced speech dynamic and static feature, specifically including: The enhanced dynamic global feature and the static local feature are added together and Operation, to obtain the fused dynamic and static features; among them, , where To integrate the dynamic and static features, is the normalized exponential function, To enhance dynamic global features, It is a static local feature; The fused dynamic and static features and the speech features are Product, to obtain enhanced speech dynamic and static features; where, , where To enhance the dynamic and static characteristics of speech, To integrate the dynamic and static features, for product, For the linear layer, For voice features.

5. The method for driving a virtual human mouth shape by fusing dynamic and static features according to claim 1, characterized in that: The encoding block, the intermediate block and the decoding block are connected in sequence; The encoding block includes a convolutional layer, a residual layer and a Transformer layer connected in sequence; The intermediate block includes a residual layer and a Transformer layer connected in sequence; The decoding block includes a residual layer, a Transformer layer and a normalization layer connected in sequence; The coding block, the intermediate block and the decoded residual layer are all embedded with a time step; The model of the noise reduction module is: ; In the formula, is the predicted hybrid deformation parameter sequence, Indicates noise reduction, is a hyperparameter, is a mixed deformation parameter sequence with noise, To enhance the dynamic and static characteristics of speech, is the total time step, represents the weight of the linear combination, Indicates that the condition is empty.

6. The method for driving a virtual human mouth shape by fusing dynamic and static features according to claim 1, characterized in that: The loss function of the virtual human mouth shape driving method is: ; ; ; In the formula, is the deep smoothing feature loss function, is the sample loss function, is the speed matching loss function, Express expectations, Represents the true sequence of blended deformation parameters for a single frame, represents the predicted hybrid deformation parameter sequence obtained conditionally by the parameters in brackets, is a mixed deformation parameter sequence with noise, To enhance the dynamic and static characteristics of speech, is the local time step, is the total number of frames of facial animation data, is the frame number, Indicates Frame noise mixed deformation parameter sequence, Indicates Frame noise mixed deformation parameter sequence, Represents the predicted hybrid deformation parameter sequence obtained according to the parameters in brackets.

7. A method and device for driving a virtual human mouth shape by fusing dynamic and static features, characterized in that: Include: A speech feature extraction module is used to obtain speech data and extract speech features; A global feature extraction module, used to extract global information from the speech features to obtain enhanced dynamic global features; A local feature extraction module, used to extract static local features from the speech features to obtain static local features; A feature fusion module, used to fuse the speech feature, the enhanced dynamic global feature and the static local feature to obtain enhanced speech dynamic and static features; A mapping module, used to obtain a mixed deformation parameter sequence with noise based on the facial mesh deformation parameters; A denoising module, used for inputting the enhanced speech dynamic and static features and the mixed deformation parameter sequence with noise into the denoising module to obtain a predicted mixed deformation parameter sequence; wherein the denoising module includes a coding block, an intermediate block and a decoding block; the mixed deformation parameter sequence with noise is input into the convolution layer of the coding block; the enhanced speech dynamic and static features are input into the Transformer layers of the coding block, the intermediate block and the decoding block; The mapping module is used to perform the following steps: Obtain a mouth animation dataset based on mesh vertex data as a target model; A three-dimensional virtual human face model with mixed deformation parameters of a blended deformation deformer representing speech mouth shape drive is obtained as a source model; Using the virtual human facial animation data mapping model, the shape of the source model is adjusted to the shape of the target model to obtain the final hybrid deformation parameter sequence; Adding noise to the final mixed deformation parameter sequence to obtain a mixed deformation parameter sequence with noise; The virtual human facial animation data mapping model is: ; In the formula, is the sequence of mixed deformation parameters after mapping, is the total number of frames of facial animation data, is the frame number, For the Frame facial mesh vertex data, is the standard face mesh vertex vector, is the total number of blend deformation parameters, is the sequence number of the blend deformation parameter, For the Frame No. The value of the blend shape parameter, For the The facial mesh vertex vector corresponding to the blend deformation parameters; The noise addition model is: ; In the formula, is a mixed deformation parameter sequence with noise, represents the noise adding module, is the sequence of mixed deformation parameters after mapping, is the noise intensity, is standard Gaussian noise.

8. A method and device for driving a virtual human mouth shape by fusing dynamic and static features, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement a virtual human mouth shape driving method with dynamic and static feature fusion as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a virtual human mouth shape driving method for fusion of dynamic and static features as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Linear complexity model architecture for speech recognition

    CN119252232A

  • Virtual image expression generation method and system based on real feeling technology

    CN119295683A