Audio-driven three-dimensional facial animation model generation method, generation apparatus, and apparatus

The method enhances audio-driven three-dimensional face video technology by training a model to accurately render facial expressions, addressing inaccuracies in mouth shape and movement, suitable for movie and TV production.

JP2025105542AActive Publication Date: 2025-07-10NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024226664
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-23
Publication Date
2025-07-10
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing audio-driven three-dimensional face video technologies suffer from inaccuracies in mouth shape and expression movement, leading to reduced accuracy and naturalness, particularly in applications like movie and TV production where line modifications require costly reshooting or rendering.

Method used

A method involving feature extraction, convolution, interval calculation, one-hot encoding, and superimposition of audio and speech style data to train an audio-driven three-dimensional face video model, optimizing BS values for precise facial expression rendering.

Benefits of technology

The trained model outputs high-precision BS values, enabling accurate and natural mouth shape and facial expression reproduction, suitable for high-accuracy applications such as movie and TV production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105542000001_ABST
    Figure 2025105542000001_ABST
Patent Text Reader

Abstract

To provide an audio-driven three-dimensional facial animation model generation method and a generation apparatus, and an apparatus.SOLUTION: A generation method includes the steps of: acquiring sample data including sample audio data, sample speaking style data, and a sample blend deformation value; performing feature extraction on the sample audio data to obtain a sample audio feature; performing convolution on the sample audio feature based on a to-be-trained audio-driven three-dimensional facial animation model to obtain an initial audio feature, and performing encoding on the sample speaking style data based on the to-be-trained audio-driven three-dimensional facial animation model to obtain a sample speaking style feature; performing encoding on the initial audio feature and the sample speaking style feature based on the to-be-trained audio-driven three-dimensional facial animation model, to obtain an output blend deformation value; and performing calculation on the sample blend deformation value and the output blend deformation value to obtain a loss function value.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular, to a method, apparatus, and device for generating an audio-driven three-dimensional face video model.

Background Art

[0002] The research on the technology of driving expressions with voice is an important content in the field of human-computer interaction. Specifically, a digital image is driven by voice, and the corresponding facial expression of this voice is used to express the content of the corresponding voice, and ultimately, the corresponding video is generated. The core of the technology of driving expressions with voice lies in the calculation and output of blend shape values (abbreviated as BS values). Specifically, the voice recorded by the user or the voice synthesized by TTS is preprocessed to output the BS value, and based on this BS value, the digital image is driven to generate a facial expression video corresponding to the voice, and then it can be displayed on various display devices by rendering technology.

[0003] In related technologies, during the production of works such as videos, movies, and TV dramas, it is often necessary to modify the lines in the later stage. However, due to the asynchronous phenomenon of voice and video caused by the modification of the lines, the effects of movies and TV dramas are significantly reduced. Therefore, in related technologies, methods such as reshooting or manufacturing and rendering in the later stage are used for processing, but the above methods all have the problem of extremely high costs. To solve the above problems, in related technologies, the modification of lines is realized by audio-driven three-dimensional face video technology, that is, the original image is driven by the audio corresponding to the modified lines to generate a facial expression video corresponding to the modified lines.

[0004] However, the BS values output by the above technology of driving expressions with voice have a series of problems such as the shape and movement of the mouth not corresponding to the voice and the unnaturalness of the movement of the expression, thereby reducing the accuracy of the audio-driven three-dimensional face video technology.

Summary of the Invention

Problems to be Solved by the Invention

[0005] In order to solve the above technical problems, the present invention provides a method, an apparatus and a device for generating an audio-driven three-dimensional face video model, which can improve the accuracy of the audio-driven three-dimensional face video technology. This technical solution is as follows.

Means for Solving the Problems

[0006] According to the present invention, a method for generating an audio-driven three-dimensional face video model is provided. This generation method includes: obtaining sample data including sample audio data, sample speech style data for depicting the facial expressions of a user, and a sample mixed deformation value obtained by preprocessing the sample audio data, wherein the sample audio data and the sample speech style data belong to the same user; performing feature extraction on the sample audio data to obtain sample audio features; performing convolution on the sample audio features based on an audio-driven three-dimensional face video model to be trained to obtain initial audio features, and performing encoding on the sample speech style data based on the audio-driven three-dimensional face video model to be trained to obtain sample speech style features; performing encoding on the initial audio features and the sample speech style features based on the audio-driven three-dimensional face video model to be trained to obtain an output mixed deformation value; calculating the sample mixed deformation value and the output mixed deformation value to obtain a loss function value; and updating the model parameters of the audio-driven three-dimensional face video model to be trained based on the loss function value.

[0007] In the first aspect, the step of performing feature extraction on the sample audio data to obtain sample audio features includes: performing feature extraction on the sample audio data based on a preset model, and using the features of the intermediate layer of the preset model as the sample audio features.

[0008] In the first aspect, the step of performing convolution on the sample audio features based on the audio-driven 3D face video model to be trained to obtain initial audio features includes: performing convolution on the sample audio features based on the audio-driven 3D face video model to be trained to obtain at least one intermediate audio feature; and performing interval calculation on the at least one intermediate audio feature to obtain the sample audio features.

[0009] In the first aspect, the step of performing interval calculation on the at least one intermediate audio feature to obtain the sample audio features includes: performing matching on every two of the at least one intermediate audio features based on the audio-driven 3D face video model to be trained to obtain an intermediate audio feature set corresponding to every two of the at least one intermediate audio features, where each intermediate audio feature among the at least one intermediate audio features corresponds to one convolution calculation channel, and the sequence values of the two convolution calculation channels corresponding to the intermediate audio feature set are not adjacent; and performing merging on every two of the intermediate audio features within each intermediate audio feature set among the at least one intermediate audio feature sets based on the audio-driven 3D face video model to be trained to obtain an intermediate merged feature corresponding to each intermediate audio feature set among the at least one intermediate audio feature sets. Performing calculations on each intermediate merging feature among at least one intermediate merging feature based on an audio-driven three-dimensional face video model to be trained to obtain sample audio features, and the like.

[0010] In the first aspect, the step of encoding sample speech style data based on an audio-driven three-dimensional face video model to be trained to obtain sample speech style features includes the step of performing one-hot encoding on sample speech style data based on an audio-driven three-dimensional face video model to be trained to obtain sample speech style features.

[0011] In the first aspect, the step of encoding initial audio features and sample speech style features based on an audio-driven three-dimensional face video model to be trained to obtain an output mixing deformation value includes the step of superimposing sample audio features and sample speech style features based on an audio-driven three-dimensional face video model to be trained to obtain sample superimposed features; and the step of encoding the sample superimposed features based on an audio-driven three-dimensional face video model to be trained to obtain an output mixing deformation value.

[0012] In the first aspect, the step of encoding sample superimposed features based on an audio-driven three-dimensional face video model to be trained to obtain an output mixing deformation value includes the step of encoding sample superimposed features based on an audio-driven three-dimensional face video model to be trained to obtain sample encoded features; and the step of decoding the sample encoded features based on an audio-driven three-dimensional face video model to be trained to obtain an output mixing deformation value.

[0013] According to the second aspect of the present invention, there is provided a generating device for an audio-driven three-dimensional face video model, and this generating device A module used to obtain sample data including sample audio data, sample speech style data for depicting a user's facial expression, and sample mixed deformation values obtained by preprocessing the sample audio data, where the sample audio data and the sample speech style data belong to the same user, an acquisition module, A feature extraction module for performing feature extraction on the sample audio data to obtain sample audio features, A first training module for performing convolution on the sample audio features based on an audio-driven 3D face video model to be trained to obtain initial audio features, and for encoding the sample speech style data based on the audio-driven 3D face video model to be trained to obtain sample speech style features, A second training module for encoding the initial audio features and the sample speech style features based on the audio-driven 3D face video model to be trained to obtain output mixed deformation values, A calculation module for calculating the sample mixed deformation values and the output mixed deformation values to obtain a loss function value, An update module for updating the model parameters of the audio-driven 3D face video model to be trained based on the loss function value, and includes.

[0014] In the second aspect, the feature extraction module Includes a feature extraction unit for performing feature extraction on the sample audio data based on a preset model and using the features of the intermediate layer of the preset model as the sample audio features.

[0015] In the second aspect, the first training module Includes a convolution unit for performing convolution on the sample audio features based on the audio-driven 3D face video model to be trained to obtain at least one intermediate audio feature, A distance calculation unit for performing distance calculation on at least one of the intermediate audio features to obtain sample audio features is included.

[0016] In a second aspect, this distance calculation unit A unit used to perform matching on every two of at least one intermediate audio feature among the intermediate audio features based on an audio-driven three-dimensional face video model to be trained, to obtain an intermediate audio feature set corresponding to every two of at least one intermediate audio feature, where each intermediate audio feature among at least one intermediate audio feature corresponds to one convolutional calculation channel, and the sequence values of the two convolutional calculation channels corresponding to the intermediate audio feature set are not adjacent, a first matching subunit A merging subunit for performing merging on every two of the intermediate audio features within each intermediate audio feature set among at least one intermediate audio feature set based on an audio-driven three-dimensional face video model to be trained, to obtain an intermediate merged feature corresponding to each intermediate audio feature set among at least one intermediate audio feature set A calculation subunit for performing calculation on each intermediate merged feature among at least one intermediate merged feature based on an audio-driven three-dimensional face video model to be trained, to obtain sample audio features, is included.

[0017] In a second aspect, the first training module Includes a one-hot encoding module for performing one-hot encoding on sample speaking style data based on an audio-driven three-dimensional face video model to be trained, to obtain sample speaking style features.

[0018] In a second aspect, the second training module Based on the audio-driven three-dimensional face video model to be trained, perform superposition on the sample audio features and the sample speaking style features to obtain a superposition unit for obtaining sample superposition features, and Based on the audio-driven three-dimensional face video model to be trained, perform encoding on the sample superposition features to obtain an encoding unit for obtaining an output mixing deformation value, and includes the above.

[0019] In the second aspect, the encoding unit Based on the audio-driven three-dimensional face video model to be trained, perform encoding on the sample superposition features to obtain an encoding sub-unit for obtaining sample encoding features, and Based on the audio-driven three-dimensional face video model to be trained, perform decoding on the sample encoding features to obtain a decoding sub-unit for obtaining an output mixing deformation value, and includes the above.

[0020] According to the third aspect of the present invention, an electronic device is provided. The electronic device includes a processor and a memory for storing at least one segment of a program. The at least one segment of the program is loaded and executed by the processor to implement the method for generating the audio-driven three-dimensional face video model provided in the above first aspect.

[0021] According to the fourth aspect of the present invention, a computer-readable storage medium is provided, which is characterized in that at least one segment of a program loaded and executed by a processor for implementing the method for generating the audio-driven three-dimensional face video model provided in the above first aspect is stored.

[0022] According to the fifth aspect of the present invention, a computer program product including computer instructions is provided. When the computer instructions are executed on an electronic device, the method for generating the audio-driven three-dimensional face video model provided in the above first aspect is implemented by the electronic device.

[0023] In the present invention, the above names do not limit the device or functional module itself. In actual implementation, these devices or functional modules may be represented by other names. As long as the functions of each device or functional module are similar to those of the present invention, they belong to the scope of the claims of the present invention and its technical equivalents.

[0024] The present invention will be made clearer and easier to understand as follows.

[0025] The present invention provides a method for generating an audio-driven three-dimensional face video model, which acquires user sample data including sample audio data, sample speaking style data, and sample mixing deformation values corresponding to the sample audio data. Here, in order to fuse the user's personalized data with the sample audio data and retain the user's personalized data as much as possible, the sample speaking style data is data that depicts the user's facial expressions. The trained audio-driven three-dimensional face video model obtained by training the audio-driven three-dimensional face video model to be trained based on the sample data can output a high-precision BS value, thereby improving the accuracy of the audio-driven three-dimensional face video technology. Furthermore, by inputting the output high-precision BS value into a preset unreal engine, the preset unreal engine can render a high-precision mouth shape and facial expressions on a video display device, thereby simultaneously realizing a high-precision reproduction of the mouth shape and facial expressions. As a result, the audio-driven three-dimensional face video technology can be widely applied to scenes with high requirements for expression accuracy, such as the production of movies and TV works.

Brief Description of the Drawings

[0026] To describe the present invention more clearly, the drawings in the following description are only examples of the embodiments of the present invention. A person skilled in the art can obtain other attached drawings based on these drawings without creative work.

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 2F

Figure 3

Figure 4

Figure 5

MODE FOR CARRYING OUT THE INVENTION

[0027] In order to make the object, technical solution and advantages of the present invention more clear, the present invention will be described in more detail below with reference to the drawings.

[0028] Examples in this specification are described in detail, and the examples are shown in the drawings. When the following description relates to the drawings, unless otherwise specified, the same numbers in different drawings indicate the same or similar elements. The following embodiments do not represent all embodiments that are consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with the present invention, as detailed in the claims.

[0029] Terms such as "first," "second," etc. in the present invention are used to distinguish the same or similar items whose actions and functions are basically the same. It should be understood that "first," "second," "nth" have no logical or chronological dependency and do not limit the quantity and execution order. In the following description, various elements are described using terms such as first, second, etc. It should also be understood that these elements are not limited by the terms.

[0030] These terms are merely used to distinguish one element from another. For example, without departing from the scope of various examples, the first operation can be called the second operation, and similarly, the second operation can also be called the first operation. Both the first operation and the second operation may be operations, and in some cases, they may be separate and different operations.

[0031] Here, "at least one" means one or more. For example, at least one operation can be one operation, two operations, three operations, etc., any operation that is one or an integer greater than one. "A plurality" means two or more. For example, a plurality of operations can be two operations, three operations, etc., any operation that is two or an integer greater than two.

[0032] It should be noted that the data (including, but not limited to, training data and prediction data such as user data, terminal-side data, etc.) and signals related to the present invention are all either permitted by the user or sufficiently permitted by each party, and it is necessary to collect, use, and process the related data in compliance with the relevant laws, regulations, and standards of the relevant countries and regions.

[0033] FIG. 1 is a schematic diagram of an implementation environment according to the present invention, and this implementation environment may include a terminal 101 and a server 102.

[0034] The terminal 101 is provided with an audio receiving device and a video display device. Here, the audio receiving device and the video display device may be independent devices respectively, or may be integrated into one hardware device. This hardware device has, for example, an audio collection function and a video display function integrated, such as an LED (Light Emitting Diode) screen with a voice recognition function. The terminal 101 may be a wearable device, a personal computer, a laptop portable computer, a tablet computer, a smart TV, an in-vehicle terminal, etc.

[0035] The server 102 may be a single server, a server cluster composed of multiple servers, or a cloud processing center.

[0036] The terminal 101 is connected to the server 102 via a wired or wireless network.

[0037] And standard communication technologies and / or protocols are used in this wireless network or wired network. The network is generally the Internet, but may be any network including, but not limited to, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network, or any combination thereof. It also represents data exchanged via a network using technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. Furthermore, all or part of the link can be encrypted using conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. Here, instead of or in addition to the above data communication technologies, customized and / or private data communication technologies can also be used.

[0038] FIG. 2A is a schematic flowchart of a method for generating an audio-driven three-dimensional face motion video model according to the present invention. As shown in FIG. 2A, the present invention will be described by taking as an example its application to a device having an audio receiving device and a video display device. This method includes steps 201 to 206.

[0039] In step 201, the terminal acquires sample data including sample audio data, sample speech style data for depicting the user's facial expression, and sample mixed deformation values obtained by preprocessing the sample audio data. The sample audio data and the sample speech style data belong to the same user.

[0040] Note that the mixed deformation value can also be expressed as a BS value.

[0041] As shown in FIG. 2B, the above step 201 includes steps 2011 to 2013.

[0042] In step 2011, the terminal obtains initial data.

[0043] Then, by presetting at least one fixed speech sample (for example, 100 pieces) and allocating this at least one fixed speech sample to at least one user, each user among the at least one user reads the speech sample allocated to himself / herself towards the data collection device in the same environment, expresses the preset facial expression, and further obtains the initial data through the data collection device.

[0044] In addition, based on the structured light in the data collection device and the built-in augmented reality development platform, the terminal performs real-time facial capture on each user among the at least one user, obtains at least one frame of image corresponding to each user among the at least one user, and obtains and records the audio corresponding to each frame of image among the at least one frame of images and at least one BS value. Here, each frame of image among the at least one frame of images is used to depict the facial expression of the corresponding user. Here, each frame of image among the at least one frame of images, the audio corresponding to each frame of image among the at least one frame of images, and at least one BS value are taken as one piece of initial data. That is, at least one piece of initial data can be obtained through the above-mentioned data collection device. In the present invention, the data collection device is not specifically limited.

[0045] Optionally, any one of the at least one BS value can be obtained by performing a mixing process on the audio to obtain the relative ratio or weight between audio signals. Here, the relative ratio or weight is the BS value.

[0046] Note that based on the conventional algorithm, mixing processing can be performed on the audio to obtain the BS value corresponding to the audio. In the present invention, the description of the specific process of the mixing processing is omitted.

[0047] In step 2012, the terminal performs filtering on the initial data at least once to obtain filtered data.

[0048] Then, based on preset parameters such as environmental error and / or user error, filtering is performed on the obtained initial data, at least one frame of the image where user error and / or environmental error occurs is deleted, and the audio and BS value corresponding to at least one frame of the image are also deleted to obtain high-quality filtered data.

[0049] In step 2013, the terminal optimizes the filtered data to obtain sample data.

[0050] Then, the images with inaccurate facial expression representations of some users are adjusted in the video production method to realize the optimization of the images with inaccurate facial expression representations of the users. Furthermore, relatively accurate sample facial expression data is obtained, and based on the sample facial expression data, sample speech style data is obtained, thereby obtaining relatively accurate sample data.

[0051] Optionally, when obtaining sample speech style data based on the sample facial expression data, taking each user as an example, each frame image of at least one frame of the image depicting the sample facial expression data of the user is judged to determine the user's speech style, and the sample speech style data of the user is obtained. Furthermore, the sample facial expression data of the user is retained as much as possible. For example, the sample speech style data includes "exaggerated" and "gentle".

[0052] In step 202, the terminal performs feature extraction on the sample audio data to obtain sample audio features.

[0053] And the above step 202 includes the step that the terminal performs feature extraction on the sample audio data based on the preset model and uses the features of the intermediate layer of the preset model as the sample audio features. Optionally, the preset model includes an input layer, an intermediate layer, and an output layer. Here, the intermediate layer includes a plurality of data processing sub-layers. Optionally, when obtaining the sample audio features, the features of one data processing sub-layer in the intermediate layer can be used as the sample audio features.

[0054] Optionally, since the audio receiving devices and the user's audio sources used to collect the user's data are diverse, there is at least one type of audio data format in the collected sample audio data, and furthermore, it is disadvantageous for the unified processing of the sample audio data, thereby causing great inconvenience in the processing of the sample audio data. To solve the above technical problems, a general-purpose audio feature extraction method is used to perform feature extraction on the sample audio data. For example, an asr (Automatic Speech Recognition) model is used to perform feature extraction on the sample audio data to obtain sample audio features.

[0055] Optionally, when performing feature extraction on the sample audio data, any one of models such as masr (Magical Automatic Speech Recognition) and deepspeech (Deep Speech) can be used to perform feature extraction on the sample audio data.

[0056] Optionally, since there is at least one language corresponding to at least one audio receiving device and at least one user in the sample audio data, when performing feature extraction on the sample audio data, in order to improve the generality of feature extraction and further facilitate the improvement of the efficiency of feature extraction, the features of the intermediate layer of the masr model are used as the sample audio features.

[0057] In step 203, the terminal performs convolution on the sample audio features based on the audio-driven 3D face video model to be trained to obtain initial audio features, and encodes the sample speech style data based on the audio-driven 3D face video model to be trained to obtain sample speech style features.

[0058] Furthermore, before step 203, the step of the terminal inputting the sample audio features, the sample speech style data and the corresponding sample BS value into the audio-driven 3D face video model to be trained is included.

[0059] And, as shown in FIG. 2C, the above step 203 includes step 2031 and step 2032.

[0060] In step 2031, the terminal performs convolution on the sample audio features based on the audio-driven 3D face video model to be trained to obtain at least one intermediate audio feature.

[0061] And when the terminal performs convolution on the sample audio features based on the audio-driven 3D face video model to be trained, the terminal obtains at least one convolution set, and each convolution set in the at least one convolution set shares hyperparameters and includes at least one convolutional layer. In each convolution set of the at least one convolution set, a convolution calculation channel for calculating the channels of the input features (a part of the sample audio features) is provided.

[0062] In step 2032, the terminal performs interval calculation on at least one intermediate audio feature to obtain a sample audio feature.

[0063] And each convolutional layer among at least one convolutional layer includes at least one convolutional kernel, and different convolutional kernels correspond to different convolutional effects, but the convolutional parameters are valid. Therefore, there may be a certain correlation between the intermediate audio features corresponding to different convolutional sets among at least one convolutional set. Furthermore, when calculating the intermediate audio features corresponding to different convolutional sets, it will lead to an overfitting effect, thereby reducing the accuracy of the calculation. In the present invention, in order to solve the above problem and improve the accuracy of calculation, the terminal performs interval calculation on at least one intermediate audio feature to obtain a sample audio feature.

[0064] And as shown in FIG. 2D, the above step 2032 includes steps 20321 to 20323.

[0065] In step 20321, the terminal performs matching on two intermediate audio features out of at least one intermediate audio feature based on an audio-driven 3D face video model to be trained, and obtains an intermediate audio feature set corresponding to two intermediate audio features out of at least one intermediate audio feature. Here, each intermediate audio feature among at least one intermediate audio feature corresponds to one convolutional calculation channel, and the sequence values of the two convolutional calculation channels corresponding to the intermediate audio feature set are not adjacent.

[0066] Then, taking one of at least one convolutional set as an example for explanation, to perform convolution on corresponding input features, the convolutional set calculates the corresponding convolutional calculation channel and the channel of the corresponding input features to obtain the convolution result of the input features. Optionally, taking a part of the sample audio features as an example for explanation, when the convolutional set performs convolution on a part of the sample audio features, the convolutional calculation channel corresponding to the convolutional set and the channel corresponding to a part of the sample audio features are calculated, and further, intermediate audio features corresponding to a part of the sample audio features are obtained.

[0067] Optionally, taking two of at least one convolutional set as an example for explanation, when the sequence values of the convolutional calculation channels corresponding to these two convolutional sets are adjacent, when performing calculations on the two intermediate audio features corresponding to these two convolutional sets, the problem of overfitting occurs. Optionally, to solve the problem of overfitting, in the present invention, based on the sequence values of the convolutional calculation channels of the convolutional sets corresponding to each intermediate audio feature among at least one intermediate audio feature, two-by-two matching is performed on each intermediate audio feature among at least one intermediate audio feature to obtain at least one intermediate audio feature set.

[0068] Then, taking one of at least one intermediate audio feature set as an example for interpretation and explanation, the intermediate audio feature set corresponds to two convolutional sets, these two convolutional sets correspond to two convolutional calculation channels, and these two convolutional calculation channels are not adjacent, that is, the difference in the sequence values of these two convolutional calculation channels is greater than "1". When performing calculations on the two intermediate audio features in the intermediate audio feature set, overfitting does not occur, so the problem of overfitting is solved.

[0069] When performing matching on two intermediate audio features out of at least one intermediate audio feature based on the preset first matching rule, the difference in the sequence values of the convolutional calculation channels corresponding to two intermediate audio features out of at least one intermediate audio feature is greater than "1". Optionally, taking two intermediate audio feature sets out of at least one intermediate audio feature set as an example for interpretation and explanation, the sequence value of the convolutional calculation channel of the convolutional set corresponding to one intermediate audio feature is "1", and the sequence value of the convolutional calculation channel of the convolutional set corresponding to another intermediate audio feature is "6". Match the two intermediate audio feature sets to obtain one intermediate audio feature set.

[0070] In step 20322, based on the audio-driven 3D face video model for which the terminal is the training target, perform merging on two intermediate audio features within each intermediate audio feature set out of at least one intermediate audio feature set to obtain intermediate merged features corresponding to each intermediate audio feature set out of at least one intermediate audio feature set.

[0071] As can be seen from the analysis in step 20321, each intermediate merged feature among at least one intermediate merged feature obtained in step 20322 has no overfitting problem, and furthermore, the accuracy of at least one intermediate merged feature is improved.

[0072] In step 20323, perform calculation on each intermediate merged feature among at least one intermediate merged feature based on the audio-driven 3D face video model for which the terminal is the training target to obtain this sample audio feature.

[0073] And when training the audio-driven 3D face video model for which the terminal is the training target, in order to improve the accuracy of the trained audio-driven 3D face video model, it is necessary to perform iterative training. Optionally, in order to make it possible to more uniformly fuse the convolution results of different convolution sets among at least one convolution set, after step 20323 is completed and before the next training is performed, based on the preset second matching rule, perform matching again on every two of the at least one intermediate audio feature, so as to obtain an intermediate audio feature set corresponding to every two of the at least one intermediate audio feature.

[0074] Optionally, after step 20323 is completed, during the process of performing the next training, based on the preset third matching rule, perform matching again on every two of the at least one intermediate audio feature, so as to obtain an intermediate audio feature set corresponding to every two of the at least one intermediate audio feature.

[0075] Optionally, taking two of the at least one intermediate audio feature set as an example for interpretation and explanation, match two intermediate audio feature sets where the sequence value of the convolution calculation channel of the convolution set corresponding to one intermediate audio feature is "1" and the sequence value of the convolution calculation channel of the convolution set corresponding to another intermediate audio feature is "3", so as to obtain one intermediate audio feature set.

[0076] In step 20323, since there is no overfitting in each of the at least one intermediate merging feature, there is also no overfitting in the sample audio features obtained based on each of the at least one intermediate merging feature. Furthermore, the accuracy of the sample audio features is improved.

[0077] And, step 203 above includes the step of performing one-hot encoding on the sample speech style data based on the audio-driven three-dimensional face video model for which the terminal is the training target to obtain sample speech style features.

[0078] Note that one-hot encoding is a common encoding method, and in the embodiments of the present invention, a detailed description of the process of one-hot encoding is omitted.

[0079] In step 204, encoding is performed on the initial audio features and the sample speech style features based on the audio-driven three-dimensional face video model for which the terminal is the training target to obtain an output mixing deformation value.

[0080] And, as shown in FIG. 2E, step 204 above includes step 2041 and step 2042.

[0081] In step 2041, superimposing is performed on the sample audio features and the sample speech style features based on the audio-driven three-dimensional face video model for which the terminal is the training target to obtain sample superimposed features.

[0082] In step 2041, superimposing is performed on the sample audio features and the sample speech style features in the channel dimension based on the audio-driven three-dimensional face video model for which the terminal is the training target, ensuring that the sample superimposed features contain rich audio information and further contain the individualized information of the user corresponding to the sample superimposed features, and further, it is used to realize good reproduction of the mouth shape video and facial expressions.

[0083] In step 2042, encoding is performed on the sample superimposed features based on the audio-driven three-dimensional face video model for which the terminal is the training target to obtain an output mixing deformation value.

[0084] And, as shown in FIG. 2F, step 2042 above includes step 20421 and step 20422.

[0085] In step 20421, the terminal encodes the sample superimposed features based on the audio-driven three-dimensional face video model to be trained, and obtains sample encoded features.

[0086] Then, the terminal encodes the sample superimposed features based on the audio-driven three-dimensional face video model to be trained with the preset encoding model.

[0087] Optionally, in the process of encoding the sample superimposed features based on the preset encoding model, the sample superimposed features are divided into different layers, and each layer represents different granularities. For example, downsampling can be gradually performed on the sample superimposed features to obtain displays of different layers.

[0088] Optionally, for one layer, the preset encoding model uses an independent self-attention mechanism to capture the internal relationship of the layer. This hierarchical self-attention calculates a weight matrix based on the position, and then multiplies it with the value matrix within the layer to obtain the self-attention display within the layer. To enable the preset encoding model to simultaneously focus on information of different granularities, the self-attention displays of different layers are concatenated together. Optionally, the self-attention displays of different layers can be concatenated together by splicing or weighted averaging.

[0089] Optionally, different layers can interact with each other, so that the preset encoding model can simultaneously capture features of at least one granularity, and the representation of the preset encoding model can be improved.

[0090] Optionally, in the preset encoding model, while maintaining performance, the number of parameters is reduced as much as possible, and in order to learn a more refined representation from the teacher model, the model parameters are compressed in a patience-intensive knowledge distillation manner. Next, by reducing the feed-forward neural network layers and increasing more self-attention mechanisms (e.g., hierarchical self-attention mechanisms, processed information with different granularities), the preset encoding model can better understand the input features. Furthermore, by reducing the redundant encoder and decoder layers, the size of the preset encoding model becomes smaller, the inference speed is improved, and thus the influence relationship between marks can be better understood.

[0091] Optionally, the preset encoding model may be an X-Transformer (a Transformer fine-tuned in depth for large-scale multi-label learning (XMC)).

[0092] Note that the patience-intensive knowledge distillation method and the teacher model may be general methods or models, and in the embodiments of the present invention, detailed descriptions are omitted.

[0093] As can be seen from the above analysis, when encoding the sample superimposed features based on the improved preset encoding model, high-precision sample encoding features can be obtained more quickly.

[0094] In step 20422, the terminal decodes the sample encoding features based on the audio-driven 3D face video model to be trained to obtain an output mixed deformation value.

[0095] And based on the decoder of the Transformer, the sample encoding features can be decoded to obtain an output mixed deformation value.

[0096] Note that the decoder of the Transformer decodes the sample encoding features based on a general decoding algorithm, and in the embodiments of the present invention, detailed descriptions of the decoding process of the sample encoding features are omitted.

[0097] In step 205, the terminal calculates for the sample mixing deformation value and the output mixing deformation value to obtain a loss function value.

[0098] Then, calculations are performed on the sample mixing deformation value and the output mixing deformation value based on a preset loss function to obtain a loss function value. Optionally, the preset loss function may be an L1 loss function (absolute value error loss function).

[0099] Optionally, in order to improve the accuracy of the trained audio-driven 3D face video model, first to third order errors are calculated. Here, the first to third order errors include a reconstruction error, a speed error, and an acceleration error.

[0100] In step 206, the terminal updates the model parameters of the audio-driven 3D face video model to be trained based on the loss function value.

[0101] And in step 201, at least one sample data can be further obtained. After obtaining at least one sample data, it facilitates performing iterative training on the audio-driven 3D face video model to be trained with at least one sample data, and to further improve the accuracy of the trained audio-driven 3D face video model, the above steps 202 to 206 are repeatedly executed.

[0102] After obtaining the trained audio-driven 3D face video model, the audio receiver device acquires the audio to be processed, performs feature extraction on the audio to be processed based on a preset model to obtain the audio features to be processed, acquires preset speech style data, inputs the audio features to be processed and the preset speech style data into the trained audio-driven 3D face video model, and outputs at least one corresponding BS value. Here, the image of each frame corresponds to at least one BS value, the complexity of the image is different, and the quantity of the corresponding BS values is also different.

[0103] Then, after obtaining the corresponding BS value, the BS value is input into a preset unreal engine. There are already various preset scenes in the preset unreal engine. After receiving the BS value, the preset unreal engine renders at least one frame target image (including the shape of a high-precision mouth and facial expressions simultaneously) corresponding to the audio feature to be processed to a video display device based on the various preset scenes and the BS value. Here, the facial expression of the person in the target image matches the preset speaking style data. For example, when the preset speaking style data is "exaggerated", the person in the target image corresponds to an exaggerated facial expression.

[0104] Note that the above embodiments have been described by taking their application to a terminal as an example. Also, the embodiments of the present invention can also be applied to a server.

[0105] The embodiments of the present invention acquire user sample data including sample audio data, sample speaking style data, and sample mixing deformation values corresponding to the sample audio data. Here, in order to fuse the user's personalized data with the sample audio data and retain the user's personalized data as much as possible, the user's facial expression data is depicted with the sample speaking style data. The trained audio-driven 3D face animation model obtained by training the audio-driven 3D face animation model to be trained based on the sample data can output a high-precision BS value, thereby improving the accuracy of the audio-driven 3D face animation technology. Furthermore, by inputting the output high-precision BS value into the preset unreal engine, the preset unreal engine can render the high-precision mouth shape and facial expressions to the video display device, thereby realizing the simultaneous high-precision reproduction of the mouth shape and facial expressions. As a result, the audio-driven 3D face video technology can be widely applied to scenes with high requirements for expression accuracy, such as the production of movies and TV works.

[0106] FIG. 3 is a schematic structural diagram of a generating device 300 for an audio-driven 3D face video model according to an embodiment of the present invention. The device includes: An acquisition module 301 used to obtain sample data including sample audio data, sample speaking style data, and sample mixed deformation values obtained by preprocessing the sample audio data, where the sample audio data and the sample speaking style data belong to the same user, and A feature extraction module 302 for performing feature extraction on the sample audio data to obtain sample audio features, and A first training module 303 for performing convolution on the sample audio features based on the audio-driven 3D face video model to be trained to obtain initial audio features, and performing encoding on the sample speaking style data based on the audio-driven 3D face video model to be trained to obtain sample speaking style features, and A second training module 304 for performing encoding on the initial audio features and the sample speaking style features based on the audio-driven 3D face video model to be trained to obtain output mixed deformation values, and A calculation module 305 for calculating the sample mixed deformation values and the output mixed deformation values to obtain a loss function value, and An update module 306 for updating the model parameters of the audio-driven 3D face video model to be trained based on this loss function value.

[0107] In one possible implementation form, the feature extraction module 302 Based on a preset model, it includes a feature extraction unit for performing feature extraction on sample audio data and using the features of the intermediate layer of the preset model as sample audio features.

[0108] In one possible implementation form, the first training module 303 includes a convolution unit for performing convolution on sample audio features based on an audio-driven 3D face video model to be trained to obtain at least one intermediate audio feature, and an interval calculation unit for performing interval calculation on at least one intermediate audio feature to obtain sample audio features.

[0109] In one possible implementation form, the interval calculation unit is a unit used to perform matching on two intermediate audio features out of at least one intermediate audio feature based on an audio-driven 3D face video model to be trained, so as to obtain an intermediate audio feature set corresponding to two intermediate audio features out of at least one intermediate audio feature. Here, the channel sequence values of two intermediate audio features in the intermediate audio feature set are not adjacent. It is the first matching sub-unit, a merging sub-unit for performing merging on two intermediate audio features in each intermediate audio feature set out of at least one intermediate audio feature set based on this audio-driven 3D face video model to be trained, so as to obtain an intermediate merged feature corresponding to each intermediate audio feature set out of at least one intermediate audio feature set, and a calculation sub-unit for performing calculation on each intermediate merged feature out of at least one intermediate merged feature based on an audio-driven 3D face video model to be trained to obtain sample audio features.

[0110] In one possible implementation form, the first training module 303 It includes a one-hot encoding module that performs one-hot encoding on sample speech style data based on an audio-driven three-dimensional face video model to be trained, so as to obtain sample speech style features for depicting the user's speech style.

[0111] In one possible implementation form, the second training module 304 includes a superimposing unit that superimposes sample audio features and sample speech style features based on an audio-driven three-dimensional face video model to be trained, so as to obtain sample superimposed features, and an encoding unit that encodes the sample superimposed features based on an audio-driven three-dimensional face video model to be trained, so as to obtain an output mixing deformation value.

[0112] In one possible implementation form, the encoding unit includes an encoding sub-unit that encodes the sample superimposed features based on an audio-driven three-dimensional face video model to be trained, so as to obtain sample encoded features, and a decoding sub-unit that decodes the sample encoded features based on an audio-driven three-dimensional face video model to be trained, so as to obtain an output mixing deformation value.

[0113] It should be noted that when the generation device of the audio-driven three-dimensional face video model provided in the above embodiments executes the corresponding steps, the division of each of the above function modules is only taken as an example for explanation. In actual applications, if necessary, the above functions can be assigned to different function modules to be completed, that is, the internal structure of the device can be divided into different function modules to complete all or part of the functions described above. In addition, the generation device of the audio-driven three-dimensional face video model provided in the above embodiments belongs to the same concept as the embodiments of the generation method of the audio-driven three-dimensional face video model. The specific implementation process thereof can refer to the embodiments of the method, and the detailed description is omitted here.

[0114] The present invention obtains user sample data including sample audio data, sample speech style data, and sample mixing deformation values corresponding to the sample audio data. Here, in order to fuse the user's personalized data with the sample audio data and retain the user's personalized data as much as possible, the user's facial expression data is depicted with the sample speech style data. The trained audio-driven three-dimensional face animation model, which is a training target based on the sample data, can output a high-precision BS value, thereby improving the accuracy of the audio-driven three-dimensional face animation technology. Furthermore, by inputting the output high-precision BS value into a preset unreal engine, the preset unreal engine can render a high-precision mouth shape and facial expression on a video display device, thereby realizing the simultaneous high-precision reproduction of the mouth shape and facial expression. As a result, the audio-driven three-dimensional face animation technology can be widely applied to scenes with high requirements for expression accuracy, such as the production of movies and TV works.

[0115] The present invention further provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for generating the above-mentioned audio-driven three-dimensional face animation model is implemented.

[0116] Taking the example that the electronic device is a terminal, FIG. 4 is a schematic structural diagram of the terminal provided by the present invention. Referring to FIG. 4, the terminal 400 may be a smartphone, a tablet, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 400 may also be called by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0117] Generally, the terminal 400 includes a processor 401 and a memory 402.

[0118] The processor 401 may include one or more processing cores such as a 4-core processor. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 401 may include a main processor and a coprocessor. The main processor is a processor for processing data in the wake-up state and is also called a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in the standby state. Here, the processor 401 may integrate a GPU (Graphics Processing Unit) for rendering and drawing the content that needs to be displayed on the display screen. Furthermore, the processor 401 may further include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.

[0119] The memory 402 may include one or more non-transitory computer-readable storage media. The memory 402 may further include a high-speed random access memory and non-volatile memory such as one or more disk storage devices and flash memory storage devices. The non-transitory computer-readable storage media in the memory 402 is used to store at least one program code, and this at least one program code is used to be executed by the processor 401 so as to implement the process executed by the terminal among the methods for generating an audio-driven three-dimensional face animation model provided by the method of the present invention.

[0120] Furthermore, the terminal 400 selectively includes a peripheral device interface 403 and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface 403 can be connected by a bus or signal lines. Each peripheral device can be connected to the peripheral device interface 403 by a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 404, a display screen 405, a camera assembly 406, an audio circuit 407, and a power source 408.

[0121] The peripheral device interface 403 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 401 and the memory 402. The processor 401, the memory 402, and the peripheral device interface 403 are integrated on the same chip or circuit board. In some other embodiments, any one or two of the processor 401, the memory 402, and the peripheral device interface 403 may be implemented on a separate chip or circuit board, and the present invention is not limited thereto.

[0122] The radio frequency circuit 404 is used to receive and transmit RF (Radio Frequency) signals, which are also called electromagnetic signals. The radio frequency circuit 404 communicates with a communication network and other communication devices via electromagnetic signals. The radio frequency circuit 404 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. The radio frequency circuit 404 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user ID module card, etc. The radio frequency circuit 404 can communicate with other terminals according to at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to, a metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. The radio frequency circuit 404 may further include a circuit related to NFC (Near Field Communication), and the present invention is not limited thereto.

[0123] The display screen 405 is used to display a UI (User Interface). This UI can include graphics, text, icons, videos, and any combination thereof. When the display screen 405 is a touch panel display screen, the display screen 405 further has the ability to collect touch signals on or above the surface of the display screen 405. This touch signal can be input to the processor 401 as a control signal and processed. At this time, the display screen 405 can be further used to provide virtual buttons and / or a virtual keyboard, which are also called soft buttons and / or a soft keyboard. And the display screen 405 may be one and is provided on the front panel of the terminal 400. In some other embodiments, the display screen 405 may be at least two and are respectively provided on different surfaces of the terminal 400, or present a folded design. In some other embodiments, the display screen 405 may be a flexible display screen provided on the curved surface or the folded surface of the terminal 400. Furthermore, the display screen 405 may also be provided as an irregular shape other than a rectangle, that is, an irregular screen. The display screen 405 can be manufactured using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0124] The camera assembly 406 is used to collect images or videos. And the camera assembly 406 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. Here, there are at least two rear cameras, each of which is any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera. Accordingly, the main camera and the depth-of-field camera are fused to realize a background blur function, and the main camera and the wide-angle camera are fused to realize a panorama shooting and VR (Virtual Reality) shooting function or other fused shooting functions. And the camera assembly 406 may further include a flash. The flash may be a single-color temperature flash or a two-color temperature flash. The two-color temperature flash refers to a combination of a warm-color light flash and a cold-color light flash that can be used for light correction at different color temperatures.

[0125] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, convert the sound waves into electrical signals, and input the electrical signals into the processor 401 for processing, or input the electrical signals into the radio frequency circuit 404 to realize voice communication. For the purpose of stereo sound collection or noise reduction, there may be a plurality of microphones, which are respectively provided at different parts of the terminal 400. The microphone may also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signal from the processor 401 or the radio frequency circuit 404 into a sound wave. The speaker may be a conventional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into a sound wave audible to humans, but also convert the electrical signal into a sound wave inaudible to humans for applications such as distance measurement. In some embodiments, the audio circuit 407 may further include a headphone jack.

[0126] The power supply 408 is used to supply power to each component in the terminal 400. The power supply 408 may be an alternating current, direct current, disposable battery, or rechargeable battery. When the power supply 408 includes a rechargeable battery, this rechargeable battery can support wired charging or wireless charging. This rechargeable battery can also be used to support fast charging technology.

[0127] And the terminal 400 further includes one or more sensors 409. The one or more sensors 409 include, but are not limited to, an acceleration sensor 410, a gyro sensor 411, a pressure sensor 412, an optical sensor 413, and a proximity sensor 414.

[0128] The acceleration sensor 410 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established in the terminal 400. For example, the acceleration sensor 410 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 401 can control the display screen 405 to display the user screen in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 410. The acceleration sensor 410 can also be used for collecting game or user motion data.

[0129] The gyro sensor 411 can detect the orientation and rotation angle of the machine body of the terminal 400, and the gyro sensor 411 can cooperate with the acceleration sensor 410 to obtain the three-dimensional motion of the user with respect to the terminal 400. The processor 401 can realize motion sensing (for example, changing the UI according to the user's tilting operation), the image stabilization function during shooting, the game control function, and the inertial navigation function based on the data collected by the gyro sensor 411.

[0130] The pressure sensor 412 may be provided under the side frame of the terminal 400 and / or the display screen 405. When the pressure sensor 412 is provided on the side frame of the terminal 400, it can detect the gripping signal of the terminal 400 by the user, and based on the gripping signal collected by the pressure sensor 412, the left and right hand recognition or shortcut operation can be performed by the processor 401. When the pressure sensor 412 is provided under the display screen 405, the processor 401 realizes the control of the operable controls on the UI screen based on the user's pressure operation on the display screen 405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0131] The optical sensor 413 is used to collect the intensity of ambient light. In one embodiment, the processor 401 can control the display brightness of the display screen 405 based on the intensity of the ambient light collected by the optical sensor 413. Specifically, when the intensity of the ambient light is high, the display brightness of the display screen 405 is adjusted to be high, and when the intensity of the ambient light is low, the display brightness of the display screen 405 is adjusted to be low. In another embodiment, the processor 401 can further dynamically adjust the shooting parameters of the camera assembly 406 based on the intensity of the ambient light collected by the optical sensor 413.

[0132] The proximity sensor 414, also called a distance sensor, is usually provided on the front panel of the terminal 400. The proximity sensor 414 is used to collect the distance between the user and the front of the terminal 400. In one embodiment, when the proximity sensor 414 detects that the distance between the user and the front of the terminal 400 is gradually decreasing, the processor 401 controls the display screen 405 to switch from a bright screen state to a dark screen (screen off), and when the proximity sensor 414 detects that the distance between the user and the front of the terminal 400 is gradually increasing, the processor 401 controls the display screen 405 to switch from a dark screen state to a bright screen.

[0133] The structure shown in FIG. 4 does not limit the terminal 400, and may include more or fewer components than shown, some components may be combined, or different component arrangements may be adopted.

[0134] Taking the example that the electronic device is a server, FIG. 5 is a structural schematic diagram of a server provided in an embodiment of the present invention. Although this server 500 may vary greatly in terms of configuration or performance, it may include one or more central processing units (CPUs) 501 and one or more memories 502. Here, at least one computer program is included in the one or more memories 502, and at least one computer program is loaded and executed by one or more processors 501 to implement the above method for generating an audio-driven 3D face animation model. This server 500 may further include components such as a wired or wireless network interface, a keyboard, and an input / output interface to facilitate input and output. This server 500 may further include other components for realizing device functions, and details are omitted here.

[0135] Embodiments of the present invention further provide a computer-readable storage medium including a stored computer program. When the computer program is executed, a device with the computer-readable storage medium is controlled to execute the above method for generating an image processing model. Optionally, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, optical data storage device, etc.

[0136] All or part of the steps of the above embodiments may be implemented in hardware, or may be implemented by instructing related hardware with a program. The program may be stored in a computer-readable storage medium, and the storage medium mentioned above may be a read-only memory, magnetic disk, compact disc, etc.

[0137] The above content is only an embodiment of the present invention and does not limit the present invention. Modifications, equivalent substitutions, improvements, etc. made within the spirit and principle scope of the present invention are included in the protection scope of this application.

Claims

1. Obtaining sample data including sample audio data, sample speech style data for depicting a user's facial expression, and a sample mixing deformation value obtained by preprocessing the sample audio data, wherein the sample audio data and the sample speech style data belong to the same user; Performing feature extraction on the sample audio data to obtain sample audio features; Performing convolution on the sample audio features based on an audio-driven three-dimensional face video model to be trained to obtain initial audio features, and performing encoding on the sample speech style data based on the audio-driven three-dimensional face video model to be trained to obtain sample speech style features; Performing encoding on the initial audio features and the sample speech style features based on the audio-driven three-dimensional face video model to be trained to obtain an output mixing deformation value; Calculating the sample mixing deformation value and the output mixing deformation value to obtain a loss function value; Updating the model parameters of the audio-driven three-dimensional face video model to be trained based on the loss function value. A method for generating an audio-driven three-dimensional face video model, characterized by comprising the above steps.

2. The step of performing feature extraction on the sample audio data to obtain sample audio features comprises: Performing feature extraction on the sample audio data based on a preset model, and using the features of the intermediate layer of the preset model as the sample audio features. The method according to claim 1.

3. The step of performing convolution on the sample audio features based on the audio-driven three-dimensional face video model to be trained to obtain initial audio features comprises: Performing convolution on the sample audio features based on the audio-driven three-dimensional face video model to be trained to obtain at least one intermediate audio feature; Performing interval calculation on the at least one intermediate audio feature to obtain the sample audio features. The method according to claim 1.

4. The step of obtaining the sample audio feature by performing interval calculation on the at least one intermediate audio feature is: A step of performing matching on two intermediate audio features out of the at least one intermediate audio feature based on the audio-driven three-dimensional face video model to be trained, and obtaining an intermediate audio feature set corresponding to two intermediate audio features out of the at least one intermediate audio feature, wherein each intermediate audio feature out of the at least one intermediate audio feature corresponds to one convolution calculation channel, and the sequence values of the two convolution calculation channels corresponding to the intermediate audio feature set are not adjacent. A step of performing merging on two intermediate audio features within each intermediate audio feature set out of the at least one intermediate audio feature set based on the audio-driven three-dimensional face video model to be trained, and obtaining an intermediate merged feature corresponding to each intermediate audio feature set out of the at least one intermediate audio feature set. A step of performing calculation on each intermediate merged feature out of the at least one intermediate merged feature based on the audio-driven three-dimensional face video model to be trained, and obtaining the sample audio feature. The method according to claim 3 is characterized by including the above steps.

5. The step of obtaining the sample speaking style feature by encoding the sample speaking style data based on the audio-driven three-dimensional face video model to be trained is: The method according to claim 1 is characterized by including the step of obtaining the sample speaking style feature by performing one-hot encoding on the sample speaking style data based on the audio-driven three-dimensional face video model to be trained.

6. The step of obtaining the output mixing deformation value by encoding the initial audio feature and the sample speaking style feature based on the audio-driven three-dimensional face video model to be trained is: A step of superimposing the sample audio feature and the sample speaking style feature based on the audio-driven three-dimensional face video model to be trained, and obtaining a sample superimposed feature. performing encoding on the sample superposition features based on the audio-driven three-dimensional face video model to be trained to obtain the output mixing deformation value, and the method according to claim 1, characterized in that it includes this step.

7. the step of performing encoding on the sample superposition features based on the audio-driven three-dimensional face video model to be trained to obtain the output mixing deformation value is performing encoding on the sample superposition features based on the audio-driven three-dimensional face video model to be trained to obtain sample encoded features, performing decoding on the sample encoded features based on the audio-driven three-dimensional face video model to be trained to obtain the output mixing deformation value, and the method according to claim 6, characterized in that it includes this step.

8. a module used to obtain sample data including sample audio data, sample talking style data for depicting the facial expressions of the user, and sample mixing deformation values obtained by preprocessing the sample audio data, wherein the sample audio data and the sample talking style data belong to the same user, an acquisition module, a feature extraction module for performing feature extraction on the sample audio data to obtain sample audio features, a first training module for performing convolution on the sample audio features based on the audio-driven three-dimensional face video model to be trained to obtain initial audio features, and performing encoding on the sample talking style data based on the audio-driven three-dimensional face video model to be trained to obtain sample talking style features, a second training module for performing encoding on the initial audio features and the sample talking style features based on the audio-driven three-dimensional face video model to be trained to obtain the output mixing deformation value, a calculation module for calculating the sample mixing deformation value and the output mixing deformation value to obtain a loss function value, an update module for updating the model parameters of the audio-driven three-dimensional face video model to be trained based on the loss function value, and a generating device for the audio-driven three-dimensional face video model, characterized in that it includes this module.

9. An electronic device including a processor and a memory for storing a program of at least one segment, wherein the program of the at least one segment is loaded and executed by the processor, and realizes the method for generating an audio-driven three-dimensional face video model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that it stores at least one segment of a program that is loaded and executed by a processor and realizes the method for generating an audio-driven three-dimensional face video model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Mixed deformation value output method and device, storage medium and electronic device

    CN113592985A

  • Method, device, equipment, and computer program for driving the movement of a target object

    JP2023545642A