Facial data generation method and device, computer device, and storage medium

By extracting motion features of the lower half of the face, the upper half of the face, and the head under audio-driven conditions, facial offset is generated, which solves the problems of insufficient lip matching and inconsistent expressions in the existing technology and achieves high-quality 3D facial animation generation.

CN116246328BActive Publication Date: 2026-01-23TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310220611.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2026-01-23
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

In existing audio-driven 3D facial animation generation technologies, the matching degree between lip shape and input audio is insufficient, resulting in inconsistent and unnatural facial expressions.

Method used

By extracting the audio features of the input audio and the facial features of the standard facial data of the virtual object, motion features of the lower half of the face, the upper half of the face, and the head are generated respectively. Based on these features, facial offsets are generated, and finally, predicted facial data is synthesized to ensure that the lip shape is highly matched with the input audio and the expression is coherent and natural.

Benefits of technology

It improves the matching degree between lip shape and audio in 3D facial animation, and ensures the continuity and naturalness of expressions, thereby improving the quality of facial animation generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246328B_ABST
    Figure CN116246328B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a face data generation method and device, computer equipment and a storage medium, and belongs to the technical field of computers. The method comprises: generating lower half face movement features, upper half face movement features and head movement features of a virtual object when the virtual object emits input audio based on audio features of the input audio and face features of standard face data of the virtual object; generating a face offset of the virtual object when the virtual object emits the input audio based on the lower half face movement features, the upper half face movement features and the head movement features; and generating predicted face data of the virtual object when the virtual object emits the input audio based on the face offset and the standard face data. The present disclosure takes into account different reflection modes of different regions of the face for the input audio, and the movement features of each region have better sensitivity and adaptability to the input audio, thereby ensuring that the expression changes of the virtual object in the face animation generated by the multiple frames of predicted face data are coherent and natural.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to a face data generation method and device, computer equipment and storage medium. BACKGROUND

[0002] With the development of computer technology, audio-driven 3D (three-dimensional) face animation generation technology is widely used in the fields of movies, games, e-commerce, AR (Augmented Reality) and VR (Virtual Reality).

[0003] The audio-driven 3D face animation generation technology refers to that, given an input audio and a face template, the face template is driven to generate a 3D face animation with a lip shape matching the input audio under the action of the input audio. At present, the matching degree of the lip shape and the input audio in the synthesized 3D face animation needs to be improved, and the expression is not natural enough. SUMMARY

[0004] The present disclosure provides a face data generation method and device, computer equipment and storage medium to at least improve the matching degree of the lip shape and the input audio in the predicted face data, so that the expression is natural and coherent. The technical solutions of the present disclosure are as follows:

[0005] According to an aspect of an embodiment of the present disclosure, a face data generation method is provided, comprising:

[0006] extracting audio features of the input audio and face features of standard face data of a virtual object;

[0007] generating lower half face motion features, upper half face motion features and head motion features of the virtual object when the virtual object emits the input audio based on the audio features and the face features;

[0008] generating a face offset of the virtual object when the virtual object emits the input audio based on the lower half face motion features, the upper half face motion features and the head motion features, the face offset indicating a face offset and a head offset of the virtual object relative to the standard face data when the virtual object emits the input audio;

[0009] generating predicted face data of the virtual object when the virtual object emits the input audio based on the face offset and the standard face data.

[0010] In some embodiments, the generating the lower half face motion features of the virtual object when the virtual object emits the input audio based on the audio features and the face features comprises:

[0011] fusing the audio features and the face features to obtain initial fusion features;

[0012] encoding based on the initial fusion feature, to obtain the lower half face motion feature, the lower half face motion feature representing a motion feature of a lower half face region of the virtual object when the virtual object pronounces the input audio.

[0013] In some embodiments, the encoding based on the initial fusion feature, to obtain the lower half face motion feature includes:

[0014] inputting the initial fusion feature into a lower half face encoder, and encoding the initial fusion feature by the lower half face encoder to output the lower half face motion feature, the lower half face encoder being configured to extract a motion feature of a lower half face region of the virtual object when the virtual object pronounces.

[0015] In some embodiments, the generating the upper half face motion feature of the virtual object when the virtual object pronounces the input audio based on the audio feature and the face feature includes:

[0016] extracting the upper half face motion feature based on the lower half face motion feature, the upper half face motion feature representing a motion feature of an upper half face region of the virtual object when the virtual object pronounces the input audio.

[0017] In some embodiments, the extracting the upper half face motion feature based on the lower half face motion feature includes:

[0018] inputting an initial fusion feature into a first encoder, and encoding the initial fusion feature by the first encoder to output an intermediate encoding feature, the first encoder being configured to extract an intermediate level of encoding feature based on an initial fusion feature of an initial level between audio and face, the initial fusion feature being fused based on the audio feature and the face feature;

[0019] fusing the intermediate encoding feature and the lower half face motion feature to obtain a lower half face fusion feature;

[0020] inputting the lower half face fusion feature into an upper half face encoder, and encoding the lower half face fusion feature by the upper half face encoder to output the upper half face motion feature, the upper half face encoder being configured to extract a motion feature of an upper half face region of the virtual object when the virtual object pronounces.

[0021] In some embodiments, the generating the head motion feature of the virtual object when the virtual object pronounces the input audio based on the audio feature and the face feature includes:

[0022] extracting the head motion feature based on the upper half face motion feature, the head motion feature representing a motion feature of a head region of the virtual object when the virtual object pronounces the input audio.

[0023] In some embodiments, the extracting the head motion feature based on the upper half face motion feature comprises:

[0024] inputting the intermediate encoded feature into a second encoder, encoding the intermediate encoded feature through the second encoder to output a final layer encoded feature, the second encoder being configured to extract a highest level encoded feature based on encoded features at an intermediate level between audio and face, the intermediate encoded feature being obtained based on an initial fusion feature, the initial fusion feature being obtained based on fusion of the audio feature and the face feature;

[0025] fusing the final layer encoded feature and the upper half face motion feature to obtain an upper half face fusion feature;

[0026] inputting the upper half face fusion feature into a head encoder, encoding the upper half face fusion feature through the head encoder to output the head motion feature, the head encoder being configured to extract a motion feature of a head region of the virtual object when the virtual object utters.

[0027] In some embodiments, the generating the face offset of the virtual object when the virtual object utters the input audio based on the lower half face motion feature, the upper half face motion feature and the head motion feature comprises:

[0028] generating a global motion feature based on an initial fusion feature, the lower half face motion feature, the upper half face motion feature and the head motion feature, the initial fusion feature being obtained based on fusion of the audio feature and the face feature;

[0029] generating a lower half face offset based on the global motion feature and the lower half face motion feature;

[0030] generating an upper half face offset based on the global motion feature and the upper half face motion feature;

[0031] generating a head offset based on the global motion feature and the head motion feature;

[0032] generating the face offset based on the lower half face offset, the upper half face offset and the head offset.

[0033] In some embodiments, the generating the global motion feature based on the initial fusion feature, the lower half face motion feature, the upper half face motion feature and the head motion feature comprises:

[0034] fusing the initial fusion feature, the lower half face motion feature, the upper half face motion feature and the head motion feature to obtain a global fusion feature;

[0035] input the global fusion feature into a global encoder, encode the global fusion feature through the global encoder, and output the global motion feature, the global encoder being configured to extract a motion feature of a lower half face region, an upper half face region, and a head region of the virtual object in the global when the virtual object is speaking.

[0036] In some embodiments, the generating the lower half face offset based on the global motion feature and the lower half face motion feature comprises:

[0037] fusing the global motion feature and the lower half face motion feature to obtain a lower half face enhanced feature;

[0038] inputting the lower half face enhanced feature into a lower half face refining encoder, encoding the lower half face enhanced feature through the lower half face refining encoder, and outputting a lower half face refining feature, the lower half face refining encoder being configured to refine the lower half face enhanced feature of the virtual object when the virtual object is speaking;

[0039] inputting the lower half face refining feature into a lower half face decoder, decoding the lower half face refining feature through the lower half face decoder, and outputting the lower half face offset, the lower half face decoder being configured to decode the lower half face refining feature into the lower half face offset of the virtual object relative to the standard face data when the virtual object is speaking.

[0040] In some embodiments, the generating the upper half face offset based on the global motion feature and the upper half face motion feature comprises:

[0041] fusing the global motion feature and the upper half face motion feature to obtain an upper half face enhanced feature;

[0042] inputting the upper half face enhanced feature into an upper half face refining encoder, encoding the upper half face enhanced feature through the upper half face refining encoder, and outputting an upper half face refining feature, the upper half face refining encoder being configured to refine the upper half face enhanced feature of the virtual object when the virtual object is speaking;

[0043] inputting the upper half face refining feature into an upper half face decoder, decoding the upper half face refining feature through the upper half face decoder, and outputting the upper half face offset, the upper half face decoder being configured to decode the upper half face refining feature into the upper half face offset of the virtual object relative to the standard face data when the virtual object is speaking.

[0044] In some embodiments, the generating the head offset based on the global motion feature and the head motion feature comprises:

[0045] fusing the global motion feature and the head motion feature to obtain a head enhanced feature;

[0046] The head enhancement features are input into the head refinement encoder, which encodes the head enhancement features and outputs the head refinement features. The head refinement encoder is used to refine the head enhancement features of the virtual object when it is speaking.

[0047] The head refinement features are input into the head decoder, which decodes the head refinement features and outputs the head offset. The head decoder is used to decode the head refinement features into a head offset that represents the virtual object's head offset relative to the standard facial data when it is speaking.

[0048] In some embodiments, generating the facial offset based on the lower face offset, the upper face offset, and the head offset includes:

[0049] Based on the lower half face mask image of the standard facial data, the lower half face offset is weighted to obtain the lower half face weighted offset.

[0050] Based on the upper half face mask image of the standard facial data, the upper half face offset is weighted to obtain the upper half face weighted offset.

[0051] Based on the head mask image of the standard facial data, the head offset is weighted to obtain the head weighted offset.

[0052] The weighted offset of the lower half of the face, the weighted offset of the upper half of the face, and the weighted offset of the head are combined to obtain the facial offset.

[0053] In some embodiments, generating predicted facial data for the virtual object when it emits the input audio, based on the facial offset and the standard facial data, includes:

[0054] Based on the facial offset, the facial feature points in the standard facial data are offset to obtain the predicted facial data.

[0055] In some embodiments, the facial features for extracting standard facial data of the virtual object include:

[0056] The standard facial data is input into a point cloud encoder, which encodes the standard facial data and outputs the facial features. The point cloud encoder is used to extract facial features from the facial data of the virtual object.

[0057] In some embodiments, the extraction of audio features from the input audio includes:

[0058] The input audio is fed into a temporal convolution model, and the temporal convolution model performs temporal convolution on at least one audio frame in the input audio to obtain the audio features of the at least one audio frame.

[0059] In some embodiments, the method further includes:

[0060] Based on the audio features of each audio frame in the input audio and the facial features, a frame of predicted facial data associated with the audio frame is generated;

[0061] The facial animation of the virtual object is synthesized based on multi-frame predicted facial data.

[0062] According to another aspect of the embodiments of this disclosure, a facial data generation apparatus is provided, comprising:

[0063] The extraction unit is configured to extract audio features from the input audio and facial features from standard facial data of the virtual object;

[0064] The motion feature generation unit is configured to generate, based on the audio features and the facial features, lower half face motion features, upper half face motion features, and head motion features when the virtual object emits the input audio;

[0065] The offset generation unit is configured to generate a facial offset when the virtual object emits the input audio based on the lower half face motion features, the upper half face motion features, and the head motion features. The facial offset indicates the facial and head offsets of the virtual object relative to the standard facial data when it emits the input audio.

[0066] A facial data generation unit is configured to generate predicted facial data when the virtual object emits the input audio, based on the facial offset and the standard facial data.

[0067] In some embodiments, the motion feature generation unit includes:

[0068] The fusion subunit is configured to perform the fusion of the audio features and the facial features to obtain an initial fused feature;

[0069] The encoding subunit is configured to perform encoding based on the initial fusion features to obtain the lower face motion features, which characterize the motion features of the lower face region when the virtual object emits the input audio.

[0070] In some embodiments, the coding subunit is configured to perform:

[0071] The initial fusion features are input into the lower face encoder, which encodes the initial fusion features and outputs the lower face motion features. The lower face encoder is used to extract the motion features of the lower face region of the virtual object when it speaks.

[0072] In some embodiments, the motion feature generation unit includes:

[0073] The first extraction subunit is configured to extract the upper face motion features based on the lower face motion features, wherein the upper face motion features characterize the motion features of the upper face region when the virtual object emits the input audio.

[0074] In some embodiments, the first extraction subunit is configured to perform:

[0075] An initial fusion feature is input into a first encoder, which encodes the initial fusion feature and outputs intermediate encoded features. The first encoder is used to extract intermediate encoded features based on the initial fusion features at the initial level between audio and facial features. The initial fusion feature is obtained by fusing the audio features and the facial features.

[0076] The intermediate encoded features and the lower half of the face motion features are fused to obtain the lower half of the face fusion features;

[0077] The lower half face fusion feature is input into the upper half face encoder, and the upper half face encoder encodes the lower half face fusion feature to output the upper half face motion feature. The upper half face encoder is used to extract the motion feature of the upper half face region of the virtual object when it speaks.

[0078] In some embodiments, the motion feature generation unit includes:

[0079] The second extraction subunit is configured to extract the head motion features based on the upper half face motion features, wherein the head motion features characterize the motion features of the head region when the virtual object emits the input audio.

[0080] In some embodiments, the second extraction subunit is configured to perform:

[0081] The intermediate coding features are input into the second encoder, which encodes the intermediate coding features and outputs the final layer coding features. The second encoder is used to extract the highest layer coding features based on the intermediate layer coding features between the audio and the face. The intermediate coding features are obtained based on the initial fusion features, which are obtained by fusing the audio features and the facial features.

[0082] The final layer coding features and the upper half face motion features are fused to obtain the upper half face fusion features;

[0083] The upper half face fusion features are input into the head encoder, which encodes the upper half face fusion features and outputs the head motion features. The head encoder is used to extract the motion features of the head region of the virtual object when it speaks.

[0084] In some embodiments, the offset generation unit includes:

[0085] The first generation subunit is configured to generate global motion features based on initial fusion features, the lower half face motion features, the upper half face motion features, and the head motion features, wherein the initial fusion features are obtained by fusing the audio features and the facial features;

[0086] The second generation subunit is configured to generate a lower face offset based on the global motion features and the lower face motion features.

[0087] The third generation subunit is configured to generate an upper face offset based on the global motion features and the upper face motion features.

[0088] The fourth generation subunit is configured to generate a head offset based on the global motion features and the head motion features;

[0089] The fifth generation subunit is configured to generate the facial offset based on the lower half face offset, the upper half face offset, and the head offset.

[0090] In some embodiments, the first generation subunit is configured to perform:

[0091] The initial fusion feature, the lower half face motion feature, the upper half face motion feature, and the head motion feature are fused to obtain the global fusion feature;

[0092] The global fusion feature is input into the global encoder, and the global fusion feature is encoded by the global encoder to output the global motion feature. The global encoder is used to extract the global motion features of the lower half of the face, the upper half of the face, and the head region of the virtual object when it speaks.

[0093] In some embodiments, the second generation subunit is configured to perform:

[0094] The global motion features and the lower face motion features are fused to obtain the lower face enhancement features;

[0095] The lower face enhancement features are input into the lower face refinement encoder, which encodes the lower face enhancement features and outputs the lower face refinement features. The lower face refinement encoder is used to refine the lower face enhancement features of the virtual object when it is speaking.

[0096] The lower face refinement features are input into the lower face decoder, which decodes the lower face refinement features and outputs the lower face offset. The lower face decoder is used to decode the lower face refinement features into a representation of the lower face offset of the virtual object relative to the standard facial data when it is speaking.

[0097] In some embodiments, the third generation subunit is configured to perform:

[0098] The global motion features and the upper half face motion features are fused to obtain the upper half face enhancement features;

[0099] The upper half face enhancement features are input into the upper half face refinement encoder, the upper half face enhancement features are encoded by the upper half face refinement encoder, and the upper half face refinement features are output. The upper half face refinement encoder is used to refine the upper half face enhancement features of the virtual object when it speaks.

[0100] The upper half face refinement features are input into the upper half face decoder, and the upper half face refinement features are decoded by the upper half face decoder to output the upper half face offset. The upper half face decoder is used to decode the upper half face refinement features into an upper half face offset that represents the virtual object relative to the standard facial data when it is speaking.

[0101] In some embodiments, the fourth generation subunit is configured to perform:

[0102] The global motion features and the head motion features are fused to obtain head enhancement features;

[0103] The head enhancement features are input into the head refinement encoder, which encodes the head enhancement features and outputs the head refinement features. The head refinement encoder is used to refine the head enhancement features of the virtual object when it is speaking.

[0104] The head refinement features are input into the head decoder, which decodes the head refinement features and outputs the head offset. The head decoder is used to decode the head refinement features into a head offset that represents the virtual object's head offset relative to the standard facial data when it is speaking.

[0105] In some embodiments, the fifth generation subunit is configured to perform:

[0106] Based on the lower half face mask image of the standard facial data, the lower half face offset is weighted to obtain the lower half face weighted offset.

[0107] Based on the upper half face mask image of the standard facial data, the upper half face offset is weighted to obtain the upper half face weighted offset.

[0108] Based on the head mask image of the standard facial data, the head offset is weighted to obtain the head weighted offset.

[0109] The weighted offset of the lower half of the face, the weighted offset of the upper half of the face, and the weighted offset of the head are combined to obtain the facial offset.

[0110] In some embodiments, the facial data generation unit is configured to perform:

[0111] Based on the facial offset, the facial feature points in the standard facial data are offset to obtain the predicted facial data.

[0112] In some embodiments, the extraction unit is configured to perform:

[0113] The standard facial data is input into a point cloud encoder, which encodes the standard facial data and outputs the facial features. The point cloud encoder is used to extract facial features from the facial data of the virtual object.

[0114] In some embodiments, the extraction unit is further configured to perform:

[0115] The input audio is fed into a temporal convolution model, and the temporal convolution model performs temporal convolution on at least one audio frame in the input audio to obtain the audio features of the at least one audio frame.

[0116] In some embodiments, the apparatus further includes an animation compositing unit configured to perform:

[0117] Based on the audio features of each audio frame in the input audio and the facial features, a frame of predicted facial data associated with the audio frame is generated;

[0118] The facial animation of the virtual object is synthesized based on multi-frame predicted facial data.

[0119] According to another aspect of the embodiments of this disclosure, a computer device is provided, comprising:

[0120] One or more processors;

[0121] One or more memories for storing the one or more processor-executable instructions;

[0122] The one or more processors are configured to perform the facial data generation method in any of the possible implementations of the above-described aspects.

[0123] According to another aspect of the present disclosure, a computer-readable storage medium is provided such that, when at least one instruction in the computer-readable storage medium is executed by one or more processors of a computer device, the computer device is enabled to perform a facial data generation method in any possible implementation of the above aspect.

[0124] According to another aspect of the present disclosure, a computer program product is provided, including one or more instructions that can be executed by one or more processors of a computer device, enabling the computer device to perform the facial data generation method in any possible implementation of the above aspect.

[0125] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0126] Based on the audio features of the input audio and the facial features of the standard facial data of the virtual object, the face is divided into three regions: the lower half of the face, the upper half of the face, and the head. Motion features of each region are extracted to obtain lower half face motion features, upper half face motion features, and head motion features. This takes into account the different ways different facial regions respond to the input audio, making the motion features of each region more sensitive and adaptable to the input audio. Then, the motion features of the lower half of the face, the upper half of the face, and the head are integrated and finely adjusted to generate a facial offset, which indicates the facial and head offsets of the virtual object relative to the standard facial data when it emits the input audio. Finally, predicted facial data is generated based on the standard facial data and the facial offset, ensuring that the predicted facial data, especially the lip shape, upper half of the face, and head, have a high degree of adaptability to the input audio. This ensures that the facial expressions of the virtual object in the facial animation generated based on multi-frame predicted facial data are coherent and natural.

[0127] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0128] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0129] Figure 1 This is a schematic diagram illustrating the implementation environment of a facial data generation method according to an embodiment of the present disclosure;

[0130] Figure 2 This is a flowchart illustrating a facial data generation method according to an embodiment of the present disclosure;

[0131] Figure 3 This is an interactive flowchart illustrating a facial data generation method according to an embodiment of the present disclosure;

[0132] Figure 4 This is a schematic diagram illustrating the generation principle of predictive facial data according to an embodiment of the present disclosure;

[0133] Figure 5 This is a logical structure block diagram of a facial data generation device according to an embodiment of the present disclosure;

[0134] Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present disclosure. Detailed Implementation

[0135] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0136] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0137] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the input audio and standard facial data involved in this disclosure were obtained with full authorization.

[0138] In some embodiments, the meaning of A and / or B includes three cases: A and B, and A and B.

[0139] Audio-driven 3D facial animation generation technology refers to the process where, given an input audio clip and a facial template, the facial template is driven to generate a 3D facial animation that matches the lip shape of the input audio clip. This technology is widely used in film, games, e-commerce, AR, VR and other fields, providing automated assistance and support for practitioners to reduce their workload.

[0140] Currently, the matching degree between lip shapes and input audio in machine-generated 3D facial animations needs improvement, and the expressions are not always coherent and natural. Considering that users' visual systems are highly sensitive, they can quickly capture and grasp the generation quality of each frame of facial data in a facial animation, the synchronization of lip sounds, and the accuracy and naturalness of facial expression changes. Even slight changes in details can affect the overall quality of the generated facial animation. Therefore, improving the generation quality of audio-driven 3D facial animations is a challenging task.

[0141] In view of this, the present disclosure provides a facial data generation method that can synthesize predicted facial data for each frame in a 3D facial animation under the drive of input audio, ensuring that the lip shape in the predicted facial data is highly matched with the sound in the input audio, and that the facial expression changes of the virtual object in the 3D facial animation synthesized from multiple frames of predicted facial data are coherent and natural.

[0142] The system architecture of the embodiments of this disclosure will be described below.

[0143] Figure 1 This is a schematic diagram illustrating the implementation environment of a facial data generation method according to an embodiment of this disclosure. See also... Figure 1 This implementation environment may include at least one terminal 101 and a server 102, as detailed below:

[0144] Terminal 101 has applications that support 3D animation installed and running. Optionally, such applications include, but are not limited to, animation applications, audio and video applications, game applications, live streaming applications, social applications, content sharing applications, browser applications, etc. This embodiment of the disclosure does not specifically limit the type of application.

[0145] Terminal 101 communicates directly or indirectly with server 102 via wired or wireless communication.

[0146] Server 102 includes at least one of a single server, multiple servers, a cloud computing platform, or a virtualization center. Server 102 is used to provide backend services for applications supporting 3D animation. Optionally, server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, terminal 101 and server 102 use a distributed computing architecture for collaborative computing.

[0147] Optionally, server 102 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0148] The device type of terminal 101 includes, but is not limited to, at least one of the following: smartphone, tablet computer, smart speaker, smartwatch, smart handheld device, portable gaming device, in-vehicle terminal, laptop computer, and desktop computer. The following embodiments use a smartphone as an example.

[0149] Those skilled in the art will understand that the number of terminals 101 described above may be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more. This disclosure does not limit the number of terminals or the type of device.

[0150] The basic process of the facial data generation scheme according to the present disclosure will be briefly introduced below.

[0151] Figure 2 This is a flowchart illustrating a facial data generation method according to an embodiment of this disclosure, see [link to flowchart]. Figure 2 The facial data generation method is performed by a computer device, which will be explained below using the computer device as a server as an example.

[0152] In step 201, the server extracts the audio features of the input audio and the facial features of the standard facial data of the virtual object.

[0153] In some embodiments, the server obtains input audio and standard facial data of a virtual object, extracts audio features from the input audio, and extracts facial features of the virtual object from the standard facial data of the virtual object.

[0154] In some embodiments, the server obtains the input audio from a local audio database, or the server downloads the input audio from a cloud audio database, or the server obtains the input audio sent by the terminal. This disclosure does not specifically limit the source of the input audio.

[0155] In some embodiments, the server obtains standard facial data of the virtual object from a local facial database, or the server downloads standard facial data of the virtual object from a cloud facial database, or the server obtains standard facial data of the virtual object sent by the terminal. This disclosure does not specifically limit the source of the standard facial data.

[0156] The aforementioned standard facial data can be a facial template drawn by an expert for the virtual object, or a facial template simulated by a machine for the virtual object. Furthermore, for a virtual object, its standard facial data refers to the facial data of the virtual object without any expression. The virtual object can be a virtual character, an anime character, a virtual animal, etc., and this disclosure does not specifically limit this. For example, if the virtual object is a three-dimensional model, and this three-dimensional model is a three-dimensional character constructed based on three-dimensional human skeleton technology, then the standard facial data of the virtual object is the facial data of the three-dimensional model without any expression.

[0157] In some embodiments, the server may preprocess the input audio, such as performing Voice Activity Detection (VAD), pre-emphasis, framing, windowing, etc., so that the input audio is divided into at least one audio frame, and then audio features are extracted from each audio frame in the at least one audio frame. For example, the frequency features of each audio frame are extracted as audio features after converting each audio frame from the time domain to the frequency domain, or an end-to-end temporal convolutional model is used to directly extract the audio features of each audio frame. The embodiments of this disclosure do not specifically limit the method of audio feature extraction.

[0158] In some embodiments, when the standard facial data of a virtual object is stored in the form of point cloud data, the point cloud features of the point cloud data are extracted as the facial features of the standard facial data. For example, an end-to-end point cloud encoder is used to encode the standard facial data to obtain the final facial features. The embodiments of this disclosure do not specifically limit the method of facial feature extraction.

[0159] In step 202, the server generates lower half face movement features, upper half face movement features, and head movement features of the virtual object when it emits the input audio, based on the audio features and the facial features.

[0160] Among them, the lower half face motion feature represents the motion characteristics of the lower half face region when the virtual object emits the input audio.

[0161] Among them, the upper half face motion feature represents the motion characteristics of the upper half face region when the virtual object emits the input audio.

[0162] The head motion feature represents the motion characteristics of the head region when the virtual object emits the input audio.

[0163] It should be noted that, in this application embodiment, the lower half of the face refers to the mouth area and its surrounding area (such as the chin, lower half of the cheeks, etc.) of the virtual object, the upper half of the face refers to the remaining facial area of ​​the virtual object excluding the lower half, and the head area refers to the facial contour and hair area of ​​the virtual object. The mouth area of ​​the virtual object specifically includes the lips and its surrounding area (such as the texture of the muscles around the lips, wrinkles, dimples, etc.). Thus, the combination of the lower half and upper half of the virtual object's face can encompass the entire facial area of ​​the virtual object, while the facial area and head area of ​​the virtual object encompass and indicate the head posture and facial expressions of the virtual object. Therefore, the lower half, upper half, and head areas of the virtual object can completely cover and represent the overall head posture of the virtual object, as well as the facial expressions presented by the specific facial features.

[0164] In one example, a horizontal line passing through the philtrum is used as the dividing line between the upper and lower face regions. This distinction based on the philtrum allows the motion features of the lower face region to focus on extracting the motion characteristics of the mouth region. The most important aspect of the mouth region is reflecting the shape of the lips (in addition to the lips, the mouth region also includes the texture of the surrounding muscles, wrinkles, dimples, etc.). The motion features of the upper face region reflect the associated motion features of the upper face region in response to the expressions of the lower face region. The head motion features of the head region reflect the swaying and shaking of the head, allowing for the targeted extraction of different motion characteristics from each region.

[0165] In some embodiments, the server divides each audio frame into three regions—lower face region, upper face region, and head region—based on the audio features and facial features extracted in step 201, and extracts motion features for each region to obtain lower face motion features, upper face motion features, and head motion features. This takes into account the different ways different facial regions respond to the input audio, making the motion features of each region more sensitive and adaptable to the input audio. For example, the content of the input audio directly affects the motion of the lower face region (i.e., different pronunciations directly affect the shape of the lips in the lower face), and the rhythm of the input audio directly affects the motion of the head region (i.e., different rhythms directly affect head swaying). By extracting motion features for three different regions separately, it is convenient to accurately and meticulously synthesize predictive facial data with higher adaptability.

[0166] In step 203, the server generates a facial offset amount when the virtual object emits the input audio based on the lower half face motion features, the upper half face motion features, and the head motion features. The facial offset amount indicates the facial and head offset of the virtual object relative to the standard facial data when it emits the input audio.

[0167] In some embodiments, the server predicts the facial offset relative to standard facial data when the virtual object emits the sound of each audio frame of the input audio, based on the lower half face motion features, upper half face motion features, and head motion features generated in step 202. This allows the facial offset to integrate the motion features of the lower half face region, upper half face region, and head region. Since different regions have different motion patterns, such as the lip motion pattern in the lower half face region being significantly different from the head motion pattern in the head region, the facial offset obtained by finely adjusting and correcting the motion features of different regions can accurately reflect the fine motion patterns of each region.

[0168] In step 204, the server generates predicted facial data for the virtual object when it emits the input audio, based on the facial offset and the standard facial data.

[0169] In some embodiments, using standard facial data and facial offsets generated frame by frame in step 203, predicted facial data matching each audio frame can be reconstructed frame by frame, thereby facilitating the generation of 3D facial animation of virtual objects based on multi-frame predicted facial data. This ensures that the lip shape of the 3D facial animation is highly consistent with the input audio, and also ensures that the expressions of the 3D facial animation are coherent and natural.

[0170] The method provided in this disclosure divides the input audio into three regions—lower face, upper face, and head—based on the audio features of the input audio and the facial features of the standard facial data of the virtual object. Motion features are extracted from each region to obtain lower face motion features, upper face motion features, and head motion features. This takes into account the different responses of different facial regions to the input audio, making the motion features of each region more sensitive and adaptable to the input audio. Furthermore, the motion features of the lower face, upper face, and head regions are integrated and finely adjusted to generate a facial offset, indicating the facial and head offsets of the virtual object relative to the standard facial data when emitting the input audio. Finally, predicted facial data is generated based on the standard facial data and the facial offset, ensuring that the predicted facial data, especially the lower face (lips), upper face, and head, have a high degree of adaptation to the input audio. This guarantees that the facial expressions of the virtual object in the facial animation generated based on multi-frame predicted facial data are coherent and natural.

[0171] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0172] Figure 3 This is an interactive flowchart illustrating a facial data generation method according to an embodiment of the present disclosure, such as... Figure 3 As shown, the facial data generation method is executed by a computer device. Taking the computer device as a server as an example, this embodiment includes the following steps.

[0173] In step 301, the server inputs the standard facial data of the virtual object into the point cloud encoder, encodes the standard facial data through the point cloud encoder, and outputs the facial features of the standard facial data.

[0174] The point cloud encoder is used to extract facial features from the facial data of virtual objects.

[0175] In some embodiments, after the server obtains the standard facial data of a virtual object, if the standard facial data is in the form of point cloud data, the point cloud data is input into a point cloud encoder, and the point cloud encoder encodes the input point cloud data (i.e., the standard facial data) to obtain the facial features of the standard facial data. These facial features can identify which virtual object the standard facial data belongs to, that is, they can identify the identity of the virtual object, aiming to extract important identity information of the virtual object from the standard facial data. The point cloud encoder can be an encoder used to encode point cloud data; for example, the point cloud encoder may be implemented as an MLP (Multilayer Perceptron) or other possible architectures. This disclosure does not specifically limit this embodiment.

[0176] In step 301 above, a possible implementation method for the server to extract facial features from standard facial data of virtual objects is provided, so that the facial features obtained by the point cloud encoder can contain the identity information of the virtual object. Optionally, the server can also extract the HOG (Histogram of Oriented Gradient) features of multiple facial feature points in the standard facial data as facial features, or extract the HOG features of multiple facial feature points in different scale spaces as facial features, so that the facial features can contain the image features of the facial features of the virtual object. The embodiments of this disclosure do not specifically limit the method of facial feature extraction.

[0177] In step 302, the server inputs the input audio into a temporal convolution model, and performs temporal convolution on at least one audio frame in the input audio to obtain the audio features of the at least one audio frame.

[0178] The temporal convolution model is used to extract audio features from each audio frame in the input audio.

[0179] In some embodiments, after the server obtains the input audio, it preprocesses the input audio, such as performing VAD, pre-emphasis, framing, windowing, etc., so that the input audio is divided into at least one audio frame. Then, the at least one audio frame is input into a temporal convolution model, and the at least one audio frame is temporally convolved by the temporal convolution model to obtain the audio features of each of the at least one audio frame.

[0180] For example, the aforementioned temporal convolutional model is provided as a TCN (Temporal Convolutional Network) that supports processing time series data. The TCN involves multiple hidden layers, each employing causal convolution with a set dilation rate (i.e., a combination of causal and dilated convolution). Causal convolution means that the audio features of each audio frame depend only on that audio frame and the historical audio frames preceding it, without incorporating future audio frames after that frame. By setting a dilation rate for the causal convolution, it ensures that the information extraction from the previous layer is cascading, and the dilation rates between hidden layers can differ. For example, the dilation rate of each layer can be twice that of the previous layer, reflecting the characteristics of dilated convolution. This increases the receptive field, allowing the feature maps output by each hidden layer to contain a wider range of information. It also ensures that the audio features extracted by the TCN for each audio frame take into account information from a broader range of historical audio frames, thereby improving the expressive power and accuracy of the audio features for each audio frame. After inputting at least one audio frame into the TCN, for each audio frame, causal convolution with a set dilation rate is performed layer by layer through multiple hidden layers based on the audio frame and the historical audio frames preceding it. The audio features of the audio frame are output by the last hidden layer. The audio features of each audio frame in the input audio can be obtained by analogy.

[0181] In step 302 above, a possible implementation of the server extracting audio features of the input audio is provided. The expressive power and accuracy of each audio frame can be guaranteed by TCN. Optionally, the server can also extract the frequency features of each audio frame as audio features after converting each audio frame from the time domain to the frequency domain, or extract the MFCC (Mel Frequency Cepstrum Coefficient) features of each audio frame as audio features after converting each audio frame from the time domain to the frequency domain. The embodiments of this disclosure do not specifically limit the method of audio feature extraction.

[0182] In steps 301-302 above, the server provides a possible implementation for extracting audio features of the input audio and facial features of standard facial data of virtual objects on a per-frame basis. Optionally, instead of extracting audio features from each audio frame of the input audio, audio features can be extracted from key audio frames of the input audio. The setting method of key audio frames is defined by a technician, or multiple audio frames can be sampled at equal or non-equal intervals to extract audio features. This can reduce the computational load of the audio feature extraction process. This disclosure does not specifically limit this aspect.

[0183] In some embodiments, for each audio frame in the input audio, the server can generate a frame of predicted facial data associated with the audio frame based on the audio features and facial features of the audio frame, using the method provided in steps 303-312 below. The generation method of a frame of predicted facial data associated with any audio frame in the input audio will be described below. The association of the audio frame with the predicted facial data means that the lip shape of the virtual object in the predicted facial data matches the mouth shape that should be present when the audio frame is emitted in the input audio frame. That is, when the virtual object simulates speaking, the mouth shape that should be present in the audio frame matches the lip shape in the predicted facial data.

[0184] In step 303, for any audio frame in the input audio, the server fuses the audio features of the audio frame with the facial features to obtain an initial fused feature.

[0185] In some embodiments, for each audio frame in the input audio, the server fuses the audio features extracted from the audio frame in step 302 with the facial features extracted in step 301 to obtain an initial fused feature.

[0186] Optionally, the merging methods include, but are not limited to: concat, element-wise addition, element-wise multiplication, bilinear merging, etc. There are no specific limitations on the merging methods here.

[0187] In one example, the audio features of the audio frame and the facial features are concatenated to obtain the initial fused features.

[0188] In step 304, the server encodes the initial fusion features to obtain the lower face motion features.

[0189] Among them, the lower half face motion feature represents the motion characteristics of the lower half face region when the virtual object emits the audio frame in the input audio.

[0190] In some embodiments, the server may utilize a lower face encoder to encode the initial fused features, wherein the lower face encoder is used to extract motion features of the lower face region of the virtual object when it speaks. Optionally, the server inputs the initial fused features into the lower face encoder, encodes the initial fused features through the lower face encoder, and outputs the lower face motion features. The lower face encoder may be implemented as an MLP, or other architectures such as a fully connected network, a convolutional neural network, or a residual convolutional network; this disclosure does not specifically limit the architecture of the lower face encoder.

[0191] In one example, the lower face encoder is a lower face MLP containing at least one hidden layer, which is kept in series. The initial fused features from step 303 are input into at least one hidden layer of the lower face MLP. The initial fused features are weighted by the weight matrix in the first hidden layer to obtain a feature map. The feature map is then input into the second hidden layer, and so on, until the last hidden layer of the lower face MLP outputs the lower face motion features. This achieves the encoding of lower face motion features based on the initial fused features.

[0192] In the above process, the lower face motion features are obtained by directly encoding based on the initial fusion features. This is consistent with the driving principle that the input audio directly acts on the lower face region, especially the mouth region. This ensures that the initial fusion features after the low-level audio features are fused with the facial features can be directly used to guide the lower face motion features, thus ensuring the strong expressive power and high accuracy of the lower face motion features.

[0193] In step 305, the server extracts the upper face motion features based on the lower face motion features.

[0194] Among them, the upper half face motion feature represents the motion features of the upper half face region when the virtual object emits the audio frame in the input audio.

[0195] In some embodiments, the server may utilize a first encoder to encode the initial fusion features, wherein the first encoder is used to extract intermediate-level encoded features based on the initial fusion features at an initial level between audio and face. That is, when the initial fusion features are considered as features at the initial level, in order to further extract deeper features, the first encoder is used to encode the initial fusion features to obtain an intermediate encoded feature, which can be regarded as a feature at an intermediate level.

[0196] Optionally, the server inputs the initial fused feature into a first encoder, which encodes the initial fused feature and outputs intermediate encoded features. The first encoder can be implemented as an MLP, or other architectures such as a fully connected network, a convolutional neural network, or a residual convolutional network. This disclosure does not specifically limit the architecture of the first encoder.

[0197] In one example, the first encoder is a first MLP containing at least one hidden layer, which is kept in series. The initial fused features from step 303 are input into at least one hidden layer of the first MLP. The initial fused features are weighted by the weight matrix in the first hidden layer to obtain a feature map. The feature map is input into the second hidden layer, and so on, until the last hidden layer of the first MLP outputs the intermediate encoded features, thus realizing the encoding of intermediate encoded features based on the initial fused features.

[0198] It should be noted that the first MLP and the lower half MLP in the previous step can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This disclosure does not specifically limit them.

[0199] In some embodiments, if intermediate coding features and lower face motion features are extracted based on the initial fusion features, the server can fuse the intermediate coding features and the lower face motion features to obtain lower face fusion features.

[0200] Optionally, the fusion method includes, but is not limited to: concatenation, element-wise addition, element-wise multiplication, bilinear merging, etc., and no specific limitation is made on the fusion method here. In one example, the intermediate encoded feature and the lower half of the face motion feature are concatenated to obtain the lower half of the face fused feature.

[0201] In some embodiments, the server may utilize an upper-face encoder to encode the lower-face fusion features, wherein the upper-face encoder is used to extract the motion features of the upper face region of the virtual object when it speaks. Optionally, the server inputs the lower-face fusion features into the upper-face encoder, encodes the lower-face fusion features through the upper-face encoder, and outputs the upper-face motion features. The upper-face encoder may be implemented as an MLP, or other architectures such as a fully connected network, convolutional neural network, or residual convolutional network; this disclosure does not specifically limit the architecture of the upper-face encoder.

[0202] In one example, the upper face encoder is an upper face MLP containing at least one hidden layer, which is kept in series. The lower face fusion features are input into at least one hidden layer of the upper face MLP. The lower face fusion features are weighted by the weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the upper face MLP outputs the upper face motion features. This achieves the encoding of upper face motion features based on the lower face fusion features.

[0203] It should be noted that the upper half face MLP, the first MLP mentioned above, and the lower half face MLP in the previous step can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This disclosure does not specifically limit them.

[0204] In the above process, considering that the movement of the upper half of the virtual object's face is affected by both the movement of the lower half of the face (i.e., lower half face movement features) and the intermediate audio features (i.e., intermediate encoding features), the intermediate encoding features are obtained by abstracting the initial fused features during the generation of the upper half face movement features. The intermediate encoding features and the lower half face movement features are then fused and encoded to obtain the final upper half face movement features. This results in the upper half face movement features containing the guidance of the lower half face movement features and carrying the key information of the intermediate audio features, thus possessing strong expressive power and high accuracy.

[0205] The above description provides a possible implementation for extracting upper face motion features based on lower face motion features. This ensures that the upper face motion features contain both the key information of the lower face motion features and the key information of the intermediate encoded features. Optionally, the server can also use only an upper face MLP to encode the lower face motion features to obtain the upper face motion features. This simplifies the extraction process of upper face motion features and saves computational resources. This disclosure does not specifically limit the implementation of this method.

[0206] In step 306, the server extracts head motion features based on the upper half of the face motion features.

[0207] The head motion feature characterizes the motion features of the head region when the virtual object emits the audio frame in the input audio.

[0208] In some embodiments, the server may utilize a second encoder to encode intermediate coded features, wherein the second encoder is used to extract the highest-level coded features based on the intermediate-level coded features between audio and facial features. That is, when intermediate coded features are considered as features at an intermediate level, in order to further extract deeper features, the second encoder is used to encode the intermediate coded features to obtain a final-level coded feature, which can be considered as a feature at the deepest level.

[0209] Optionally, the server inputs the intermediate encoded features into a second encoder, which encodes the intermediate encoded features and outputs the final layer encoded features. The second encoder can be implemented as an MLP, or other architectures such as a fully connected network, a convolutional neural network, or a residual convolutional network. This disclosure does not specifically limit the architecture of the second encoder.

[0210] In one example, the second encoder is a second MLP containing at least one hidden layer, which is kept in series. The intermediate encoded features from step 305 are input into at least one hidden layer of the second MLP. The intermediate encoded features are weighted by the weight matrix in the first hidden layer to obtain a feature map. The feature map is input into the second hidden layer, and so on, until the last hidden layer of the second MLP outputs the final layer encoded features, thus realizing the encoding of the final layer encoded features based on the intermediate encoded features.

[0211] It should be noted that the second MLP here and the aforementioned lower half MLP, first MLP, and upper half MLP can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This disclosure does not specifically limit this.

[0212] In some embodiments, when the final layer coding features and the upper half face motion features are extracted separately, the server can fuse the final layer coding features and the upper half face motion features to obtain the upper half face fused features.

[0213] Optionally, the fusion method includes, but is not limited to: concatenation, element-wise addition, element-wise multiplication, bilinear merging, etc., and no specific limitation is made on the fusion method here. In one example, the last layer encoded feature and the upper half face motion feature are concatenated to obtain the upper half face fusion feature.

[0214] In some embodiments, the server may utilize a head encoder to encode the upper face fusion features, wherein the head encoder is used to extract motion features of the head region of the virtual object when it speaks. Optionally, the server inputs the upper face fusion features into the head encoder, encodes the upper face fusion features through the head encoder, and outputs the head motion features. The head encoder may be implemented as an MLP, or other architectures such as a fully connected network, a convolutional neural network, or a residual convolutional network; this disclosure does not specifically limit the architecture of the head encoder.

[0215] In one example, the head encoder is a head MLP containing at least one hidden layer, which is kept in series. The above-mentioned upper face fusion features are input into at least one hidden layer of the head MLP. The upper face fusion features are weighted by the weight matrix in the first hidden layer to obtain a feature map. The feature map is input into the second hidden layer, and so on, until the last hidden layer of the head MLP outputs the head motion features, thus realizing the encoding of head motion features based on the upper face fusion features.

[0216] It should be noted that the head MLP and the aforementioned lower face MLP, first MLP, second MLP, and upper face MLP can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This disclosure does not specifically limit them.

[0217] In the above process, considering that the movement of the virtual object's head is affected by both the movement of the upper half of the face (i.e., upper half face movement features) and the deepest audio features (i.e., the final layer encoding features), the intermediate encoding features are abstracted to obtain the final layer encoding features during the generation of head movement features. The final layer encoding features and the upper half face movement features are then fused and encoded to obtain the final head movement features. This makes the head movement features contain the guidance of the upper half face movement features and also carry the key information of the deepest audio features, thus having strong expressive power and high accuracy.

[0218] The above description provides a possible implementation method for extracting head motion features based on upper face motion features. This method ensures that the head motion features contain both the key information of the upper face motion features and the key information of the final layer encoded features. Optionally, the server can also use only one head MLP to encode the upper face motion features to obtain the head motion features. This simplifies the head motion feature extraction process and saves computational resources. This disclosure does not specifically limit the implementation of this method.

[0219] In steps 303-306 above, a possible implementation is provided to generate the lower half face motion features, upper half face motion features, and head motion features of the virtual object when it emits the input audio for each audio frame based on the audio features and facial features of each audio frame. This allows the lower half face motion features, upper half face motion features, and head motion features of each frame to be extracted frame by frame based on the audio features and facial features, which facilitates the prediction of the virtual object's facial data frame by frame and the synthesis of the virtual object's facial animation.

[0220] Furthermore, through three cascaded feature encoding stages, features from three different regions are extracted: lower face motion features, upper face motion features, and head motion features. This forms a hierarchical region encoding module. This hierarchical region encoding module can analyze and decompose facial features, obtaining the motion features of three key regions with different motion patterns. By extracting motion features from different regions, targeted analysis and feature encoding of motion patterns in different facial regions can be performed. At the same time, the working methods of audio features and facial features are abstracted, and the lower face motion features, upper face motion features, and head motion features are extracted hierarchically. This facilitates the simulation and analysis of the motion patterns of different regions under their respective motion patterns, and facilitates the generation of the final predicted facial data.

[0221] In step 307, the server generates global motion features based on the initial fusion features, the lower half face motion features, the upper half face motion features, and the head motion features.

[0222] In some embodiments, the server fuses the initial fusion feature generated in step 303, the lower half face motion feature generated in step 304, the upper half face motion feature generated in step 305, and the head motion feature generated in step 306 to obtain a global fusion feature.

[0223] Furthermore, a global encoder is used to encode the global fusion features. This global encoder extracts the global motion features of the lower half of the face, the upper half of the face, and the head region of the virtual object during vocalization. Optionally, the server inputs the global fusion features into the global encoder, which encodes the global fusion features and outputs the global motion features. The global encoder can be implemented as an MLP, or other architectures such as a fully connected network, a convolutional neural network, or a residual convolutional network. This embodiment does not specifically limit the architecture of the global encoder.

[0224] In one example, the global encoder is a global MLP containing at least one hidden layer, which is kept in series. The global fused features are input into at least one hidden layer of the global MLP, and the global fused features are weighted by the weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the global MLP outputs the global motion features, thus realizing the encoding of global motion features based on global fused features.

[0225] It should be noted that the global MLP here and the aforementioned lower half face MLP, first MLP, second MLP, upper half face MLP and head MLP can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This disclosure does not specifically limit them.

[0226] In the above process, by first splicing and then encoding the initial fusion features, lower half face motion features, upper half face motion features and head motion features, a global motion feature is obtained. This global motion feature contains the key information of the motion features of each region, thereby guiding the calculation of the offset of each region relative to the standard facial data.

[0227] Figure 4 This is a schematic diagram illustrating the generation principle of predictive facial data according to an embodiment of the present disclosure, such as... Figure 4 As shown, the various models involved in the entire generation process are collectively referred to as the Speech2MeshRegion model, which mainly includes an encoding module 410 based on hierarchical regions and a decoding module 420 based on local refinement.

[0228] In the input phase, the standard facial data (Template) of the virtual object is input into a point cloud encoder to extract facial features (ID Feature), and the input audio (Speech) is input into a temporal convolutional model to extract audio features (Speech Feature).

[0229] During the encoding stage, facial features and audio features are concatenated by encoding module 410 to obtain an initial fused feature f. low The initial fusion features f of the initial layer low Input the first encoder E1, output the intermediate encoded features f of the intermediate layer mid ; the intermediate encoded features f of the intermediate level mid Input the second encoder E2, output the highest level's final layer encoded features f high .

[0230] Based on this, for the lower half of the face region, the aforementioned initial fusion feature f is applied. low Input lower half face encoder E m Output the lower half of the face motion features f mouth For the upper half of the face, the motion features of the lower half of the face are... mouth and intermediate encoded features f mid The features are concatenated to obtain a lower half face fusion feature, which is then input into the upper half face encoder E. uf Output upper half face motion features f ufFor the head region, the upper half of the face motion features f uf and the last layer encoding features f high The features are concatenated to obtain an upper half face fusion feature, which is then input into the head encoder E. h Output head motion features f head .

[0231] Based on this, the initial fusion feature f low Lower half facial movement characteristics f mouth Upper facial movement characteristics f uf Head movement characteristics f head The features are concatenated to obtain a global fusion feature, which is then input into the global encoder E. g Output global motion features f global .

[0232] In one example, the lower face encoder E in encoding module 410 m Upper face encoder E uf and head encoder E h All three are MLPs with the same architecture and shared parameters, but the first encoder E1 and the second encoder E2 are independent MLPs that do not share parameters with the former, and each encoder and decoder in the decoding module 420 is also an independent MLP that does not share parameters with the former.

[0233] In the aforementioned encoding module 410, the face of the virtual object is divided into three distinct regions (lower face region, upper face region, and head region). Facial feature points in each region exhibit strong motion consistency within that region; that is, under audio-driven conditions, the facial feature points in each region share the same motion pattern. The identical motion pattern within each region means that different facial feature points within each region respond to audio with similar motion patterns. Thus, by dividing the virtual object's face into three regions with different motion patterns, and combining these three regions, a complete face can be reconstructed. This allows for region-specific motion analysis under the drive of input audio, enabling the targeted extraction of unique motion features from each region.

[0234] Specifically, the input audio is first applied to the lower half of the face region, and the initial fused features f low The input audio contains information such as phonemes, pitch, and emotion, which has the most direct impact on the movement of the lower face region, thus affecting the initial fusion feature f. low The encoded lower face motion features f mouth It will be directly guided by information such as phonemes, pitch, and emotion in the initial level, ensuring the lower half of the face's motor features f mouthThe accuracy of the data; furthermore, movement in the lower half of the face will influence movement in the upper half of the face. For example, when a person speaks excitedly, the lower half of the face, especially the mouth area, moves significantly, and the upper half of the face will also be affected, producing a significant movement. Therefore, the lower half of the face movement feature f mouth and based on the initial fusion feature f low The intermediate encoded features f obtained from encoding mid Together, they are introduced into the upper half of the face motion features f uf During the synthesis process, the motion features of the upper half of the face were preserved. uf The accuracy of the data; furthermore, the movement of the upper and lower face regions determines the overall facial movement of the virtual object, and facial movement, in turn, affects head movement. For example, head movement is usually closely related to facial movement in terms of amplitude, rhythm, and emotional expression. Therefore, the upper face movement features f uf and based on intermediate encoding features f mid The final layer encoded feature f is obtained through encoding. high Together they are introduced into the head motion feature f head During the synthesis process, the head motion features f were preserved. head The accuracy of.

[0235] Overall, driven by the input audio, the lower half of the face motion features f are generated first in the lower half of the face region. mouth Regenerate the upper half of the face motion features f in the upper half of the face region. uf Finally, head motion features f of the head region are generated. head Furthermore, the motion features of the regions generated first can be used to guide the motion features of the regions generated later. For example, the motion features of the lower half of the face generated first can be used to guide the motion features of the regions generated later. mouth upper face motion features f used for guidance in the generation of post-generative features uf Similarly, the first generated upper half face motion features f uf Head motion features f generated after guidance head .

[0236] Since audio features function differently at different stages and for different regions of the virtual object's face, hierarchical and regional abstraction allows for the assignment of more customized meanings to the motion features of each region at each level. In the initial level, the initial fusion feature f, resulting from the concatenation of audio and facial features, is... low It directly acts on the lower half of the face, thereby generating lower half facial motion features f. mouth In the intermediate layer, the initial fused feature f low Abstraction is performed to obtain intermediate encoded features f mid Encoding features f in the middle of the upper half of the face region midand lower face movement features f mouth Generate upper half facial motion features f under the guidance of uf At the highest level, the intermediate encoded features are abstracted to obtain the final encoded features f. high Encoding features f in the head region at the last layer high and upper facial movement features f uf Head motion features f are generated under the guidance of [the relevant authority]. head .

[0237] Because the virtual object's face is divided into three different regions, the initial fusion feature f low After undergoing multi-level coding, the generation process of motion features in different regions is guided accordingly. In this way, by observing the motion patterns of different facial regions, the generation relationship and generation order of motion features between different regions can be simulated by machine, thereby improving the expressive power and accuracy of the motion features of each region.

[0238] In step 308, the server generates a lower face offset based on the global motion feature and the lower face motion feature.

[0239] The lower face offset indicates the offset of the lower face feature points of the virtual object relative to standard facial data when it speaks.

[0240] In some embodiments, the server fuses the global motion features generated in step 307 and the lower face motion features generated in step 304 to obtain a lower face enhancement feature, which aims to enhance the motion features of the lower face region based on the global motion features.

[0241] Optionally, the merging methods include, but are not limited to: concatenation, element-wise addition, element-wise multiplication, bilinear merging, etc. No specific limitations are imposed on the merging method here. In one example, such as... Figure 4 As shown, the global motion features f generated in the encoding module 410 are... global and lower face movement features f mouth By concatenating the features, we obtain the enhanced lower half of the face using concat(f). mouth f global ).

[0242] In some embodiments, the server may utilize a lower face refinement encoder to encode lower face enhancement features, wherein the lower face refinement encoder is used to refine the lower face enhancement features of the virtual object when it speaks. Optionally, the server inputs the lower face enhancement features into the lower face refinement encoder, encodes the lower face enhancement features through the lower face refinement encoder, and outputs lower face refinement features. The lower face refinement encoder may be implemented as an MLP, or other architectures such as a fully connected network, convolutional neural network, or residual convolutional network. This disclosure does not specifically limit the architecture of the lower face refinement encoder.

[0243] In one example, the lower face refinement encoder is a lower face refinement MLP containing at least one hidden layer, which is cascaded. The lower face enhancement features are input into at least one hidden layer of the lower face refinement MLP, and the lower face enhancement features are weighted by the weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the lower face refinement MLP outputs the lower face refinement features. This achieves the encoding of lower face refinement features based on the lower face enhancement features.

[0244] For example, such as Figure 4 As shown, the lower half of the face enhancement features are concat(f) mouth f global The input is given to a lower face retouching MLP, and the output is the lower face retouching features. Assuming E is used m If we're referring to MLP (Multi-Layered Retouching) for the lower half of the face, then:

[0245] It should be noted that, as Figure 4 As shown, the lower half face refinement MLP in the decoding module 420 and the lower half face MLP in the encoding module 410 can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This embodiment does not specifically limit them.

[0246] In some embodiments, the server may utilize a lower face decoder to decode the lower face refinement features, wherein the lower face decoder is used to decode the lower face refinement features into a representation of the lower face offset of the virtual object relative to the standard facial data when it speaks. Optionally, the server inputs the lower face refinement features into the lower face decoder, decodes the lower face refinement features through the lower face decoder, and outputs the lower face offset. The lower face decoder may be implemented as an MLP, or other architectures such as a fully connected network, convolutional neural network, or residual convolutional network; this disclosure does not specifically limit the architecture of the lower face decoder.

[0247] In one example, the lower face decoder is a lower face decoding MLP containing at least one hidden layer, which is cascaded. The aforementioned lower face refinement features are input into at least one hidden layer of the lower face decoding MLP. The lower face refinement features are weighted by a weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the lower face decoding MLP outputs the lower face offset. This achieves decoding the lower face offset based on the lower face refinement features. For example, the lower face offset is a three-dimensional matrix, where the three-dimensional matrix represents the offset direction and offset distance of facial feature points in the x, y, and z directions, respectively.

[0248] like Figure 4 As shown, refine the features of the lower half of the face. The input is fed into the lower half face decoding MLP, and the output is the lower half face offset O. mouth Assuming we use D m If we represent MLP decoding of the lower half of the face, then we have:

[0249] In the above process, for the lower half of the face region, the motion features of the lower half of the face in the encoding stage are first concatenated with the global motion features and then encoded to obtain the refined features of the lower half of the face. Based on the refined features of the lower half of the face, the offset of the lower half of the face after refinement is obtained. Compared with the method of directly reconstructing the facial offset of the whole face, the lower half of the face offset generated in this step 308 has higher accuracy and stronger expressive power in the offset direction and offset distance of the predicted facial feature points in the lower half of the face region.

[0250] In step 309, the server generates an upper face offset based on the global motion feature and the upper face motion feature.

[0251] The upper face offset indicates the offset of the upper face feature points of the virtual object relative to standard facial data when it speaks.

[0252] In some embodiments, the server fuses the global motion features generated in step 307 and the upper face motion features generated in step 305 to obtain an upper face enhancement feature, which aims to enhance the motion features of the upper face region based on the global motion features.

[0253] Optionally, the merging methods include, but are not limited to: concatenation, element-wise addition, element-wise multiplication, bilinear merging, etc. No specific limitations are imposed on the merging method here. In one example, such as... Figure 4 As shown, the global motion features f generated in the encoding module 410 are... global and upper facial movement features f uf By concatenating the features, we obtain the enhanced upper face features concat(f) uf f global ).

[0254] In some embodiments, the server may utilize an upper-face refinement encoder to encode upper-face enhancement features, wherein the upper-face refinement encoder is used to refine the upper-face enhancement features of the virtual object when it speaks. Optionally, the server inputs the upper-face enhancement features into the upper-face refinement encoder, encodes the upper-face enhancement features through the upper-face refinement encoder, and outputs upper-face refinement features. The upper-face refinement encoder may be implemented as an MLP, or other architectures such as a fully connected network, convolutional neural network, or residual convolutional network. This disclosure does not specifically limit the architecture of the upper-face refinement encoder.

[0255] In one example, the upper face refinement encoder is an upper face refinement MLP containing at least one hidden layer, which is cascaded. The aforementioned upper face enhancement features are input into at least one hidden layer of the upper face refinement MLP. The upper face enhancement features are weighted by the weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the upper face refinement MLP outputs the upper face refinement features. This achieves the encoding of upper face refinement features based on the upper face enhancement features.

[0256] For example, such as Figure 4 As shown, the enhanced features of the upper half of the face are concat(f) uf f global The input is given to an upper half-face retouching MLP, and the output is the upper half-face retouching features. Assuming E is used uf If we're talking about MLP (Multi-Layered Retouching) for the upper half of the face, then:

[0257] It should be noted that, as Figure 4As shown, the upper half face refinement MLP in the decoding module 420 and the upper half face MLP in the encoding module 410 can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This embodiment does not specifically limit them.

[0258] In some embodiments, the server may utilize an upper face decoder to decode the upper face refinement features, wherein the upper face decoder is used to decode the upper face refinement features into an upper face offset representing the virtual object's upper face offset relative to the standard facial data when uttering a sound. Optionally, the server inputs the upper face refinement features into the upper face decoder, decodes the upper face refinement features through the upper face decoder, and outputs the upper face offset. The upper face decoder may be implemented as an MLP, or other architectures such as a fully connected network, convolutional neural network, or residual convolutional network; this disclosure does not specifically limit the architecture of the upper face decoder.

[0259] In one example, the upper face decoder is an upper face decoding MLP containing at least one hidden layer, which is cascaded. The aforementioned upper face refinement features are input into at least one hidden layer of the upper face decoding MLP. The upper face refinement features are weighted by a weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the upper face decoding MLP outputs the upper face offset. This achieves decoding the upper face offset based on the upper face refinement features. For example, the upper face offset is a three-dimensional matrix, where the three-dimensional matrix represents the offset direction and offset distance of facial feature points in the x, y, and z directions, respectively.

[0260] like Figure 4 As shown, the upper half of the face is refined with special features. The input is fed into the upper half face decoding MLP, and the output is the upper half face offset O. uf Assuming we use D uf If we represent decoding the upper half of the face using MLP, then we have:

[0261] In the above process, for the upper half of the face region, the upper half of the face motion features and the global motion features are first concatenated and then encoded to obtain the upper half of the face refined features. Based on the upper half of the face refined features, the upper half of the face offset after refinement is obtained. Compared with the method of directly reconstructing the facial offset of the whole face, the upper half of the face offset generated in this step 309 has higher accuracy and stronger expressive power in the offset direction and offset distance of the predicted facial feature points in the upper half of the face region.

[0262] In step 310, the server generates a head offset based on the global motion feature and the head motion feature.

[0263] The head offset indicates the offset of the head feature points in the head region of the virtual object relative to standard facial data when it speaks.

[0264] In some embodiments, the server fuses the global motion features generated in step 307 and the head motion features generated in step 306 to obtain a head enhancement feature, which aims to enhance the motion features of the head region based on the global motion features.

[0265] Optionally, the merging methods include, but are not limited to: concatenation, element-wise addition, element-wise multiplication, bilinear merging, etc. No specific limitations are imposed on the merging method here. In one example, such as... Figure 4 As shown, the global motion features f generated in the encoding module 410 are... global and head movement characteristics f head Concatenate the features to obtain the head enhancement features concat(f) head f global ).

[0266] In some embodiments, the server may utilize a head refinement encoder to encode head enhancement features, wherein the head refinement encoder is used to refine the head enhancement features of the virtual object when it speaks. Optionally, the server inputs the head enhancement features into the head refinement encoder, encodes the head enhancement features through the head refinement encoder, and outputs head refinement features. The head refinement encoder may be implemented as an MLP, or other architectures such as a fully connected network, a convolutional neural network, or a residual convolutional network; this disclosure does not specifically limit the architecture of the head refinement encoder.

[0267] In one example, the head refinement encoder is a head refinement MLP containing at least one hidden layer, which is cascaded. The head enhancement features are input into at least one hidden layer of the head refinement MLP, and the head enhancement features are weighted by the weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the head refinement MLP outputs the head refinement features. This achieves the encoding of head refinement features based on head enhancement features.

[0268] For example, such as Figure 4 As shown, the head enhancement features are concat(f) head f global The input is given to a head refinement MLP, and the output is the head refinement features. Assuming E is usedh If we mean head-level fine-tuning MLP, then:

[0269] It should be noted that, as Figure 4 As shown, the header refinement MLP in the decoding module 420 and the header MLP in the encoding module 410 can be MLP models with the same architecture but different parameters, or they can be MLP models with the same architecture and shared parameters. This embodiment does not specifically limit them.

[0270] In some embodiments, the server may utilize a head decoder to decode refined head features, wherein the head decoder decodes the refined head features into a representation of the head offset of the virtual object relative to the standard facial data when uttering a sound. Optionally, the server inputs the refined head features into the head decoder, decodes the refined head features through the head decoder, and outputs the head offset. The head decoder may be implemented as an MLP, or other architectures such as a fully connected network, a convolutional neural network, or a residual convolutional network; this disclosure does not specifically limit the architecture of the head decoder.

[0271] In one example, the head decoder is a head decoding MLP containing at least one hidden layer, which is cascaded. The aforementioned refined head features are input into at least one hidden layer of the head decoding MLP. The refined head features are weighted by a weight matrix in the first hidden layer to obtain a feature map. This feature map is then input into the second hidden layer, and so on, until the last hidden layer of the head decoding MLP outputs the head offset. This achieves the decoding of the head offset based on the refined head features. For example, the head offset is a three-dimensional matrix, representing the offset direction and offset distance of facial feature points in the x, y, and z directions, respectively.

[0272] like Figure 4 As shown, refine the features of the head. The input is fed into the header decoding MLP, and the output is the header offset O. head Assuming we use D h If we represent header decoding of an MLP, then we have:

[0273] In the above process, for the head region, the head motion features from the encoding stage are first concatenated with the global motion features and then encoded to obtain the head refinement features. Based on the head refinement features, the head offset after refinement of the head region is obtained by decoding. Compared with the method of directly reconstructing the facial offset of the whole face, the head offset generated in this step 310 has higher accuracy and stronger expressive power in the offset direction and offset distance of the predicted facial feature points in the head region.

[0274] In step 311, the server generates the facial offset of the virtual object when it emits the audio frame based on the lower half face offset, the upper half face offset, and the head offset.

[0275] The facial offset indicates the facial and head offset relative to the standard facial data when the virtual object emits the audio frame in the input audio.

[0276] In some embodiments, the server can perform a weighted summation of the lower face offset generated in step 308, the upper face offset generated in step 309, and the head offset generated in step 310 to obtain the final global face offset.

[0277] In some embodiments, the server weights the lower face offset based on the lower face mask image of the standard facial data to obtain a weighted lower face offset. Optionally, the server obtains the lower face mask image M of the standard facial data. m Lower half face mask image M m This is used to distinguish between the lower half of the face and non-lower half of the face in standard facial data for virtual objects. For example, facial feature points in the lower half of the face are set to 1, while other facial feature points are set to 0; this is not strictly limited here. Next, the server generates the lower half face mask image M. m Offset of the lower half of the face O mouth Perform matrix multiplication to obtain the weighted offset M of the lower half of the face. m O mouth .

[0278] In some embodiments, the server weights the upper face offset based on the upper face mask image of the standard facial data to obtain a weighted upper face offset. Optionally, the server obtains the upper face mask image M. uf Upper half face mask image M uf This is used to distinguish between the upper and lower face regions and non-upper face regions in standard facial data for virtual objects. For example, facial feature points in the upper face region are set to 1, while other facial feature points are set to 0; this is not strictly limited here. Next, the server generates the upper face mask image M... uf Offset O of the upper half of the face uf Perform matrix multiplication to obtain the weighted offset M of the upper half of the face. uf O uf .

[0279] In some embodiments, the server weights the head offset based on the head mask image of the standard facial data to obtain a weighted head offset. Optionally, the server obtains the head mask image M. h Header mask image M hThis is used to distinguish between head and non-head regions in standard facial data for virtual objects. For example, facial feature points in the head region are set to 1, while other facial feature points are set to 0; this is not strictly limited here. Next, the server sends the head mask image M... h With head offset O head Perform matrix multiplication to obtain the head weighted offset M. h O head .

[0280] In some embodiments, the server fuses the lower face weighted offset, the upper face weighted offset, and the head weighted offset to obtain the facial offset. For example, the lower face weighted offset M... m O mouth Upper half face weighted offset M uf O uf Head weighted offset M h O head Element-wise addition yields the final global facial offset O. global That is, O global =M n O mouth +M uf O uf +M h O head .

[0281] It should be noted that each facial feature point in the standard facial data will be uniquely assigned to one of the following regions: the lower half of the face, the upper half of the face, or the head region. That is, the lower half of the face mask M... m Upper half face mask image M uf Header mask image M h The union of the facial feature points of the three is exactly equal to all the facial feature points of the virtual object.

[0282] In the above process, by weighted summing of the lower half face offset, upper half face offset, and head offset, the global facial offset incorporates a refined offset optimized for each region. With the help of the mask map of each region, the refined offset of each region is accurately added to the regions of the facial offset at the same position, making the global facial offset more expressive and accurate.

[0283] In steps 308-311 above, a possible implementation is provided to generate the facial offset of the virtual object when it emits the input audio based on the lower half face motion features, the upper half face motion features and the head motion features. This allows the facial offset of the virtual object relative to the standard facial data to be reconstructed frame by frame based on the motion features of the three regions and the global motion features. This makes it convenient to use the facial offset to apply offset to the facial feature points in the standard facial data, thereby reconstructing the predicted facial data.

[0284] In the above process, based on the idea of ​​extracting motion features from different regions of the virtual object's face, the facial offset is also constructed by dividing the face into different regions to build local fine offsets. This makes the fine offset of each region more accurate and more compatible with the input audio. As a result, the global facial offset obtained by weighted summation of the fine offsets of each region is also more accurate and more compatible with the input audio.

[0285] In other embodiments, the fine-tuning offset of each region can be decoded without dividing the region into regions during the decoding stage. Instead, a global decoder can be used to decode the global motion features generated in step 307 and output the global facial offset. This can improve the decoding rate of the facial offset and save computational resources in the decoding stage.

[0286] In step 312, the server generates predicted facial data for the virtual object when it emits the audio frame, based on the facial offset and the standard facial data.

[0287] In some embodiments, the server can apply an offset to facial feature points in the standard facial data based on this facial offset to obtain the predicted facial data. For example, such as Figure 4 As shown, the facial offset O of a frame is... global The predicted facial data (Output) is obtained by adding it element-wise to the standard facial data (Template).

[0288] In the above process, a possible implementation is provided to generate predicted facial data when the virtual object emits the input audio based on the facial offset and the standard facial data. By extracting the audio features of each audio frame, the facial offset can be calculated frame by frame, and then the predicted facial data can be synthesized frame by frame. Finally, 3D facial animation is synthesized based on the predicted facial data of multiple frames, achieving frame-level audio-driven 3D facial animation synthesis.

[0289] In step 313, the server synthesizes the facial animation of the virtual object based on multi-frame predicted facial data.

[0290] In some embodiments, the server repeatedly executes steps 303-312 above multiple times for multiple audio frames in the input audio to obtain multi-frame predicted facial data. A predicted facial sequence can be synthesized from the multi-frame predicted facial data. This predicted facial sequence can synthesize the facial animation of a virtual object. Each predicted facial data in the predicted facial sequence has the same timestamp as an audio frame in the input audio, achieving automatic alignment of sound and image. In the synthesized facial animation, because the motion features of each region of the face are accurately predicted during the encoding stage, the fine-tuning offset of each region is more detailed and reasonable. This makes the final synthesized predicted facial data more refined and coordinated, and the facial animation composed of multi-frame predicted facial data has more coherent and natural expression changes.

[0291] The method provided in this disclosure divides the input audio into three regions—lower face, upper face, and head—based on the audio features of the input audio and the facial features of the standard facial data of the virtual object. Motion features are extracted from each region to obtain lower face motion features, upper face motion features, and head motion features. This takes into account the different responses of different facial regions to the input audio, making the motion features of each region more sensitive and adaptable to the input audio. Furthermore, the motion features of the lower face, upper face, and head regions are integrated and finely adjusted to generate a facial offset, indicating the facial and head offsets of the virtual object relative to the standard facial data when emitting the input audio. Finally, predicted facial data is generated based on the standard facial data and the facial offset, ensuring that the predicted facial data, especially the lower face (lips), upper face, and head, have a high degree of adaptation to the input audio. This guarantees that the facial expressions of the virtual object in the facial animation generated based on multi-frame predicted facial data are coherent and natural.

[0292] Furthermore, by employing a more efficient multi-level feature encoding module during the encoding stage, the learned audio features at different levels can guide the motion features of different regions, fully leveraging the role of audio features. This allows for the use of a large amount of information from different levels of the input audio, such as content, pitch, frequency, and rhythm, to provide guidance in different regions. For example, in the initial level, low-level content information in the audio is fully utilized to drive the synthesis of motion features in the lower half of the face; in the intermediate level, mid-level rhythm and frequency information is fully utilized to drive the synthesis of motion features in the upper half of the face; and in the highest level, high-level semantic information is fully utilized to drive the synthesis of motion features in the head region. This ensures that the information from different levels of audio features is fully utilized in the motion features of different regions. Ultimately, driven by the input audio, high-quality 3D facial animation is generated regionally, ensuring optimal facial reconstruction and lip-sync animation effects, demonstrating better performance and generalization.

[0293] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0294] In an exemplary test scenario, based on Face reconstruction metric (e-5) The model architecture described above was tested on the VOCASET test set. A face reconstruction metric and a lip-sync metric were used as two performance metrics. The face reconstruction metric calculates the L1 distance between the point clouds of the generated predicted face data and the reference (Ground Truth, GT) face data, thus evaluating the overall quality of the generated face. The lip-sync metric calculates the L1 distance between the lips of the generated predicted face data and the GT face data, thus evaluating the accuracy of the prediction in the lower face region. The experimental group was the Speech2MeshRegion method of this publication, and the control group included the baseline method VOCA and the control method Faceformer (face reconstruction). The test results are shown in Table 1 below:

[0295] Table 1

[0296] Lip-sync metric (e-4) VOCA Faceformer 42.99 15.91 Speech2MeshRegion 42.93 13.7 Figure 5 40.58 15.07

[0297] Here, e-5 refers to the numerical unit of the face reconstruction index being 10. -5 e-4 refers to the lip-sound synchronization index, with a unit of 10. -4 .

[0298] As can be seen from Table 1, the Speech2MeshRegion scheme disclosed in this paper achieves the best generation effect in face reconstruction metrics, and also achieves good results in lip-sync metrics, maintaining its superiority over the baseline VOCA method.

[0299] Figure 5 This is a logical structure block diagram of a facial data generation device according to an embodiment of the present disclosure. (Refer to...) Figure 5 The device includes: an extraction unit 501, a motion feature generation unit 502, an offset generation unit 503, and a facial data generation unit 504. These will be described below:

[0300] Extraction unit 501 is configured to extract audio features from the input audio and facial features from standard facial data of the virtual object;

[0301] The motion feature generation unit 502 is configured to generate, based on the audio feature and the facial feature, lower half face motion features, upper half face motion features and head motion features when the virtual object emits the input audio.

[0302] The offset generation unit 503 is configured to generate a facial offset when the virtual object emits the input audio based on the lower half face motion features, the upper half face motion features and the head motion features. The facial offset indicates the facial offset and head offset of the virtual object relative to the standard facial data when it emits the input audio.

[0303] The facial data generation unit 504 is configured to generate predicted facial data when the virtual object emits the input audio, based on the facial offset and the standard facial data.

[0304] The apparatus provided in this disclosure divides the input audio into three regions—lower face, upper face, and head—based on the audio features of the input audio and the facial features of the standard facial data of the virtual object. It then extracts the motion features of each region, resulting in lower face motion features, upper face motion features, and head motion features. This approach takes into account the different responses of different facial regions to the input audio, ensuring that the motion features of each region have better sensitivity and adaptability to the input audio. Furthermore, it integrates the motion features of the lower face, upper face, and head regions, and after fine-tuning, generates a facial offset to indicate the facial and head offsets of the virtual object relative to the standard facial data when emitting the input audio. Finally, it generates predicted facial data based on the standard facial data and the facial offset, ensuring that the predicted facial data, especially the lower face (lips), upper face, and head, have a high degree of adaptability to the input audio. This guarantees that the facial expressions of the virtual object in the facial animation generated based on multi-frame predicted facial data are coherent and natural.

[0305] In some embodiments, based on Figure 5 The device comprises a motion feature generation unit 502 including:

[0306] The fusion subunit is configured to perform the fusion of the audio feature and the facial feature to obtain an initial fused feature;

[0307] The encoding subunit is configured to perform encoding based on the initial fusion feature to obtain the lower face motion feature, which characterizes the motion features of the lower face region when the virtual object emits the input audio.

[0308] In some embodiments, the coding subunit is configured to perform:

[0309] The initial fusion feature is input into the lower face encoder, which encodes the initial fusion feature and outputs the lower face motion feature. The lower face encoder is used to extract the motion features of the lower face region of the virtual object when it speaks.

[0310] In some embodiments, based on Figure 5 The device comprises a motion feature generation unit 502 including:

[0311] The first extraction subunit is configured to extract the upper face motion features based on the lower face motion features, wherein the upper face motion features characterize the motion features of the upper face region when the virtual object emits the input audio.

[0312] In some embodiments, the first extraction subunit is configured to perform:

[0313] The initial fusion feature is input into the first encoder, which encodes the initial fusion feature and outputs intermediate encoded features. The first encoder is used to extract intermediate encoded features based on the initial fusion feature at the initial level between audio and face. The initial fusion feature is obtained by fusing the audio feature and the facial feature.

[0314] The intermediate encoded feature and the lower half of the face motion feature are fused together to obtain the lower half of the face fused feature;

[0315] The lower half face fusion feature is input into the upper half face encoder, which encodes the lower half face fusion feature and outputs the upper half face motion feature. The upper half face encoder is used to extract the motion features of the upper half face region of the virtual object when it speaks.

[0316] In some embodiments, based on Figure 5 The device comprises a motion feature generation unit 502 including:

[0317] The second extraction subunit is configured to extract the head motion features based on the upper half of the face motion features, the head motion features representing the motion features of the head region when the virtual object emits the input audio.

[0318] In some embodiments, the second extraction subunit is configured to perform:

[0319] The intermediate coding features are input into the second encoder, which encodes the intermediate coding features and outputs the final layer coding features. The second encoder is used to extract the highest level coding features based on the intermediate level coding features between the audio and the face. The intermediate coding features are encoded based on the initial fusion features, which are obtained by fusing the audio features and the facial features.

[0320] The final layer encoding features and the upper half face motion features are fused to obtain the upper half face fusion features;

[0321] The upper half face fusion feature is input into the head encoder, which encodes the upper half face fusion feature and outputs the head motion feature. The head encoder is used to extract the motion features of the head region of the virtual object when it speaks.

[0322] In some embodiments, based on Figure 5 The device comprises an offset generation unit 503 including:

[0323] The first generation subunit is configured to generate global motion features based on the initial fusion features, the lower half face motion features, the upper half face motion features, and the head motion features, wherein the initial fusion features are obtained by fusing the audio features and the facial features;

[0324] The second generation subunit is configured to generate a lower face offset based on the global motion feature and the lower face motion feature.

[0325] The third generation subunit is configured to generate an upper face offset based on the global motion feature and the upper face motion feature.

[0326] The fourth generation subunit is configured to generate a head offset based on the global motion feature and the head motion feature;

[0327] The fifth generation subunit is configured to generate the facial offset based on the lower half face offset, the upper half face offset, and the head offset.

[0328] In some embodiments, the first generation subunit is configured to perform:

[0329] The initial fusion feature, the lower half face motion feature, the upper half face motion feature, and the head motion feature are fused together to obtain the global fusion feature;

[0330] The global fusion feature is input into the global encoder, which encodes the global fusion feature and outputs the global motion feature. The global encoder is used to extract the global motion features of the lower half of the face, the upper half of the face, and the head region of the virtual object when it speaks.

[0331] In some embodiments, the second generation subunit is configured to perform:

[0332] The global motion feature and the lower half of the face motion feature are fused together to obtain the lower half of the face enhancement feature;

[0333] The lower face enhancement feature is input into the lower face refinement encoder, which encodes the lower face enhancement feature and outputs the lower face refinement feature. The lower face refinement encoder is used to refine the lower face enhancement feature of the virtual object when it speaks.

[0334] The lower face refinement feature is input into the lower face decoder, which decodes the lower face refinement feature and outputs the lower face offset. The lower face decoder is used to decode the lower face refinement feature into a representation of the lower face offset of the virtual object relative to the standard facial data when it is speaking.

[0335] In some embodiments, the third generation subunit is configured to perform:

[0336] The global motion feature and the upper half face motion feature are fused together to obtain the upper half face enhancement feature;

[0337] The upper half face enhancement feature is input into the upper half face refinement encoder, which encodes the upper half face enhancement feature and outputs the upper half face refinement feature. The upper half face refinement encoder is used to refine the upper half face enhancement feature of the virtual object when it speaks.

[0338] The upper half face refinement feature is input into the upper half face decoder. The upper half face decoder decodes the upper half face refinement feature and outputs the upper half face offset. The upper half face decoder is used to decode the upper half face refinement feature into a representation of the upper half face offset of the virtual object relative to the standard facial data when it is speaking.

[0339] In some embodiments, the fourth generation subunit is configured to perform:

[0340] The global motion feature and the head motion feature are fused together to obtain the head enhancement feature;

[0341] The head enhancement feature is input into the head refinement encoder, which encodes the head enhancement feature and outputs the head refinement feature. The head refinement encoder is used to refine the head enhancement feature of the virtual object when it is speaking.

[0342] The head refinement feature is input into the head decoder, which decodes the head refinement feature and outputs the head offset. The head decoder is used to decode the head refinement feature into a head offset that represents the virtual object's head offset relative to the standard facial data when it is speaking.

[0343] In some embodiments, the fifth generation subunit is configured to perform:

[0344] Based on the lower half face mask image of the standard facial data, the lower half face offset is weighted to obtain the weighted lower half face offset;

[0345] Based on the upper half face mask image of the standard facial data, the upper half face offset is weighted to obtain the upper half face weighted offset.

[0346] Based on the head mask image of the standard facial data, the head offset is weighted to obtain the head weighted offset.

[0347] The weighted offset of the lower half of the face, the weighted offset of the upper half of the face, and the weighted offset of the head are combined to obtain the facial offset.

[0348] In some embodiments, the facial data generation unit 504 is configured to perform:

[0349] Based on this facial offset, the facial feature points in the standard facial data are offset to obtain the predicted facial data.

[0350] In some embodiments, the extraction unit 501 is configured to perform:

[0351] The standard facial data is input into a point cloud encoder, which encodes the standard facial data and outputs the facial features. The point cloud encoder is used to extract facial features from the facial data of virtual objects.

[0352] In some embodiments, the extraction unit 501 is further configured to perform:

[0353] The input audio is fed into a temporal convolution model, which performs temporal convolution on at least one audio frame of the input audio to obtain the audio features of that at least one audio frame.

[0354] In some embodiments, based on Figure 6 The device comprises an animation synthesis unit configured to perform:

[0355] Based on the audio features of each audio frame in the input audio and the facial features, a frame of predicted facial data associated with the audio frame is generated;

[0356] The facial animation of the virtual object is synthesized based on multi-frame predicted facial data.

[0357] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0358] Regarding the apparatus in the above embodiments, the specific manner in which each unit performs its operation has been described in detail in the embodiments relating to the facial data generation method, and will not be elaborated upon here.

[0359] ​ This is a schematic diagram of a computer device 600 according to an embodiment of the present disclosure. The computer device 600 can vary significantly due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 601 and one or more memories 602. The memory 602 stores at least one instruction, which is loaded and executed by the processor 601 to implement the facial data generation method provided in the various embodiments described above. Of course, the computer device 600 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device 600 may also include other components for implementing device functions, which will not be elaborated upon here.

[0360] In an exemplary embodiment, a computer-readable storage medium including at least one instruction is also provided, such as a memory including at least one instruction, which can be executed by a processor in a computer device to complete the facial data generation method in the above embodiments. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (CompactDisc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.

[0361] In an exemplary embodiment, a computer program product is also provided, including one or more instructions that can be executed by a processor of a computer device to perform the facial data generation method provided in the above embodiments.

[0362] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0363] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for generating facial data, characterized in that, include: Extract audio features from the input audio and facial features from the standard facial data of the virtual object; The audio features and the facial features are fused to obtain the initial fused features; Encoding is performed based on the initial fusion features to obtain the lower half face movement features of the virtual object when it emits the input audio; The initial fusion features are input into a first encoder, which encodes the initial fusion features and outputs intermediate encoded features. The first encoder is used to extract intermediate-level encoded features based on the initial fusion features at the initial level between audio and face. The intermediate encoded features are fused with the lower half-face motion features to obtain lower half-face fusion features. The lower half-face fusion features are input into an upper half-face encoder, which encodes the lower half-face fusion features and outputs upper half-face motion features. The upper half-face encoder is used to extract the motion features of the upper half-face region of the virtual object when it speaks. The intermediate encoded features are input into a second encoder, which encodes the intermediate encoded features and outputs the final layer encoded features. The second encoder is used to extract the highest level encoded features based on the intermediate layer encoded features between the audio and the face. The final layer encoded features and the upper half face motion features are fused to obtain the upper half face fusion features. The upper half face fusion features are input into a head encoder, which encodes the upper half face fusion features and outputs the head motion features. The head encoder is used to extract the motion features of the head region of the virtual object when it speaks. Based on the lower half face motion features, the upper half face motion features, and the head motion features, a facial offset is generated when the virtual object emits the input audio. The facial offset indicates the facial and head offset of the virtual object relative to the standard facial data when it emits the input audio. Based on the facial offset and the standard facial data, predictive facial data is generated when the virtual object emits the input audio.

2. The facial data generation method according to claim 1, characterized in that, The encoding based on the initial fusion features to obtain the lower face motion features includes: The initial fusion features are input into the lower face encoder, which encodes the initial fusion features and outputs the lower face motion features. The lower face encoder is used to extract the motion features of the lower face region of the virtual object when it speaks.

3. The facial data generation method according to any one of claims 1 to 2, characterized in that, The step of generating the facial offset when the virtual object emits the input audio based on the lower half face motion features, the upper half face motion features, and the head motion features includes: Global motion features are generated based on the initial fusion features, the lower half face motion features, the upper half face motion features, and the head motion features. The initial fusion features are obtained by fusing the audio features and the facial features. Based on the global motion features and the lower face motion features, a lower face offset is generated; Based on the global motion features and the upper face motion features, an upper face offset is generated; Based on the global motion features and the head motion features, a head offset is generated; The facial offset is generated based on the lower half face offset, the upper half face offset, and the head offset.

4. The facial data generation method according to claim 3, characterized in that, The generation of global motion features based on the initial fusion features, the lower half face motion features, the upper half face motion features, and the head motion features includes: The initial fusion feature, the lower half face motion feature, the upper half face motion feature, and the head motion feature are fused to obtain the global fusion feature; The global fusion feature is input into the global encoder, and the global fusion feature is encoded by the global encoder to output the global motion feature. The global encoder is used to extract the global motion features of the lower half of the face, the upper half of the face, and the head region of the virtual object when it speaks.

5. The facial data generation method according to claim 3, characterized in that, The process of generating the lower face offset based on the global motion features and the lower face motion features includes: The global motion features and the lower face motion features are fused to obtain the lower face enhancement features; The lower face enhancement features are input into the lower face refinement encoder, which encodes the lower face enhancement features and outputs the lower face refinement features. The lower face refinement encoder is used to refine the lower face enhancement features of the virtual object when it is speaking. The lower face refinement features are input into the lower face decoder, which decodes the lower face refinement features and outputs the lower face offset. The lower face decoder is used to decode the lower face refinement features into a representation of the lower face offset of the virtual object relative to the standard facial data when it is speaking.

6. The facial data generation method according to claim 3, characterized in that, The process of generating the upper face offset based on the global motion features and the upper face motion features includes: The global motion features and the upper half face motion features are fused to obtain the upper half face enhancement features; The upper half face enhancement features are input into the upper half face refinement encoder, the upper half face enhancement features are encoded by the upper half face refinement encoder, and the upper half face refinement features are output. The upper half face refinement encoder is used to refine the upper half face enhancement features of the virtual object when it speaks. The upper half face refinement features are input into the upper half face decoder, and the upper half face refinement features are decoded by the upper half face decoder to output the upper half face offset. The upper half face decoder is used to decode the upper half face refinement features into an upper half face offset that represents the virtual object relative to the standard facial data when it is speaking.

7. The facial data generation method according to claim 3, characterized in that, The process of generating the head offset based on the global motion features and the head motion features includes: The global motion features and the head motion features are fused to obtain head enhancement features; The head enhancement features are input into the head refinement encoder, which encodes the head enhancement features and outputs the head refinement features. The head refinement encoder is used to refine the head enhancement features of the virtual object when it is speaking. The head refinement features are input into the head decoder, which decodes the head refinement features and outputs the head offset. The head decoder is used to decode the head refinement features into a head offset that represents the virtual object's head offset relative to the standard facial data when it is speaking.

8. The facial data generation method according to claim 3, characterized in that, The process of generating the facial offset based on the lower half face offset, the upper half face offset, and the head offset includes: Based on the lower half face mask image of the standard facial data, the lower half face offset is weighted to obtain the lower half face weighted offset. Based on the upper half face mask image of the standard facial data, the upper half face offset is weighted to obtain the upper half face weighted offset. Based on the head mask image of the standard facial data, the head offset is weighted to obtain the head weighted offset. The weighted offset of the lower half of the face, the weighted offset of the upper half of the face, and the weighted offset of the head are combined to obtain the facial offset.

9. The facial data generation method according to claim 1, characterized in that, The step of generating predicted facial data for the virtual object when it emits the input audio, based on the facial offset and the standard facial data, includes: Based on the facial offset, the facial feature points in the standard facial data are offset to obtain the predicted facial data.

10. The facial data generation method according to claim 1, characterized in that, The facial features extracted from the standard facial data of the virtual object include: The standard facial data is input into a point cloud encoder, which encodes the standard facial data and outputs the facial features. The point cloud encoder is used to extract facial features from the facial data of the virtual object.

11. The facial data generation method according to claim 1, characterized in that, The extracted audio features of the input audio include: The input audio is fed into a temporal convolution model, and the temporal convolution model performs temporal convolution on at least one audio frame in the input audio to obtain the audio features of the at least one audio frame.

12. The facial data generation method according to claim 11, characterized in that, The method further includes: Based on the audio features of each audio frame in the input audio and the facial features, a frame of predicted facial data associated with the audio frame is generated; The facial animation of the virtual object is synthesized based on multi-frame predicted facial data.

13. A facial data generation device, characterized in that, include: The extraction unit is configured to extract audio features from the input audio and facial features from standard facial data of the virtual object; A motion feature generation unit is configured to perform the fusion of the audio features and the facial features to obtain an initial fused feature; Encoding is performed based on the initial fusion features to obtain the lower half face movement features of the virtual object when it emits the input audio; The motion feature generation unit is further configured to: input the initial fusion feature into a first encoder; encode the initial fusion feature using the first encoder; output intermediate encoded features; the first encoder is used to extract intermediate-level encoded features based on the initial fusion feature at the initial level between audio and face; fuse the intermediate encoded features and the lower half face motion feature to obtain a lower half face fusion feature; input the lower half face fusion feature into an upper half face encoder; encode the lower half face fusion feature using the upper half face encoder; output upper half face motion feature; the upper half face encoder is used to extract motion features of the upper half face region of the virtual object when it speaks. The motion feature generation unit is further configured to input the intermediate encoded features into a second encoder, encode the intermediate encoded features through the second encoder, and output the final layer encoded features. The second encoder is used to extract the highest level encoded features based on the intermediate layer encoded features between the audio and the face. The final layer encoded features and the upper half face motion features are fused to obtain the upper half face fused features. The upper half face fused features are input into a head encoder, and the head encoder is used to encode the upper half face fused features to output the head motion features. The head encoder is used to extract the motion features of the head region of the virtual object when it speaks. The offset generation unit is configured to generate a facial offset when the virtual object emits the input audio based on the lower half face motion features, the upper half face motion features, and the head motion features. The facial offset indicates the facial and head offsets of the virtual object relative to the standard facial data when it emits the input audio. A facial data generation unit is configured to generate predicted facial data when the virtual object emits the input audio, based on the facial offset and the standard facial data.

14. A computer device, characterized in that, include: One or more processors; One or more memories for storing the one or more processor-executable instructions; The one or more processors are configured to execute the instructions to implement the facial data generation method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, When at least one instruction in the computer-readable storage medium is executed by one or more processors of a computer device, the computer device is enabled to perform the facial data generation method as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Virtual human image video generation method, system and device and storage medium

    CN113192161A

  • Animation fusion method and device

    CN114898019A