Face animation generation method and device and storage medium
By combining the whisper model and wav2cev model to extract audio features, combined with the facial image features, and injecting the diffusion model with the cross attention mechanism, the problem of synchronizing digital human facial movements and audio information is solved, achieving a more natural and accurate facial movement performance.
Patent Information
- Application Number
- CN202510166105.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to achieve precise synchronization of digital human facial movements and audio information.
By encoding the input audio based on the whisper model and wav2cev model, the common features, face features, head posture features and lips features of the face image are extracted, the characteristics and audio feature vectors are fused, and the diffusion model is injected using the cross attention mechanism to generate face animation.
It realizes the precise synchronization of audio information and facial movements, improving the naturalness and accuracy of digital human facial movements.
Smart Images

Figure CN120219579A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital human technology, and in particular, to a method, device, and storage medium for generating face animations. Background Art
[0002] Digital humans are virtual characters generated with the aid of computer and artificial intelligence technologies, and have extremely high application prospects in fields such as virtual anchors, education and training, customer service, and human-computer interaction.
[0003] Currently, by taking audio features and image features as conditions and integrating them into the processing of the latent representation of the diffusion model through a cross-attention mechanism, the influence weight of the audio features on the latent representation is determined by calculating attention scores, so as to integrate audio information into the latent representation and guide the diffusion model to generate corresponding facial actions according to the audio features, enabling the digital human to make corresponding facial actions according to the audio information.
[0004] However, when actually generating the facial actions of a digital human, it is difficult to capture the synchronous mapping of audio information and facial actions at the true physical level, so it is difficult to achieve the precise synchronization of audio information and facial actions.
[0005] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main purpose of this application is to provide a method, device, and storage medium for generating face animations, aiming to solve the technical problem of the asynchrony between audio information and facial actions in face animations.
[0007] To achieve the above object, this application proposes a method for generating face animations, and the method for generating face animations includes:
[0008] Encoding the input audio based on the whisper model and the wav2cev model to obtain an audio feature vector;
[0009] Extracting common features, facial features, head pose features, and lip features from the input face image, where the common features include face mesh features, tooth features, and eye features;
[0010] Fusing a first combined feature and a second combined feature to obtain a fused feature, where the first combined feature includes common features, facial features, and head pose features, and the second combined feature includes common features and lip features;
[0011] Taking the fused feature and the audio feature vector as conditional information and injecting them into the diffusion model through a cross-attention mechanism to enable the diffusion model to generate a face animation.
[0012] In one embodiment, the steps of encoding the input audio based on the Whisper model and the wav2cev model to obtain the audio feature vector include:
[0013] Based on the Whisper model and the wav2cev model, obtain the audio feature blocks of the input audio;
[0014] Loop through the audio feature blocks to obtain the audio feature sequence;
[0015] Based on the audio feature sequence, obtain the audio feature vector.
[0016] In one embodiment, the steps of obtaining the audio feature blocks of the input audio based on the Whisper model and the wav2cev model include:
[0017] Input the input audio into the Whisper model to obtain the first audio feature block of the input audio;
[0018] Input the first audio feature block into the wav2cev model to obtain the audio feature blocks.
[0019] In one embodiment, the steps of obtaining the audio feature vector based on the audio feature sequence include:
[0020] Fuse the audio feature sequence with the temporal embedding vector and the positional embedding vector to obtain the audio feature vector.
[0021] In one embodiment, the steps of injecting the fused features and the audio feature vector as conditional information into the diffusion model through the cross-attention mechanism to enable the diffusion model to generate a face animation include:
[0022] The diffusion model receives random noise, uses the fused features and the audio feature vector as conditional information, and based on the cross-attention mechanism, takes the latent representation of the face image as the query, and the fused features and the audio feature vector as the key and value, and determines the attention weights between the latent representation and the fused features and the audio feature vector respectively;
[0023] Adjust the latent representation according to the fused features, the audio feature vector, and the attention weights;
[0024] The decoder converts the latent representation into a face image frame;
[0025] Combine all the face image frames to obtain the face animation.
[0026] In one embodiment, the steps of combining all the face image frames to obtain the face animation include:
[0027] Obtain the prior knowledge of the face image frame of the current frame;
[0028] Construct the prior knowledge vector of the prior knowledge;
[0029] Adjust the generation parameters of the face image frame of the next frame based on the prior knowledge vector to obtain the face image frame of the next frame;
[0030] Connect all the face image frames to obtain a face animation.
[0031] In one embodiment, the steps for the decoder to transform the latent representation into a face image frame include:
[0032] The decoder maps the information in the latent representation to pixel values in the image space;
[0033] Generate a face image frame according to the pixel values in the image space.
[0034] In one embodiment, the steps for extracting the common features, facial features, head pose features, and lip features based on the input face image include:
[0035] Extract the overall framework of the input face image based on the face modeling method;
[0036] Determine the common features, facial features, head pose features, and lip features of the input face image in combination with the overall framework.
[0037] In addition, to achieve the above object, the present application also proposes a device for generating a face animation, the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the above method for generating a face animation.
[0038] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the steps of the above method for generating a face animation.
[0039] The present application encodes the input audio based on the whisper model and the wav2cev model to obtain an audio feature vector; extracts the common features, facial features, head pose features, and lip features based on the input face image, wherein the common features include face mesh features, tooth features, and eye features; fuses the first combined feature and the second combined feature to obtain a fused feature, wherein the first combined feature includes the common features, facial features, and head pose features, and the second combined feature includes the common features and lip features; uses the fused feature and the audio feature vector as conditional information and injects them into the diffusion model through the cross-attention mechanism to enable the diffusion model to generate a face animation. Since the multi-modal feature fusion condition method is adopted, it can accurately convert the semantic, emotional, and rhythm information in the audio into corresponding facial expressions and actions, and achieve the accurate synchronization of audio information and facial actions. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a schematic flowchart provided for the first embodiment of the method for generating the applicant's face animation;
[0043] Figure 2 It is a detailed schematic flowchart provided for step S10 in the first embodiment of the method for generating the applicant's face animation;
[0044] Figure 3 It is a detailed schematic flowchart provided for step S110 in the first embodiment of the method for generating the applicant's face animation;
[0045] Figure 4 It is a detailed schematic flowchart provided for step S20 in the first embodiment of the method for generating the applicant's face animation;
[0046] Figure 5 It is a detailed schematic flowchart provided for step S40 in the first embodiment of the method for generating the applicant's face animation;
[0047] Figure 6 It is a detailed schematic flowchart provided for step S430 in the first embodiment of the method for generating the applicant's face animation;
[0048] Figure 7 It is a detailed schematic flowchart provided for step S440 in the first embodiment of the method for generating the applicant's face animation;
[0049] Figure 8 It is a schematic flowchart of the method for generating the applicant's face animation;
[0050] Figure 9 It is a schematic diagram of the device structure of the hardware operating environment involved in the method for generating the face animation in the embodiments of the present application.
[0051] The realization of the purpose, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0052] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0053] To better understand the technical solution of this application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.
[0054] A digital human is a virtual character generated with the help of computer and artificial intelligence technologies, and has extremely high application prospects in fields such as virtual anchors, education and training, customer service, and human-computer interaction. Currently, by taking audio features and image features as conditions and integrating them into the processing of the latent representation of the diffusion model through a cross-attention mechanism, the influence weight of the audio features on the latent representation is determined by calculating attention scores, so as to integrate audio information into the latent representation, and guide the diffusion model to generate corresponding facial actions according to the audio features, enabling the digital human to make corresponding facial actions according to the audio information. However, when actually generating the facial actions of the digital human, it is difficult to capture the synchronous mapping of audio information and facial actions at the true physical level, so it is difficult to achieve the precise synchronization of audio information and facial actions.
[0055] In response to the above problems, the main solution of this application is: encoding the input audio based on the whisper model and the wav2cev model to obtain an audio feature vector; extracting common features, facial features, head pose features, and lip features from the input face image, where the common features include face mesh features, tooth features, and eye features; fusing the first combined feature and the second combined feature to obtain a fused feature, where the first combined feature includes common features, facial features, and head pose features, and the second combined feature includes common features and lip features; injecting the fused feature and the audio feature vector as conditional information into the diffusion model through a cross-attention mechanism, so that the diffusion model generates a face animation.
[0056] The solution of this application, by adopting the method of multi-modal feature fusion conditions, can accurately convert the semantic, emotional, and rhythm information in the audio into corresponding facial expressions and actions, and achieve the precise synchronization of audio information and facial actions.
[0057] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a face animation generation device, etc. that can implement the above functions. Hereinafter, taking the face animation generation device as an example, this embodiment and the following embodiments will be described.
[0058] Based on this, the embodiment of this application provides a method for generating a face animation, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the method for generating a face animation of this applicant.
[0059] In this embodiment, the method for generating a face animation includes steps S10 to S40:
[0060] Step S10: Encode the input audio based on the whisper model and the wav2cev model to obtain an audio feature vector.
[0061] It should be noted that the audio feature vector contains rich information about the input audio from the semantic to the acoustic level. For example, information such as semantics, prosody, pitch, timbre, and audio rhythm.
[0062] Further, referring to Figure 2 , step S10 includes steps S11 to S13:
[0063] Step S11: Based on the whisper model and the wav2cev model, obtain the audio feature blocks of the input audio.
[0064] It should be noted that the audio feature block is a high-dimensional feature vector obtained by the input audio through the whisper model and the wav2cev model. It carries vectors of multi-dimensional acoustic attributes such as the pitch, timbre, and rhythm of the input audio, and is a "storage box" for the attributes of the input audio.
[0065] Further, referring to Figure 3 , step S11 includes steps S111 to S112:
[0066] Step S111: Input the input audio into the whisper model to obtain the first audio feature block of the input audio.
[0067] Step S112: Input the first audio feature block into the wav2cev model to obtain the audio feature block.
[0068] In this embodiment, the input audio is preprocessed by using Whisper (a speech recognition model of OpenAI), such as operations like sampling and framing, to convert the input audio into a format suitable for model processing. Then, the built-in multi-layer Transformer encoder is used to extract and encode the features of the input audio, obtaining the first audio feature block of the input audio. Next, Wav2vec2 (a speech recognition model of Facebook AI Research) is used to perform quantization operations on the original waveform of the first audio feature block, such as discretizing the continuous input audio, and then the built-in multi-layer Transformer architecture is used to process the quantized input audio, obtaining the audio feature block of the input audio. Since Wav2vec2 can capture the underlying audio contour features of the input audio and Whisper can extract the high-level semantic features of the input audio, combining the two speech recognition models can more accurately extract the speech emotions and intonation changes of the input audio, enabling the facial expressions of the digital human to accurately reflect the corresponding emotional state of the input audio in a timely manner, making the facial movements of the digital human look vivid and natural when speaking, and improving the robustness of the recognition results.
[0069] Step S12: Obtain an audio feature sequence by looping the audio feature blocks.
[0070] In this embodiment, due to the length limitation of the audio feature blocks, multiple audio feature blocks are copied and spliced together to form an audio feature sequence.
[0071] Step S13: Obtain an audio feature vector based on the audio feature sequence.
[0072] Further, step S13 includes step S131:
[0073] Step S131: Fuse the audio feature sequence with a time embedding vector and a position embedding vector to obtain an audio feature vector.
[0074] It should be noted that time embedding is a technique for encoding time information into a vector, which allows the model to know the position of audio features in the time series. Position embedding is used to determine the relative position of audio features in the entire audio clip. By fusing the audio feature sequence with the time embedding vector and the position embedding vector, the temporal information and position information of the audio can be combined with the feature information of the audio itself to form a conditional information containing richer information.
[0075] In addition, the fusion method can be concatenation fusion, weighted fusion, and attention mechanism-based fusion.
[0076] In this embodiment, the time embedding vector and the position embedding vector are both randomly generated. By concatenating the audio feature sequence, the time embedding vector, and the position embedding vector, an audio feature vector is obtained to form a rich conditional information and injected into the subsequent diffusion model.
[0077] Step S20: Extract common features, facial features, head pose features, and lip features from the input face image. Among them, the common features include face mesh features, tooth features, and eye features.
[0078] It should be noted that the face mesh features, tooth features, eye features, facial features, head pose features, and lip features are all relevant face features of the input face image. For example, the face mesh features include the shape features and texture features of the face.
[0079] Further, referring to Figure 4 , step S20 includes steps S21 to S22:
[0080] Step S21: Extract the overall framework of the input face image based on the face modeling method.
[0081] It should be noted that the face modeling method can be used to determine the relevant parameters for face modeling in the input face image, such as the fatness and thinness of the face, skin color, the position of the nose tip, the position of the eye corners, etc.
[0082] In a feasible implementation manner, a face modeling method based on statistical learning is used to determine the overall framework of the input face image, and a face modeling method based on key points is used to determine the local position of the input face image.
[0083] In this embodiment, 3DMM (3D Morphable Model) is used to parameterize face model parameters such as face shape and face texture. It can perform principal component analysis (PCA) based on a large amount of 3D face data. Each principal component represents a factor that significantly affects the face appearance. By adjusting the principal component coefficients, the 3D face is represented as an average face shape plus a linear combination of multiple principal components. Specifically, as shown in formula (1):
[0084]
[0085] Among them, S is the final 3D face shape, is the average face shape, ai is the shape coefficient, S i is the i-th shape principal component.
[0086] Similarly, for the face texture, its formula is as shown in formula (2):
[0087]
[0088] Among them, T is the face texture, is the average face texture, β j is the texture coefficient, T i is the j-th texture principal component.
[0089] Through 3DMM, the general outline of the face, the relative positions of facial features, etc. can be determined.
[0090] Step S22: Combine the overall framework to determine the common features, facial features, head pose features, and lip features of the input face image.
[0091] In this embodiment, combined with the overall framework provided by 3DMM in step S21, the 3D Key Point technology can better understand the overall structure of the face and avoid large deviations in key point positioning. 3D Key Point will determine the specific positions of the key points of the input face image according to the overall framework generated by 3DMM, such as the specific positioning of teeth, eyes, and lips on the input face image.
[0092] Combined with the overall framework provided by 3DMM in step S21 and the specific positions of the key points provided by 3D Key Point, the accurate key point coordinates extracted by the 3D Key Point technology can be fed back to 3DMM to optimize the parameters of 3DMM, making the modeling of the face shape and texture by 3DMM more accurate, improving the accuracy of face modeling and making the created digital human more realistic and natural. Finally, through the combined use of 3DMM and 3D Key Point, the face mesh features, tooth features, eye features, facial features, head pose features, and lip features related to the face modeling parameters in the input face image are determined. These features contain the parameter information related to face modeling in the form of vectors. Among them, the face mesh provides the topological structure information of the face in the entire input face image, teeth and eyes are important components of the facial expressions and appearances in the input face image, the head pose reflects the overall orientation and angle of the face in the input face image, the face contains the key positions of facial muscle movements and expression changes in the input face image, and the facial parameters do not include lip parameters.
[0093] Step S30: Fuse the first combined feature and the second combined feature to obtain a fused feature, where the first combined feature includes the common features, facial features, and head pose features, and the second combined feature includes the common features and lip features.
[0094] In this embodiment, the features in the input face image are combined into two parts for independent modeling and encoding. The face mesh features, tooth features, and eye features are determined as common features, and the head pose features and face features are combined with the common features to determine the first combined feature, so as to comprehensively describe the pose and expression features of the face except the lips. Secondly, the lip features are combined with the common features as the second combined feature. The lips are one of the parts with the most obvious movement changes during speech. Combining the lips with the common features can highlight the movement features of the lips during speech. The above-mentioned first combined feature and second combined feature are spliced together to obtain a fusion feature, which contains rich face information. For the first combined feature, it can reflect the overall pose of the face and the movement trend of the facial muscles. For the second combined feature, it mainly reflects the shape and movement features of the lips. Thus, the decoupling of the digital human's facial movements can be achieved, making the digital human's movement performance more delicate and natural.
[0095] Step S40: Use the fusion feature and the audio feature vector as conditional information, and inject them into the diffusion model through the cross-attention mechanism, so that the diffusion model generates a face animation.
[0096] Further, referring to Figure 5 , step S40 includes steps S41 to S44:
[0097] Step S41: The diffusion model receives random noise, uses the fusion feature and the audio feature vector as conditional information, and based on the cross-attention mechanism, takes the latent representation of the face image as the query, and the fusion feature and the audio feature vector as the key and value, and determines the attention weights between the latent representation and the fusion feature and the audio feature vector respectively.
[0098] In this embodiment, first, the diffusion model starts to receive random noise, which is the starting data basis for the subsequent generation process. Then, the pre-prepared fusion feature and the audio feature vector are introduced into the model as key conditional information. At this time, the cross-attention mechanism starts to operate. First, the latent representation obtained by encoding the face image is set as the query, and the fusion feature and the audio feature vector are set as the key and value respectively. Subsequently, the model performs calculations according to the operation logic between the query, the key, and the value. Specifically, it will calculate the dot product of the latent representation and the fusion feature and the audio feature vector one by one, and then perform normalization processing through the softmax function, thereby accurately determining the attention weights between the latent representation and the fusion feature and the audio feature vector respectively. These weights will be used to weighted-sum relevant information later to drive the model towards the accurate generation target.
[0099] Step S42: Adjust the latent representation according to the fusion feature, the audio feature vector, and the attention weights.
[0100] In this embodiment, after obtaining the fused feature, the audio feature vector, and the corresponding attention weights, the key step of adjusting the latent representation is initiated. First, according to the attention weights between the calculated latent representation and the fused feature, a weighted summation operation is performed on the fused feature, highlighting the key elements in the fused feature that are strongly correlated with the latent representation, and obtaining the fused feature information with weight adjustment. Similarly, according to the attention weights between the latent representation and the audio feature vector, a weighted summation is also performed on the audio feature vector to filter out important audio features. Subsequently, the results of these two weighted processes are added together, and the newly obtained feature is incorporated into the original latent representation. Through addition or other preset fusion strategies, the update and optimization of the latent representation are completed, making it contain more accurate and rich information, laying a solid foundation for subsequent tasks such as generating high-quality face images.
[0101] Step S43, the decoder converts the latent representation into a face image frame.
[0102] Further, referring to Figure 6 , step S43 includes steps S431 to S432:
[0103] Step S431, the decoder maps the information in the latent representation to pixel values in the image space.
[0104] In this embodiment, after the adjustment of the latent representation is completed, the decoder starts to function. The decoder is built with specific mapping rules and transformation functions and first obtains the optimized latent representation. Since the latent representation is in a highly abstract feature space after encoding, what the decoder needs to do is to break this abstraction and reverse-transform it. Based on the pre-trained network parameters, for each group of feature vectors in the latent representation, a series of linear transformations and non-linear activation functions are used to gradually disassemble and recombine these abstract features into the corresponding pixel values in the image space. From color information to spatial position layout, all are accurately restored in this process, and finally, the blurred latent representation is clearly presented as face image pixel information with a clear visual form and rich colors, completing the key leap from abstract features to intuitive images.
[0105] Step S432, generate a face image frame according to the pixel values in the image space.
[0106] In this embodiment, after obtaining the pixel values of the image space, the process of generating the face image frame is immediately launched. The system will first arrange these pixel values in order according to the established image resolution specifications to construct a two-dimensional matrix. The number of rows and columns of the matrix corresponds to the number of pixels in length and width of the image, and each element accurately carries the corresponding color information. Subsequently, the image generation library or the built-in image rendering module is used to convert this two-dimensional matrix into a visual image format, and the pixels are given a real color mode, such as the common RGB mode, so that the values of the three color channels of red, green, and blue can be combined to produce tens of millions of colors. After the color calibration, format conversion and other fine processing are completed, a clear, complete, and expected facial feature-fitting image frame is successfully generated, providing key visual materials for subsequent coherent facial animation, dynamic display and other application scenarios.
[0107] Step S44, combining all face image frames to obtain face animation.
[0108] In a feasible implementation, all face image frames are acquired and connected in series in time to obtain a face animation.
[0109] In this embodiment, after a series of face image frames are obtained, in order to generate a coherent face animation, the image frames must first be accurately arranged in chronological order. A special animation editing tool or programming framework is used to set a corresponding timestamp for each frame to clarify its playback position on the animation timeline and ensure a natural and smooth transition between frames. Subsequently, the connection between adjacent frames is optimized, and an interpolation algorithm is used to fill the visual gaps caused by slightly abrupt movements, sudden changes in light and shadow, etc., so that the screen switches without any sense of lag. After completing the sorting and connection optimization, the playback speed is uniformly adjusted according to the frame rate requirements of the target animation to ensure that the animation rhythm is appropriate. Finally, the necessary format encapsulation is added to the entire animation to make it conform to common video playback specifications. At this point, a coherent and vivid face animation is integrated and can be used in various multimedia displays and interactive scenes.
[0110] Further, see Figure 7 , step S44 includes steps S441 to S444:
[0111] Step S441, obtaining prior knowledge of the face image frame of the current frame.
[0112] In this embodiment, a trained feature extraction model, such as a convolutional neural network, is used to capture the key facial features in an image, including elements such as the outlines of facial features, expression textures, and light and shadow distributions. At the same time, from a pre-constructed face database, similar facial samples are matched, and the common features and annotation information attached to the corresponding samples are extracted, such as the regular deformation range of facial muscles under specific expressions and the facial tone rules caused by common lighting conditions. By integrating the features directly extracted from the current frame and the relevant knowledge matched from the database, the collection of prior knowledge for the current frame of the face image is completed.
[0113] Step S442: Construct a prior knowledge vector of the prior knowledge.
[0114] In this embodiment, when constructing the prior knowledge vector of the prior knowledge, various prior knowledge obtained in the previous step is first quantized. For the outline information of facial features, it is described by numerical values such as coordinates and curvatures; the expression texture is converted into texture feature values; the knowledge related to light is converted into quantization parameters of brightness and contrast. Then, these quantization data in different dimensions are arranged and concatenated in sequence to form a high-dimensional vector. On this basis, a dimensionality reduction algorithm, such as principal component analysis, is used to compress the vector dimension and eliminate redundant information, and finally a prior knowledge vector that can accurately contain the key prior knowledge and is convenient for subsequent calculations is obtained.
[0115] Step S443: Adjust the generation parameters of the next frame of the face image based on the prior knowledge vector to obtain the next frame of the face image.
[0116] In this embodiment, when adjusting the generation parameters of the next frame of the face image based on the prior knowledge vector, first, it is necessary to determine which generation parameters can be adjusted, such as the deformation parameters that control the movement amplitude of facial muscles and the rendering parameters that affect skin color and light and shadow effects. Subsequently, the prior knowledge vector is input into a preset adjustment model. This model calculates the amplitude and direction of the change of each generation parameter through built-in algorithms based on the information such as facial contours, expressions, and light trends contained in the vector. According to the calculated adjustment amount, the original generation parameters are modified one by one, and then the updated parameters are substituted into the face image generation process, so as to obtain the next frame of the face image that conforms to the prior knowledge.
[0117] Step S444: Connect all the face image frames to obtain a face animation.
[0118] In this embodiment, to obtain a face animation by connecting all face image frames, the first step is to arrange all the generated face image frames in chronological order and put them into a sequence container. Then, select a suitable animation production tool or programming framework. Using the video generation function of such tools, set the frame rate, that is, the number of images played per second, and let these frames be played in sequence at the established frame rate. At the same time, handle the smooth connection during the transition between frames to avoid screen flickering and freezing. Finally, package it into a common video format to obtain the face animation.
[0119] Exemplarily, to help understand the implementation process of the method for generating a face animation obtained by combining this embodiment with the above-mentioned Embodiment 1, please refer to Figure 8 , Figure 8 which provides a schematic flowchart of a brief process of a method for generating a face animation. Specifically:
[0120] Use whisper and wav2vec2 to extract the audio features of the input audio to obtain audio feature blocks, and combine the audio feature blocks with time embedding and position embedding as the conditional information of the diffusion model; for the facial key point extraction part, use 3DMM and 3D Key Point technologies to extract features related to facial movements from the input face, such as head pose, face, face mesh, teeth, eyes, and lips. Take the face mesh, teeth, and eyes as the common part, the head pose and face combined with the common part as the first part, and the lips combined with the common part as the second part. Encode these two parts respectively to obtain Condition 1 and Condition 2. Combine Condition 1 with Condition 2 as the conditional information of the diffusion model; the diffusion model receives random noise and outputs continuous video frames based on the above conditional information to realize the face animation.
[0121] It should be noted that Condition 1 is the first combined feature mentioned in Embodiment 1, Condition 2 is the second combined feature, and the input reference diagram is the input face image.
[0122] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method for generating a face animation in this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0123] This application provides a device for generating a face animation. The device for generating a face animation includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a face animation in the above-mentioned Embodiment 1.
[0124] Next, refer to Figure 9, which shows a schematic structural diagram of a generating device suitable for implementing the face animation in the embodiments of the present application. The generating device for face animation in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The shown generating device for face animation is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0125] As Figure 9 shown, the generating device for face animation may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the generating device for face animation are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the generating device for face animation to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a generating device for face animation having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0126] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0127] The face animation generation device provided in the present application adopts the face animation generation method in the above embodiments, and can solve the technical problem of the out-of-sync between audio information and facial movements in face animation. Compared with the prior art, the beneficial effects of the face animation generation device provided in the present application are the same as those of the face animation generation method provided in the above embodiments, and other technical features in the face animation generation device are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated herein.
[0128] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0129] The above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0130] The present application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the face animation generation method in the above embodiments.
[0131] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0132] The above computer-readable storage medium can be included in a face animation generation device; it can also exist separately without being assembled into the face animation generation device.
[0133] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by a face animation generation device, the face animation generation device is caused to: encode the input audio based on the whisper model and the wav2cev model to obtain an audio feature vector; extract common features, facial features, head pose features, and lip features from the input face image, where the common features include face mesh features, tooth features, and eye features; fuse a first combined feature and a second combined feature to obtain a fused feature, where the first combined feature includes common features, facial features, and head pose features, and the second combined feature includes common features and lip features; use the fused feature and the audio feature vector as conditional information and inject them into a diffusion model through a cross-attention mechanism to cause the diffusion model to generate a face animation.
[0134] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0136] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0137] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned method for generating facial animation, and can solve the technical problem of the asynchronous audio information and facial movements in facial animation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for generating facial animation provided in the above embodiments, and will not be elaborated here.
[0138] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for generating facial animation, characterized in that: The method for generating the facial animation comprises: Encode the input audio based on the whisper model and wav2cev model to obtain the audio feature vector; Extracting common features, facial features, head posture features and lip features based on an input face image, wherein the common features include face mesh features, tooth features and eye features; Fusing the first combined feature and the second combined feature to obtain a fused feature, wherein the first combined feature includes the common feature, the facial feature and the head posture feature, and the second combined feature includes the common feature and a lip feature; The fusion feature and the audio feature vector are used as conditional information and injected into a diffusion model through a cross-attention mechanism, so that the diffusion model generates facial animation.
2. The method for generating facial animation according to claim 1, characterized in that: The step of encoding the input audio based on the whisper model and the wav2cev model to obtain the audio feature vector comprises: Based on the whisper model and the wav2cev model, obtaining an audio feature block of the input audio; Circulating the audio feature block to obtain an audio feature sequence; The audio feature vector is obtained based on the audio feature sequence.
3. The method for generating facial animation according to claim 2, characterized in that: The step of obtaining the audio feature block of the input audio based on the whisper model and the wav2cev model comprises: Inputting the input audio into the whisper model to obtain a first audio feature block of the input audio; The first audio feature block is input into the wav2cev model to obtain the audio feature block.
4. The method for generating facial animation according to claim 2, characterized in that: The step of obtaining the audio feature vector based on the audio feature sequence comprises: The audio feature sequence is fused with a time embedding vector and a position embedding vector to obtain an audio feature vector.
5. The method for generating facial animation according to claim 1, characterized in that: The step of injecting the fusion feature and the audio feature vector as conditional information into the diffusion model through a cross attention mechanism so that the diffusion model generates a facial animation comprises: The diffusion model receives random noise, takes the fused feature and the audio feature vector as conditional information, takes the potential representation of the face image as a query, takes the fused feature and the audio feature vector as a key and a value, and determines the attention weights between the potential representation and the fused feature and the audio feature vector respectively based on the cross attention mechanism; adjusting the latent representation according to the fused feature, the audio feature vector, and the attention weight; The decoder converts the latent representation into a face image frame; The face animation is obtained by combining all the face image frames.
6. The method for generating facial animation according to claim 5, characterized in that: The step of combining all face image frames to obtain the face animation comprises: Acquire prior knowledge of the face image frame of the current frame; Constructing a priori knowledge vector of the prior knowledge; Adjusting generation parameters of a next face image frame based on the prior knowledge vector to obtain a next face image frame; All face image frames are connected to obtain the face animation.
7. The method for generating facial animation according to claim 5, characterized in that: The decoder converts the potential representation into a face image frame comprising: The decoder maps the information in the latent representation into pixel values in an image space; The face image frame is generated according to the pixel values of the image space.
8. The method for generating facial animation according to claim 1, characterized in that: The step of extracting common features, facial features, head posture features and lip features based on the input face image comprises: Extracting the overall framework of the input face image based on a face modeling method; The common features, facial features, head posture features and lip features of the input facial image are determined in combination with the overall framework.
9. A facial animation generation device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for generating a facial animation according to any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the method for generating a facial animation according to any one of claims 1 to 8 are implemented.