Method and apparatus for generating facial animation, electronic device, computer-readable storage medium, and computer program product
Patent Information
- Application Number
- PCT/CN2026/077331
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-11
- Filing Date
- 2026-02-05
- Publication Date
- 2026-09-17
Smart Images

Figure CN2026077331_17092026_PF_FP_ABST
Abstract
Description
Methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for generating facial animations.
[0001] Cross-references to related applications
[0002] This application is based on and claims priority to Chinese Patent Application No. 202510295484.5, filed on March 11, 2025, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating facial animation. Background Technology
[0004] Voice-driven facial animation uses the analysis and processing of voice data to drive changes in the facial expressions and lip movements of virtual characters. Emotional and intonation information from the voice data is used to adjust the facial model's expression parameters, ensuring the generated facial animation is synchronized with the spoken content.
[0005] In related technologies, facial animation is usually generated by controlling the overall facial expression of a virtual object through voice. This can lead to imprecise control of the facial expressions of different parts of the virtual object's head, resulting in poor compatibility between the generated facial animation and the voice. Summary of the Invention
[0006] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating facial animation, which can effectively improve the compatibility between facial animation and voice of virtual objects.
[0007] The technical solution of this application embodiment is implemented as follows:
[0008] This application provides a method for generating facial animation, including:
[0009] Noise features are obtained by extracting features from preset noise, and a first speech feature sequence is obtained by extracting features from preset speech. The first speech feature sequence includes multiple first speech features.
[0010] The noise feature is decomposed into multiple sub-noise features, each of which corresponds to a part of the head of the virtual object.
[0011] The plurality of sub-noise features are fused with the first speech feature sequence to obtain a first fused feature sequence;
[0012] Based on the first fused feature sequence, a driving parameter sequence matching the first speech feature sequence is determined. The driving parameter sequence includes multiple driving parameters, which are used to control at least one of the facial expressions and head postures of the virtual object. Based on the driving parameter sequence, a facial animation of the virtual object is generated.
[0013] This application provides a facial animation generation apparatus, comprising:
[0014] The feature extraction module is configured to extract noise features from preset noise and extract features from preset speech to obtain a first speech feature sequence, wherein the first speech feature sequence includes multiple first speech features.
[0015] The decomposition module is configured to decompose the noise feature into multiple sub-noise features, each of which corresponds to a part of the head of the virtual object.
[0016] The fusion module is configured to fuse the plurality of sub-noise features with the first speech feature sequence to obtain a first fused feature sequence;
[0017] A determining module is configured to determine a driving parameter sequence matching the first speech feature sequence based on the first fused feature sequence. The driving parameter sequence includes multiple driving parameters, which are used to control at least one of the facial expressions and head postures of the virtual object. Based on the driving parameter sequence, a facial animation of the virtual object is generated. This application embodiment provides an electronic device, including:
[0018] Memory, configured to store computer-executable instructions or computer programs;
[0019] When a processor is configured to execute computer-executable instructions or computer programs stored in the memory, it implements the facial animation generation method provided in the embodiments of this application.
[0020] This application provides a computer-readable storage medium storing computer-executable instructions configured to, when executed by a processor, implement the facial animation generation method provided in this application.
[0021] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the facial animation generation method described above in this application.
[0022] The embodiments of this application have the following beneficial effects:
[0023] Noise features are obtained by extracting features from preset noise, and a first speech feature sequence is obtained by extracting features from preset speech. The noise features are decomposed into multiple sub-noise features. Since each sub-noise feature corresponds to a part of the virtual object's head, subsequent fusion and control have higher granularity and are therefore more accurate. By fusing the sub-noise features with the first speech feature sequence, a first fused feature sequence is obtained, thus combining speech and noise information and providing more comprehensive information for determining the driving parameters. Based on the first fused feature sequence, a driving parameter sequence matching the first speech feature sequence is determined. Since the noise features correspond to different parts of the virtual object's head, the driving parameter sequence determined based on the first fused feature sequence can achieve precise control of the virtual object's facial expressions and head posture. Through the driving parameter sequence, facial animation of the virtual object is generated, enabling the driving parameter sequence to accurately control the virtual object's facial expressions and head posture, thereby effectively improving the adaptation between the virtual object's facial animation and speech. Attached Figure Description
[0024] Figure 1 is a schematic diagram of the architecture of the facial animation generation system provided in an embodiment of this application;
[0025] Figure 2 is a schematic diagram of the structure of an electronic device for generating facial animation provided in an embodiment of this application;
[0026] Figure 3 is a schematic flowchart of the facial animation generation method provided in an embodiment of this application;
[0027] Figure 4 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application.
[0028] Figure 5 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application.
[0029] Figure 6 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application;
[0030] Figure 7 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application;
[0031] Figure 8 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application;
[0032] Figure 9 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application;
[0033] Figure 10 is a schematic diagram of the principle of the facial animation generation method provided in the embodiment of this application;
[0034] Figure 11 is a schematic diagram of the principle of the facial animation generation method provided in the embodiment of this application;
[0035] Figure 12 is a schematic diagram of the facial animation effect provided in an embodiment of this application;
[0036] Figure 13 is a second schematic diagram of the facial animation effect provided in the embodiment of this application;
[0037] Figure 14 is a schematic diagram of the facial animation effect provided in the embodiment of this application. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0039] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0040] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0042] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0043] 1) Virtual Objects: Virtual objects refer to the images of various people and things that can be interacted with in a virtual scene, or movable objects in the virtual scene. These movable objects can be virtual characters, virtual animals, anime characters, etc., such as people, animals, plants, oil drums, walls, stones, etc., displayed in the virtual scene. A virtual object can be a virtual avatar representing the user within the virtual scene. A virtual scene can include multiple virtual objects, each with its own shape and volume, occupying a portion of the space in the virtual scene. Optionally, the virtual object can be a user character driven by client-side operations, an artificial intelligence (AI) trained and set up for virtual scene battles, or a non-user character (NPC) set up for virtual scene interaction. The virtual object can also be a virtual character engaging in adversarial interaction within the virtual scene. Optionally, the number of virtual objects participating in the interaction in the virtual scene can be pre-set or dynamically determined based on the number of clients joining the interaction.
[0044] 2) Facial Expressions: Facial expressions refer to the movements and changes in the facial muscles of a virtual object, used to express emotions or reactions. Expression Basis: A set of basic facial movements or expressions that can be combined to create more complex facial expressions. Expression Synthesis: The process of combining multiple basic expression or action units (AUs) to generate complex facial expressions.
[0045] 3) Head Pose: Head pose refers to the orientation and position of a virtual object's head in three-dimensional space. It is typically described using three rotation angles: Yaw (left-right rotation), Pitch (up-down rotation), and Roll (left-right tilt). Yaw: Indicates the angle at which the head turns left or right, i.e., the head faces left or right. Pitch: Indicates the angle at which the head turns up or down, such as nodding or bowing. Roll: Indicates the angle at which the head tilts around the vertical axis, similar to a left-right tilt.
[0046] 4) Facial Animation: Virtual Object Facial Animation (VAFA) refers to the process of generating realistic facial expressions and movements for virtual characters or objects using technical means in computer graphics and digital media production. The aim is to enable virtual characters to express emotions and communicate like real humans. Facial animation has wide applications in film, games, virtual reality (VR), augmented reality (AR), animation, education, and training. It not only enhances the realism of virtual character performances but also improves user experience and immersion, providing viewers with a more vivid and engaging visual experience. Animation is an art form and technique that simulates or creates dynamic visual effects by creating and continuously playing static images frame by frame. Essentially, it uses the persistence of vision effect to make viewers perceive movement within the image by rapidly displaying a series of still images.
[0047] 5) Driving parameters: In computer graphics and virtual reality, these are a set of numerical values or signals used to control the animation and behavior of virtual objects. These parameters can originate from various data sources, such as speech, noise, and motion capture data, and are converted into instructions for controlling the movement and expressions of virtual objects through algorithms or models. The role of driving parameters is to convert abstract audio features into specific motion instructions for virtual objects, enabling the animation of virtual objects to be synchronized with speech and noise input, thereby achieving a more natural and immersive interactive experience. These parameters can be automatically generated by machine learning models or manually adjusted by animators to achieve the best animation effect.
[0048] During the implementation of the embodiments of this application, the applicant discovered the following problems with the related technology:
[0049] In related technologies, facial animation generation is usually achieved by controlling the overall facial expressions of a virtual object through voice. This can lead to imprecise control of the expressions of different parts of the virtual object's head, resulting in poor compatibility between the generated facial animation and the voice.
[0050] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating facial animation, which can effectively improve the compatibility between facial animation and voice of virtual objects. The following describes an exemplary application of the facial animation generation system provided in this application.
[0051] Referring to Figure 1, which is a schematic diagram of the architecture of the facial animation generation system 100 provided in the embodiment of this application, the terminal (exemplarily shown as terminal 400) connects to the server 200 through the network 300, which can be a wide area network or a local area network, or a combination of both.
[0052] Terminal 400 is configured for users to use client 410 to display facial animations on graphical interface 410-1 (graphical interface 410-1 is shown as an example). Terminal 400 and server 200 are interconnected via wired or wireless network.
[0053] In some embodiments, server 200 can be a standalone physical server, a server cluster or business system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal 400 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, in-vehicle terminal, etc., but is not limited to these. The electronic device provided in this application embodiment can be implemented as a terminal or a server. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application embodiment.
[0054] In some embodiments, the server 200 decomposes the noise features obtained by encoding preset noise into multiple sub-noise features based on the various parts included in the virtual object's header, fuses the sub-noise features with a first speech feature sequence to obtain a first fused feature sequence, determines a driving parameter sequence based on multiple first fused feature sequences, and sends the driving parameter sequence to the terminal 400. The terminal 400 generates a facial animation of the virtual object based on the driving parameter sequence.
[0055] In other embodiments, the terminal 400 decomposes the noise features obtained by encoding preset noise into multiple sub-noise features based on the various parts included in the virtual object's head, fuses the sub-noise features with a first speech feature sequence to obtain a first fused feature sequence, determines a driving parameter sequence based on multiple first fused feature sequences, and sends the driving parameter sequence to the server 200. The server 200 generates a facial animation of the virtual object based on the driving parameter sequence.
[0056] As an example, in applications such as virtual news broadcasting or online education, the system acquires synthesized audio or live recordings corresponding to news articles as preset speech and generates random signals that satisfy a specific distribution as preset noise. The system first performs feature encoding on this random signal to obtain noise features and extracts acoustic features such as Mel-frequency cepstral coefficients from the audio to construct a first speech feature sequence. Subsequently, the system segments or maps the global noise features along the channel dimension, decomposing them into multiple sub-noise features corresponding to the virtual anchor's eye region, mouth region, and eyebrow region, respectively. The system deeply fuses these sub-noise features targeting specific facial anatomy structures with the first speech feature sequence containing semantic and prosodic information, generating a first fused feature sequence containing multimodal information. The decoding module outputs a driving parameter sequence based on this first fused feature sequence. The parameters in this sequence are used to precisely control the virtual anchor's lip opening and closing amplitude, eye blinking frequency, and subtle eyebrow movements. Finally, the rendering engine uses this driving parameter sequence to generate a virtual object facial animation that is highly synchronized with the broadcast content and has natural micro-expressions, making the virtual anchor appear lifelike during the broadcast.
[0057] As an example, in applications such as virtual social interaction or immersive online games, the system receives real-time voice input from users as preset speech and introduces Gaussian noise as preset noise to increase the diversity of movements. The system extracts features from the Gaussian noise to obtain noise features and converts the voice input into a first speech feature sequence containing phonemes and emotional features. To give the virtual avatar more realistic body language, the system decomposes the noise features into multiple sub-noise features corresponding to the rotation of the virtual object's neck joint and the opening and closing of the jawbone. By fusing these sub-noise features with the first speech feature sequence under an attention mechanism, a first fused feature sequence is obtained. The driving parameter sequence generated on this basis not only includes facial expression parameters that control facial muscle deformation but also posture parameters that control the head's pitch or lateral rotation in three-dimensional space. The system uses this driving parameter sequence to drive the virtual avatar to speak with natural head shaking and jaw movement, thereby generating high-fidelity facial animation and enhancing the user's interactive experience in the virtual environment. Referring to Figure 2, which is a schematic diagram of the structure of an electronic device 500 for generating facial animation according to an embodiment of this application, the electronic device 500 shown in Figure 2 can be the server 200 or terminal 400 in Figure 1. The electronic device 500 shown in Figure 2 includes at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together through a bus system 440. It is understood that the bus system 440 is configured to enable communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a drive bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 440 in Figure 2.
[0058] Processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0059] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 430.
[0060] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0061] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0062] Operating system 451 includes system programs configured to handle various basic system services and perform hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., configured to implement various basic business functions and handle hardware-based tasks;
[0063] The network communication module 452 is configured to reach other electronic devices via one or more (wired or wireless) network interfaces 420, such as Bluetooth, WiFi, and Universal Serial Bus.
[0064] In some embodiments, the facial animation generation apparatus provided in this application can be implemented in software. Figure 2 shows a facial animation generation apparatus 455 stored in memory 450, which can be software in the form of programs and plug-ins, including the following software modules: feature extraction module 4551, decomposition module 4552, fusion module 4553, and determination module 4554. These modules are logically related and can therefore be arbitrarily combined or further split according to their implemented functions. The functions of each module will be described below.
[0065] In other embodiments, the facial animation generation apparatus provided in this application can be implemented in hardware. As an example, the facial animation generation apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the facial animation generation method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0066] In some embodiments, the terminal or server can implement the facial animation generation method provided in this application by running a computer program or computer-executable instructions. For example, the computer program can be a native program in the operating system (e.g., a dedicated animation generation program) or a software module, such as an animation generation module that can be embedded in any program (e.g., an instant messaging client, a photo album program, an electronic map client, a navigation client); or it can be a native application (APP), i.e., a program that needs to be installed in the operating system to run. In summary, the above-mentioned computer program can be any form of application, module, or plugin.
[0067] The method for generating facial animation provided in this application will be described in conjunction with exemplary applications and implementations of the server or terminal provided in the embodiments of this application.
[0068] Referring to Figure 3, which is a flowchart of the facial animation generation method provided in this application embodiment, the method will be described in conjunction with steps 101 to 105 shown in Figure 3. The facial animation generation method provided in this application embodiment can be implemented by the server or the terminal alone, or by the server and the terminal working together. The following description will take the implementation by the server alone as an example.
[0069] In step 101, noise features are extracted from the preset noise, and a first speech feature sequence is obtained by extracting features from the preset speech.
[0070] In some embodiments, the first speech feature sequence includes multiple first speech features. Preset noise refers to any unwanted interference components in the signal that reduce its clarity and intelligibility. Noise can originate from various sources, such as environmental noise (traffic, wind, etc.), noise from electronic devices, mechanical noise, etc. In speech processing, noise typically refers to background noise, which affects the quality of the speech signal; noise can be of any form.
[0071] In some embodiments, preset speech refers to the sound produced by humans through vocal organs (such as vocal cords, tongue, lips, etc.) for language communication. Speech signals contain a large amount of information, such as speech content, speaker's identity, emotional state, etc. In speech processing, speech signals usually need to be digitized for subsequent processing and analysis.
[0072] In some embodiments, feature extraction is a crucial step in signal processing, referring to the extraction of characteristic parameters from the raw signal that represent its properties. These characteristic parameters are commonly used in tasks such as pattern recognition, classification, and clustering. The goal of feature extraction is to reduce the dimensionality of the data while preserving the important information of the signal.
[0073] In some embodiments, noise features refer to characteristic parameters extracted from a noise signal that represent the characteristics of the noise. These characteristic parameters can be used for tasks such as noise identification, noise classification, and noise suppression. Common noise features include the power spectral density of noise and the statistical properties of noise (such as mean, variance, kurtosis, etc.).
[0074] In some embodiments, the first speech feature sequence refers to a series of speech feature parameters extracted from a speech signal. These feature parameters are typically arranged in chronological order to form a sequence. Common speech features in the first speech feature sequence include Mel-frequency cepstral coefficients (MFCC), linear predictive coding (LPC) coefficients, spectral envelope, fundamental frequency (F0), etc.
[0075] In some embodiments, the above-mentioned feature extraction of preset noise to obtain noise features can be achieved in the following manner: feature extraction is performed on the preset noise to obtain the initial noise features of the noise, and feature extraction is performed on the head data of the virtual object to obtain the head features of the virtual object; the initial noise features and the head features are fused to obtain the noise features.
[0076] In some embodiments, the initial noise features and head features are fused as described above. Fusion refers to the process of combining data or features from different sources to produce a more comprehensive and accurate result. The fusion of initial noise features and head features can be weighted averaging, feature-level fusion, decision-level fusion, etc. The purpose is to combine the information of the two feature sets to form a comprehensive noise feature representation. The fused noise features should be able to better describe the relationship between noise and the virtual object head.
[0077] In some embodiments, the purpose of feature extraction on preset noise is to obtain the characteristics of the noise. Common noise feature extraction methods include calculating the power spectral density of the noise, statistical characteristics (such as mean, variance, kurtosis, etc.), and using filter banks to extract the frequency domain features of the noise. The head data of a virtual object may include the head's position, orientation, motion state, the corresponding 3D mesh model of the head, or vertex data, etc. Feature extraction on this data is to obtain the motion characteristics of the head.
[0078] In some embodiments, the above-mentioned fusion of the initial noise features and the head features to obtain the noise features can be achieved by splicing the initial noise features and the head features together to obtain the noise features.
[0079] In other embodiments, the above-mentioned fusion of the initial noise features and the head features to obtain the noise features can also be achieved by: using the initial noise features as the key vector and the head features as the query vector and value vector, calling a self-attention model to fuse the initial noise features and the head features to obtain the noise features.
[0080] In other embodiments, the above-mentioned fusion of the initial noise features and the head features to obtain the noise features can also be achieved by: using the head features as the key vector and the initial noise features as the query vector and value vector, calling a self-attention model to fuse the initial noise features and the head features to obtain the noise features.
[0081] In other embodiments, the above-described fusion of the initial noise features and the head features to obtain the noise features can also be achieved in the following ways: determining the initial noise features as a key vector and the head features as a value vector and a query vector; or determining the initial noise features as the value vector and the query vector and the head features as the key vector; and invoking an attention model to fuse the key vector, the value vector, and the query vector to obtain the noise features.
[0082] In some embodiments, initial noise features and head features are mapped to different combinations of the three basic vectors of the attention model: query vector, key vector, and value vector. The first approach uses noise as the key and head features as the query and value, meaning the model uses noise features as an index to retrieve and reconstruct information from head features, focusing on using noise to modulate head information. The second approach is the opposite, using head features as the key and focusing on extracting and reconstructing noise features based on the semantic distribution of head features. This technical solution demonstrates that the invention lies not only in using an attention model but also in revealing a bidirectional, interchangeable interaction between these two features. Regardless of which is used as the benchmark, the essence is to capture the deep correlation between the two through cross-attention, thereby generating an enhanced noise feature that integrates the contextual information of both.
[0083] In some embodiments, a self-attention model is a neural network mechanism that allows the model to dynamically focus on different parts of a sequence when processing sequential data. The core idea of a self-attention model is based on the interaction of three vectors: query, key, and value. In a self-attention model, each input position generates these three vectors: a query vector (representing the feature representation of the current input position when searching for relevant information), a key vector (representing the feature representation of each position in the entire sequence, used for matching with the query vector), and a value vector (representing the actual information content of each position in the sequence, which is what the model ultimately focuses on and outputs). The self-attention model generates an attention weight matrix by calculating the similarity between the query vector and all key vectors. This weight matrix represents the degree of attention the current input position pays to other positions in the sequence. The model then applies these weights to the value vectors, generating the output representation through a weighted summation. A significant advantage of the self-attention model is its ability to capture long-range dependencies in a sequence without processing elements one by one like an RNN. Furthermore, the introduction of the multi-head attention mechanism enables the self-attention model to capture information from multiple different representation subspaces, thereby enhancing the model's expressive power.
[0084] In some embodiments, virtual objects are images of various people and objects that can interact with each other, or movable objects in a virtual scene. These movable objects can be virtual characters, virtual animals, anime characters, etc., such as people, animals, plants, oil drums, walls, stones, etc., displayed in a virtual scene. A virtual object can be a virtual avatar representing the user within the virtual scene. A virtual scene can include multiple virtual objects, each with its own shape and volume, occupying a portion of the space within the virtual scene. Optionally, the virtual object can be a user character driven by client-side operations, an artificial intelligence (AI) trained and set up for virtual scene battles, or a non-user character (NPC) set up for virtual scene interaction. A virtual object can be a virtual character engaging in adversarial interaction within the virtual scene. The number of virtual objects participating in the interaction in the virtual scene can be pre-set or dynamically determined based on the number of clients joining the interaction.
[0085] In some embodiments, the initial noise features and the head features are fused to obtain the noise features. This can be achieved by assigning different weights to different features and calculating the fused feature values based on the weights. The vectors of different features are then directly concatenated to form a longer feature vector. Alternatively, deep learning models such as neural networks can be used to learn the complex relationships between features and fuse them. Fusion can also be performed using the self-attention layer method shown in Figure 10.
[0086] In some embodiments, the head data of a virtual object can be: head position, describing the specific coordinates of the virtual object's head in virtual space, typically represented in three-dimensional spatial coordinates (X, Y, Z); head orientation, describing the orientation or pointing of the virtual object's head, typically represented using quaternions or Euler angles; head rotation, referring to the rotational movement of the virtual object's head, which can be described by rotation angle and axis of rotation; head movement, describing the movement of the virtual object's head, including translation and rotation; head pose, combining information such as head position, orientation, and rotation to describe the overall pose of the virtual object's head; and head acceleration, describing the acceleration of the virtual object's head, typically obtained through a built-in accelerometer.
[0087] Thus, by extracting features from preset noise, we can obtain its initial characteristics, which helps us to gain a deeper understanding of the noise's nature and provides a basis for subsequent noise processing and suppression. Extracting features from the head data of virtual objects not only captures their motion in the virtual environment but also allows us to infer the user's intentions and focus based on head posture and direction. Fusing these two sets of features generates more comprehensive and accurate noise characteristics, improving the realism of sound in the virtual environment. This allows users to experience more realistic sound effects in an immersive experience, and enhances speech clarity and intelligibility in noisy environments, thereby improving the quality of communication and interaction.
[0088] In step 102, the noise feature is decomposed into multiple sub-noise features.
[0089] In some embodiments, multiple sub-noise features each correspond to a part of the virtual object's head. The noise features are decomposed based on each part of the virtual object's head, refining the overall noise features into multiple sub-noise features corresponding to different parts of the virtual object's head. The virtual object's head may contain multiple different parts, such as eyes, nose, mouth, cheeks, etc., and each part exhibits different behaviors and actions in the virtual environment.
[0090] As an example, the virtual object header includes the following parts: part A, part B, part C, and part D. The noise features include multiple row features (row feature 1, row feature 2, row feature 3, row feature 4, row feature 5, row feature 6, row feature 7, and row feature 8) or multiple column features. The noise features are decomposed into multiple sub-noise features. Taking row features as an example, the decomposition process is explained as follows: row features 1 and 2 are combined to obtain the sub-noise features corresponding to part A; row features 3 and 4 are combined to obtain the sub-noise features corresponding to part B; row features 5 and 6 are combined to obtain the sub-noise features corresponding to part C; and row features 7 and 8 are combined to obtain the sub-noise features corresponding to part D.
[0091] Continuing with the previous example, if the number of row features included in the noise feature cannot be evenly distributed among the various parts of the header, then the noise feature can be decomposed into multiple sub-noise features by adding row features with all feature elements equal to 0 to at least some of the sub-noise features.
[0092] Continuing the previous example, in the noise feature decomposition process of the virtual object's head, the head is first divided into four parts: part A, part B, part C, and part D. Each part corresponds to a different region of the head, such as the eyes, nose, mouth, and ears. Next, consider the noise features, which can be multiple row features, such as row features 1 to 8. Each row feature represents noise information in a specific direction in the image. To associate these noise features with the different parts of the head, they need to be decomposed into multiple sub-noise features. Taking row features as an example, row features 1 and 2 can be combined to obtain the sub-noise features corresponding to part A. Similarly, row features 3 and 4 are combined to form the sub-noise features of part B, row features 5 and 6 are combined to form the sub-noise features of part C, and row features 7 and 8 are combined to form the sub-noise features of part D. In this way, the original noise features are decomposed into four sub-noise features, each corresponding to a part of the head. However, if the number of noise features cannot be evenly distributed across the different parts of the head, even distribution can be achieved by adding row features with all feature elements equal to 0 to at least some of the sub-noise features. For example, if there are only 7 noise features, the first 6 row features can be assigned to parts A, B, C, and D respectively, with one or two row features assigned to each part. Then, a row feature with all feature elements equal to 0 can be added to the last part to ensure that each part has a corresponding sub-noise feature.
[0093] As an example, the extracted noise features are decomposed into multiple sub-noise features, each corresponding to a facial feature of the virtual object, such as the eyes, eyebrows, and mouth. These sub-noise features are then fused with a first speech feature sequence to obtain a first fused feature sequence. The noise features and speech features are combined to form a comprehensive feature representation. Based on the first fused feature sequence, the system determines a sequence of driving parameters that matches the first speech feature sequence. These driving parameters include multiple parameters used to control the facial expressions of the virtual object, such as the degree of eye opening, eyebrow raising or lowering, and mouth opening and closing. The driving parameter sequence is used to generate a video of the virtual object, and the driving parameters are applied to the 3D model of the virtual object to generate a facial video synchronized with the speech.
[0094] Thus, by decomposing the noise features of a virtual object's head into sub-noise features corresponding to each part, refined noise processing can be achieved, significantly enhancing the visual realism and detail of the virtual avatar. This decomposition allows for personalized noise reduction processing for different areas of the head; for example, noise in the eye area can be softened to maintain the liveliness of the gaze, while noise in the mouth area can be sharpened to enhance the clarity of the lip movements. When the number of noise features cannot be evenly distributed, adding row features with zero feature elements ensures that each part has a corresponding sub-noise feature, avoiding inconsistencies caused by uneven noise distribution and improving the consistency and stability of the processing results.
[0095] In step 103, the plurality of sub-noise features are fused with the first speech feature sequence to obtain a first fused feature sequence.
[0096] In some embodiments, the first fused feature sequence corresponds one-to-one with the sub-noise features. By fusing the visual sub-noise features with the auditory speech feature sequence, multimodal information can be integrated, making the virtual avatar not only visually realistic but also auditorily consistent with the visual presentation, enhancing the user's immersion and experience. The fused feature sequence can more naturally reflect the changes in the virtual avatar's facial expressions and movements under different speech states. For example, when the virtual avatar speaks, its facial expressions and head movements will match the speech content, making the overall expression more natural and realistic. The first fused feature sequence can be used to drive the real-time interaction of the virtual avatar, enabling the virtual object to respond accordingly to the user's voice input, improving the interactivity and responsiveness of the virtual avatar.
[0097] In some embodiments, fusion features refer to a new set of features generated by combining visual sub-noise features with a first auditory speech feature sequence. The aforementioned first fusion feature sequence includes sub-noise features and a first speech feature sequence, with each fusion feature corresponding to a sub-noise feature. It contains both visual and auditory information, aiming to achieve synchronization and coordination of visual and auditory information. In virtual avatar generation and animation production, the use of fusion features can make the expression of virtual avatars more natural and realistic. For example, when a virtual avatar speaks, its facial expressions and head movements will match the speech content, thereby enhancing the user's immersion and experience. Furthermore, fusion features can also be used to drive real-time interaction of virtual avatars, enabling them to respond accordingly to the user's voice input, improving the interactivity and responsiveness of the virtual avatar.
[0098] In some embodiments, the above-mentioned fusion refers to the process of integrating sub-noise features with the first speech feature sequence to obtain a first fused feature sequence. Fusion is a multi-level feature integration process, which aims to combine speech feature sequences and noise feature sequences (or subsets thereof) through specific operations to generate a new fused feature sequence. Fusion refers to the process of integrating sub-noise features with the first speech feature sequence to generate a first fused feature sequence through the following two methods. The fusion between the sub-noise features and the first speech feature sequence can be a process of concatenating the sub-noise features with each first speech feature in the first speech feature sequence. Fusion can be simple concatenation, weighted summation, attention mechanism, etc. The fusion process will be described below with reference to steps 1031A to 1032A and steps 1031B to 1033B shown in Figure 4.
[0099] In some embodiments, referring to FIG4, FIG4 is a schematic flowchart of the facial animation generation method provided in the present application embodiment. Step 103 shown in FIG3 can be implemented by steps 1031A to 1032A shown in FIG4.
[0100] In step 1031A, the first speech feature sequence and the noise features are fused to obtain the second speech feature sequence.
[0101] In some embodiments, the second speech feature sequence refers to a new set of speech feature sequences generated by fusing the first speech feature sequence and noise features. The first speech feature sequence is typically extracted from the original speech signal using speech recognition or speech processing techniques and contains acoustic features of the speech, such as frequency, intensity, and duration.
[0102] In some embodiments, step 1031A above can be implemented as follows: for each first speech feature in the first speech feature sequence, the first speech feature and the noise feature are concatenated to obtain a concatenated feature; the concatenated features are arranged according to the order of the first speech features in the first speech feature sequence to obtain the second speech feature sequence.
[0103] In some embodiments, for each first speech feature in the first speech feature sequence, the first speech feature and noise feature are concatenated to obtain a concatenated feature; then, the concatenated features are arranged according to the order of the first speech features in the first speech feature sequence to obtain a second speech feature sequence. For each first speech feature in the first speech feature sequence, it is concatenated with the corresponding noise feature. This concatenation can be a simple concatenation or a combination of the two using an algorithm. The purpose of concatenation is to integrate the information of the noise feature into the speech feature so that this information can be better recognized and utilized in subsequent processing. The feature sequence after concatenation and arrangement is called the second speech feature sequence. This sequence contains comprehensive information of the original speech features and noise features, which can be used to improve the accuracy of speech recognition, especially in noisy environments.
[0104] In some embodiments, feature concatenation refers to combining corresponding features from two or more feature sequences to form a new feature sequence. In speech processing, feature concatenation is often used to combine speech features with noise features to improve the quality and robustness of the speech signal. Concatenation can be achieved through simple concatenation, weighted fusion, or other feature fusion methods. Arrangement refers to sorting features in a certain order to maintain the temporal order and contextual information of the feature sequence. In speech processing, arrangement is often used to maintain the temporal relationship of speech feature sequences.
[0105] As an example, suppose we have a first speech feature sequence containing three first speech features: F1, F2, and F3. Simultaneously, we have a corresponding noise feature sequence containing three noise features: N1, N2, and N3. Concatenating F1 and N1 yields the concatenated feature P1 = [F1, N1]. Concatenating F2 and N2 yields the concatenated feature P2 = [F2, N2]. Concatenating F3 and N3 yields the concatenated feature P3 = [F3, N3]. Arranging the concatenated features P1, P2, and P3 in the order of the original first speech feature sequence results in a second speech feature sequence: [P1, P2, P3]. This gives us a second speech feature sequence containing both speech and noise features, which can be used for speech processing tasks such as speech recognition or speech enhancement.
[0106] Thus, by concatenating each first speech feature in the first speech feature sequence with its corresponding noise feature and arranging them in the original order, the resulting second speech feature sequence can more comprehensively reflect the characteristics of the speech signal. This helps improve the accuracy and robustness of speech recognition, especially in noisy environments. By combining the information from speech features and noise features, the influence of noise can be better suppressed, enhancing the intelligibility and clarity of the speech.
[0107] In step 1032A, the plurality of sub-noise features are fused with the second speech feature sequence to obtain a plurality of first fused feature sequences.
[0108] In some embodiments, the first fusion feature sequence refers to a new set of feature sequences generated by fusing sub-noise features with a second speech feature sequence. In virtual avatar generation and animation production, the use of the first fusion feature sequence can make the virtual avatar's expression more natural and realistic. For example, when a virtual avatar speaks, its facial expressions and head movements will match the speech content, thereby enhancing the user's immersion and experience. Furthermore, the first fusion feature sequence can also be used to drive real-time interaction of the virtual avatar, enabling it to respond accordingly to the user's voice input, improving the virtual avatar's interactivity and responsiveness.
[0109] In some embodiments, step 1032A above can be implemented as follows: performing two different linear transformations on the second speech feature sequence to obtain a third speech feature sequence and a fourth speech feature sequence; performing the following processing on each sub-noise feature respectively: fusing the third speech feature sequence and the sub-noise feature to obtain a second fused feature sequence; fusing the fourth speech feature sequence and the second fused feature sequence to obtain a first fused feature sequence corresponding to the sub-noise feature.
[0110] In some embodiments, the second speech feature sequence is subjected to two different linear transformations to obtain a third and a fourth speech feature sequence. Different linear transformations are used to extract and emphasize different aspects of the speech features. These linear transformations can be matrix multiplication, weighted summation, or other linear operations to adjust the weights and combinations of features to adapt to different processing requirements. The third speech feature sequence and sub-noise features are then fused to obtain a second fused feature sequence. This step aims to combine the linearly transformed speech features with the noise features to improve the quality and robustness of the speech signal. Fusion can be achieved through simple addition, multiplication, or other fusion algorithms, aiming to incorporate noise feature information into the speech features so that this information can be better identified and utilized in subsequent processing. The fourth speech feature sequence and the second fused feature sequence are then fused to obtain a first fused feature sequence. This aims to combine the speech features that have undergone different linear transformations and fusion processes to achieve a more comprehensive and refined feature representation.
[0111] In some embodiments, the second speech feature sequence is subjected to two different linear transformations to obtain a third speech feature sequence and a fourth speech feature sequence. The transformation coefficients of the two different linear transformations are different, where the difference in transformation coefficients means that each linear transformation emphasizes or extracts different aspects of the speech features. The difference in transformation coefficients can lead to differences in the feature sequences in terms of dimension, weight, and combination, thereby providing diverse feature information for the subsequent fusion process.
[0112] In some embodiments, the third speech feature sequence refers to a new set of speech features obtained by performing a first linear transformation on the original speech feature sequence (the second speech feature sequence). This transformation may be achieved by applying specific mathematical algorithms or filter banks, with the aim of changing or enhancing certain characteristics of the original speech, such as frequency distribution and resonance characteristics, to meet the needs of specific speech processing tasks. The fourth speech feature sequence refers to a set of speech features obtained by performing a second, different linear transformation on the same original speech feature sequence (the second speech feature sequence). This transformation may also employ different algorithms or filter banks, with the aim of mining speech information or extracting features complementary to the first transformation to enrich the speech feature set. The second fused feature sequence refers to a set of fused speech features obtained by combining the third speech feature sequence with sub-noise features. Sub-noise features represent noise components in the speech signal. The fusion method may be simple summation, weighted summation, or other more complex operations, with the aim of modeling the noise components to a certain extent while preserving effective speech information, in order to facilitate subsequent noise suppression or speech enhancement processing.
[0113] Thus, by performing two different linear transformations on the second speech feature sequence, we can obtain the third and fourth speech feature sequences. These two sequences reveal different aspects of the original speech features, enhancing their expressiveness and discriminative power. Fusing the third speech feature sequence with the sub-noise features yields the second fused feature sequence. This fusion effectively combines noise information with speech features, aiding in better noise recognition and suppression during subsequent processing. Fusing the fourth speech feature sequence with the second fused feature sequence results in the first fused feature sequence, which not only contains rich speech information but also integrates noise characteristics. This feature sequence helps improve the robustness of the speech recognition system, reduce noise interference, and enhance recognition accuracy and speech enhancement quality.
[0114] In some embodiments, the first speech feature sequence includes a plurality of first speech features, and the third speech feature sequence includes second speech features that correspond one-to-one with the first speech features.
[0115] In some embodiments, a one-to-one correspondence refers to a bijective relationship between two sets, meaning that every element in the first set has one and only one corresponding element in the second set, and vice versa. This correspondence is complete and reversible. A one-to-one correspondence means that every first speech feature in the first speech feature sequence has one and only one corresponding second speech feature in the third speech feature sequence, and at the same time, every second speech feature in the third speech feature sequence also has one and only one corresponding first speech feature in the first speech feature sequence.
[0116] As an example, suppose the first speech feature sequence is A, B, C, and the third speech feature sequence is XY, Z. If A corresponds to X, B corresponds to Y, and C corresponds to Z, then this is a one-to-one correspondence. Each element has exactly one corresponding element in both sequences.
[0117] In some embodiments, the above-described fusion of the third speech feature sequence and the sub-noise feature to obtain the second fused feature sequence can be achieved as follows: selecting a target speech feature from the third speech feature sequence, the target speech feature being associated with the sub-noise feature; constructing a fifth speech feature sequence based on the target speech feature and the second speech feature, the second speech feature being adjacent to the target speech feature in the third speech feature sequence; and fusing the fifth speech feature sequence and the sub-noise feature to obtain the second fused feature sequence.
[0118] In some embodiments, the third speech feature sequence is obtained by performing a linear transformation on the first speech feature sequence. Importantly, each feature in the third speech feature sequence corresponds one-to-one with a feature in the first speech feature sequence. This means that the third speech feature sequence is a transformed or enhanced version of the first speech feature sequence, preserving the structure and information of the original features. The first speech feature sequence contains multiple first speech features, which may be extracted from the original speech signal to represent certain characteristics of speech, such as Mel-frequency cepstral coefficients (MFs), linear deterministic codes (LPCs), etc. The third speech feature sequence contains second speech features that correspond one-to-one with the first speech features. This means that the third sequence is a transformation or enhancement of the first sequence, possibly generated by an algorithm or model, with the aim of improving the quality of speech features or highlighting certain features. Target speech features associated with sub-noise features are selected from the third speech feature sequence. Sub-noise features may refer to noise components contained in the speech signal or features under specific noise conditions. Selecting speech features associated with noise may be for subsequent noise suppression or enhancement processing. Based on the target speech features and the second speech features, a fifth speech feature sequence is constructed. The second speech feature is adjacent to the target speech feature, which means that the relationship between adjacent features was considered when constructing the new feature sequence, possibly utilizing local contextual information. The fifth speech feature sequence and the sub-noise features are fused to obtain the second fused feature sequence. The purpose of fusion is to combine speech features and noise features, possibly to more accurately identify useful information in the speech signal or to better suppress noise.
[0119] In some embodiments, the construction of the fifth speech feature sequence based on the target speech features and the second speech features can be achieved by constructing the fifth speech feature sequence from the target speech features and the second speech features adjacent to the target speech feature sequence in the third speech feature sequence.
[0120] As an example, the expression for the fifth speech feature sequence mentioned above can be: T = {T1, T2, T3} (1)
[0121] Where T is used to indicate the fifth speech feature sequence, T2 is used to indicate the target speech feature, and T1 and T3 are used to indicate the second speech feature in the third speech feature sequence that is adjacent to the target speech feature sequence.
[0122] In some embodiments, the associated target speech features refer to the speech features of the portion of the virtual object represented as the sub-noise feature in the third speech feature sequence. The target speech features are selected from the third speech feature sequence to specifically process those portions significantly affected by noise, or those speech features that can provide noise information. The association of these target speech features with the sub-noise features means that their corresponding positions or characteristics in the speech signal are directly related to the noise. The target speech features associated with the sub-noise features can refer to third speech features in the third speech feature sequence that correspond one-to-one with the sub-noise features, or multiple third speech features in the third speech feature sequence that correspond to the sub-noise features, or the third speech features of the portion of the virtual object in the third speech feature sequence that corresponds to the same sub-noise feature.
[0123] In some embodiments, selecting a target speech feature from a third speech feature sequence can be achieved as follows: for each third speech feature in the third speech feature sequence, the portion of the first speech feature corresponding to the virtual object and the portion of the sub-noise feature corresponding to the virtual object are compared to obtain a comparison result. When the comparison result indicates that the third speech feature and the portion of the sub-noise feature corresponding to the virtual object are the same, the third speech feature is determined as the target speech feature.
[0124] As an example, suppose there is a speech recognition system that attempts to identify a speaker's voice from a noisy recording. The recording contains a lot of background noise, such as traffic noise and wind noise. The goal is to identify clear speech while reducing the impact of noise. In this example, the first speech feature sequence is the speech features extracted from the original recording. These features may include frequency, amplitude, time-domain, and frequency-domain characteristics. The third speech feature sequence is a preprocessed speech feature sequence, which may include some enhanced or noise-reduced features. Now, it is necessary to select target speech features from the third speech feature sequence, which are associated with sub-noise features. Sub-noise features are those features that can represent the characteristics of noise, such as the frequency distribution of noise, the intensity of noise, etc. Specifically, suppose there is a sub-noise feature sequence that contains the frequency distribution of background noise. The goal is to select speech features from the third speech feature sequence that are associated with these noise frequency distributions. These target speech features may include speech features with high energy in the noise frequency range, or speech features that can provide noise information. By selecting these target speech features, it is possible to specifically process the parts that are heavily affected by noise, or to extract speech features that can provide noise information. These target speech features are associated with sub-noise features, meaning that their corresponding positions or characteristics in the speech signal are directly related to the noise. For example, suppose there is a speech feature in the third speech feature sequence whose frequency matches a frequency in a sub-noise feature. Then, this speech feature can be considered a target speech feature because it is associated with the sub-noise feature. Similarly, if multiple speech features in the third speech feature sequence match multiple frequencies in the sub-noise features, then these speech features can all be considered target speech features.
[0125] In some embodiments, the above-mentioned fusion of the fourth speech feature sequence and the second fusion feature sequence to obtain the first fusion feature sequence can be achieved in the following manner: for each fourth speech feature in the fourth speech feature sequence, the second fusion feature at the same position in the fourth speech feature sequence is concatenated to obtain the concatenated feature corresponding to the fourth speech feature; the concatenated features corresponding to each fourth speech feature are arranged according to the order of the fourth speech features in the fourth speech feature sequence to obtain the first fusion feature sequence.
[0126] In some embodiments, a one-to-one mapping relationship is established between the two feature sequences, treating the fourth speech feature sequence and the second fused feature sequence as two parallel time axes. During fusion, the system does not employ weighted summation or dot product operations that could obscure the original information; instead, it uses a physical dimension concatenation method. Specifically, for each specific time step or position index in the sequence, the algorithm directly concatenates the current speech feature vector with the fused feature vector at the same position along the channel dimension. This means that if the dimension of the speech feature is N and the dimension of the fused feature is M at a certain moment, the dimension of the concatenated feature will be expanded to N+M. This processing method ensures that the information from the two sources is juxtaposed without loss at the most microscopic feature unit, providing a wider range of data inputs for subsequent model processing.
[0127] In some embodiments, by constructing an enhanced sequence that combines the original acoustic characteristics with high-level contextual information, and reassembling these concatenated features into a first fused feature sequence according to their original order, the inherent temporal dependencies of the speech signal are strictly preserved. The advantage of this fusion mechanism is that it allows subsequent neural network layers to see both the current speech details (fourth speech feature) and refer to previous processing results or contextual information (second fused feature) when processing data in each time slice. This parallel information input method avoids mutual cancellation between features at different levels, effectively giving the model the ability to process multi-source information simultaneously, thereby significantly improving the expressive power and richness of the final generated feature sequence and laying a data foundation for achieving more accurate speech recognition or synthesis tasks.
[0128] Thus, by selecting target speech features associated with sub-noise features from the third speech feature sequence, and constructing a fifth speech feature sequence by combining it with adjacent second speech features, and then fusing this sequence with the sub-noise features to obtain the second fused feature sequence, this process can significantly improve the processing effect of speech signals. It not only enhances the discriminativeness of speech features and improves the accuracy of speech recognition, but also effectively reduces the impact of environmental noise on speech signals by fusing noise features, thereby improving the quality and robustness of speech signals.
[0129] Thus, fusing the first speech feature sequence and noise features to obtain the second speech feature sequence effectively combines speech and noise information, improving the robustness of the speech features. Subsequently, the sub-noise features are fused with the second speech feature sequence to obtain the first fused feature sequence, enhancing the speech features' adaptability to noise and significantly improving speech recognition and speech enhancement performance in noisy environments.
[0130] In some embodiments, referring to FIG5, FIG5 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application. Step 103 shown in FIG3 can be implemented by steps 1031B to 1033B shown in FIG5.
[0131] In step 1031B, a mask corresponding one-to-one with each of the first speech features in the first speech feature sequence is obtained.
[0132] In some embodiments, the mask here refers to a binary sequence or matrix used to indicate which features are valid and which need to be ignored or masked. A mask sequence is generated, which corresponds one-to-one with each feature in the first speech feature sequence. The mask value can be 0 or 1, where 1 indicates that the feature is valid, and 0 indicates that the corresponding feature needs to be ignored or masked.
[0133] In some embodiments, the above-mentioned acquisition of a mask corresponding one-to-one with each first speech feature in the first speech feature sequence can be achieved by performing the following processing on each first speech feature in the first speech feature sequence: determining the correlation between the first speech feature and the sub-noise feature; and determining the mask corresponding to the first speech feature based on the correlation.
[0134] In some embodiments, a mask is determined for each first speech feature based on correlation. If a speech feature has a high correlation with a sub-noise feature, indicating that it may be affected by noise, the corresponding mask value may be set to 0, meaning that this feature needs to be ignored or masked. Conversely, if the correlation is low, the mask value may be set to 1, indicating that this feature is valid and can be preserved.
[0135] In some embodiments, the above-mentioned determination of the mask corresponding to the first speech feature based on correlation can be achieved as follows: when the correlation indicates that the first speech feature is associated with the sub-noise feature, the mask corresponding to the first speech feature is determined to be 1; when the correlation indicates that the first speech feature is not associated with the sub-noise feature, the mask corresponding to the first speech feature is determined to be 0.
[0136] In this way, noise components in the speech feature sequence can be effectively filtered and adjusted, improving the quality and clarity of the speech signal. The correlation between each first speech feature and sub-noise feature is evaluated separately, ensuring that only those features affected by noise are labeled and processed, while other useful features remain unchanged.
[0137] In some embodiments, the determination of the association between the first speech feature and the sub-noise feature can be achieved as follows: when the first speech feature and the sub-noise feature correspond to the same part of the virtual object, the first speech feature is identified as the third speech feature, and the third speech feature is associated with the sub-noise feature; when the first speech feature is adjacent to the third speech feature in the first speech feature sequence, the first speech feature is associated with the sub-noise feature; when the first speech feature and the sub-noise feature correspond to different parts of the virtual object, and the first speech feature and the third speech feature are not adjacent, the first speech feature is not associated with the sub-noise feature. In some embodiments, the virtual object part affected by the currently processed first speech feature is identified, and it is compared with the virtual object part corresponding to the sub-noise feature. If the comparison result shows that both point to the same part of the virtual object, for example, both correspond to a facial region or a specific limb node, this indicates that the two have a direct basis for interaction in spatial semantics. In this case, the specific first speech feature is marked as the third speech feature, and an association is established between the third speech feature and the sub-noise feature. This spatially consistent determination ensures that noise information is applied only to speech features that match their physical properties, thus achieving precise local feature enhancement. Considering the continuous temporal variation of speech feature sequences, a neighborhood diffusion mechanism is introduced. After confirming the third speech feature, the adjacent positions in the first speech feature sequence before and after the third speech feature are examined. Since changes in action or pronunciation are usually accompanied by transitions, features in adjacent frames often carry crucial information about state transitions. Therefore, for first speech features that are immediately adjacent to the third speech feature in sequence position, even if their spatial orientation may have slight shifts or they are in a transitional state, they are still determined to be associated with the sub-noise feature. This temporal neighborhood-based association extension effectively avoids abrupt changes during noise application, ensuring a smooth temporal transition of the generated effect and preventing unnatural jumps.
[0138] In some embodiments, explicit exclusion conditions are set to avoid interference between irrelevant features. When it is detected that the virtual object portion corresponding to a certain first speech feature is completely different from the portion corresponding to a sub-noise feature, and the first speech feature is not adjacent to any third speech feature in the time series, it is determined that there is no correlation between the two. This logic ensures that features that are spatially non-overlapping and temporally unrelated remain independent, preventing erroneous noise features from being introduced into irrelevant control signals, thereby maintaining the purity and accuracy of the final generated virtual object representation.
[0139] In some embodiments, a first speech feature is identified as a third speech feature when it corresponds to the same part of a virtual object as a sub-noise feature. This means that if a speech feature and a sub-noise feature are associated with the same part of a virtual object (e.g., head, torso, etc.), then that feature is considered a third speech feature. Once a third speech feature is identified, it is considered associated with a sub-noise feature. This association may be based on the spatial or functional similarity between the speech feature and the sub-noise feature on the virtual object. A first speech feature is identified as associated with a sub-noise feature when it is adjacent to a third speech feature in a sequence of first speech features. This indicates that even if a speech feature does not directly correspond to a sub-noise feature, it may still be associated with a sub-noise feature if it is adjacent to a known third speech feature. This may be because adjacent speech features have some temporal and spatial correlation. A first speech feature is identified as not associated with a sub-noise feature when it corresponds to different parts of a virtual object and the first speech feature is not adjacent to a third speech feature. If a speech feature is located differently from a sub-noise feature on a virtual object and is not adjacent to any third speech feature, then it is not associated with the sub-noise feature.
[0140] As an example, suppose there is a first speech feature sequence containing multiple speech features that correspond to the mouth and head movements of a virtual character. For example, feature A corresponds to the opening of the mouth, feature B to the turning of the head, feature C to the closing of the mouth, and so on. If a first speech feature (e.g., feature A) is found to correspond to the same part of the virtual character (e.g., the mouth) as a sub-noise feature, then feature A is identified as the third speech feature. Once the third speech feature (feature A) is identified, it is considered to be associated with the sub-noise feature. This means that feature A may be affected by the sub-noise feature and needs to be processed. If feature A is adjacent to feature B in the first speech feature sequence, then feature B is also identified as associated with the sub-noise feature. This is because feature B and feature A are adjacent in the sequence and may have some correlation in time and space. If a first speech feature (e.g., feature C) corresponds to a different part of the virtual character (e.g., the head) as a sub-noise feature, and feature C is not adjacent to the third speech feature (feature A), then feature C is identified as not associated with the sub-noise feature. This means that feature C may not be affected by the sub-noise feature and can be left unchanged.
[0141] As an example, in the context of driving facial expressions in a virtual digital human, the virtual object can be a digital human model with facial bones or mixed shape weights. Different parts of the virtual object can include the mouth region, eye region, and eyebrow region. In this scenario, the first speech feature sequence is time-series data used to drive the digital human's facial expressions, and the sub-noise features can be noise data added to the mouth region to simulate subtle tremors or breathing sounds during real speech. During processing, the system first determines which facial region the current first speech feature primarily affects. If a first speech feature primarily controls the opening, closing, or deformation of the mouth region, then this first speech feature and the sub-noise feature targeting the mouth region correspond to the same part of the virtual object. At this point, the system marks the first speech feature controlling the mouth as the third speech feature and establishes an association with it, thereby superimposing the corresponding noise effect on the mouth during image synthesis. Next, for adjacent first speech features in the first speech feature sequence located in the frame before or after the third speech feature, considering the continuity and traction effect of facial muscle movements, the system directly determines that these adjacent first speech features are associated with sub-noise features, regardless of their main control area, to ensure a natural and smooth transition of mouth movements. Conversely, if there is a first speech feature in the sequence that mainly controls eye blinking, and the part corresponding to this feature is the eye region, which is different from the mouth region, and this feature is far away from the aforementioned third speech feature controlling the mouth in the time sequence, i.e., not adjacent, then the system determines that the first speech feature controlling the eyes is not associated with the sub-noise features of the mouth region, thereby avoiding interference from irrelevant mouth noise on eye movements and ensuring the accuracy of expression generation.
[0142] As an example, in the generation of body movements for 3D game characters, the virtual object can be a 3D character model in the game. Different parts of the virtual object can be divided into hand nodes, torso nodes, and leg nodes. In this scenario, the first speech feature sequence can be a sequence of motion feature vectors analyzed based on speech intonation to drive the character's body language. The sub-noise features can be random features used to simulate slight hand tremors in a natural state. When the analysis finds that a certain first speech feature in the sequence contains a clear hand gesture command such as waving or pointing, this feature corresponds to the hand node, and the part corresponding to the sub-noise feature is the same. The system then determines the third speech feature from this first speech feature and confirms that the two are related, so that the generated hand movements contain natural tremor details. Furthermore, for other first speech features adjacent to the waving gesture feature on the timeline, to prevent abruptness during action transitions, the system determines that these features are also associated with the hand sub-noise features based on the adjacency principle. However, when processing a feature in the sequence that represents the character's walking gait, this feature corresponds to the leg node, which is a different part from the hand node. If the walking feature is not immediately adjacent to the aforementioned hand waving feature in time, the system determines that the first speech feature is not related to the sub-noise feature of the hand, thereby ensuring that the stability of leg walking is not affected by random noise from the hand and achieving coordinated control of the whole body movement.
[0143] This allows for the effective identification and separation of speech features affected by noise, thereby improving the quality and clarity of the speech signal. It ensures that only speech features associated with the noise source are labeled and processed, while other irrelevant features remain unchanged, thus improving the accuracy and efficiency of speech processing.
[0144] In step 1032B, based on the mask, the first speech features in the first speech feature sequence are filtered to obtain the filtered first speech feature sequence.
[0145] In some embodiments, the above-mentioned filtering of the first speech features in the first speech feature sequence based on the mask to obtain the filtered first speech feature sequence can be achieved in the following way: for each first speech feature in the first speech feature sequence, when the value of the mask corresponding to the first speech feature is equal to a preset value (1), the first speech feature is determined as the target speech feature; the target speech features are combined to obtain the first speech feature sequence.
[0146] In some embodiments, there is a mask sequence that corresponds one-to-one with the first speech feature sequence. Each mask value indicates whether the corresponding speech feature is valid (1) or should be ignored (0). For each first speech feature in the first speech feature sequence, its corresponding mask value is checked. If the mask value is equal to 1, it means that the feature is valid and is identified as the target speech feature. All features identified as target speech features are combined to form a new first speech feature sequence. This new sequence contains only those speech features that are considered valid and unaffected by noise.
[0147] In some embodiments, during the process of filtering the first speech feature sequence based on a mask to obtain the filtered first speech feature sequence, this embodiment employs an element-by-element discrimination and extraction data processing mechanism. This mechanism aims to use a mask as a logic gate to accurately separate valid or key information segments from the original feature stream. The processing logic first traverses each first speech feature contained in the first speech feature sequence. For any first speech feature in the sequence, the value at the corresponding position in the mask is searched and obtained. In this context, a preset value is set as a specific identifier indicating validity or retention intent. The obtained mask value is compared with the preset value. When the comparison result shows that the mask value corresponding to the current first speech feature is equal to the preset value, this indicates that the feature carries the required valid information or meets specific processing conditions. Based on this determination, the first speech feature is established as a target speech feature and is marked or temporarily stored. After completing the traversal and discrimination of all features in the sequence, a combination operation is performed. This process involves aggregating all data units identified as target speech features. Typically, these target speech features are rearranged and concatenated according to their relative temporal or logical order in the original sequence to construct the first speech feature sequence after filtering. Through this method, the final sequence eliminates invalid features with mask values not equal to preset values, retaining only highly relevant target data. This simplifies feature dimensions and purifies information, providing a more accurate input foundation for subsequent data processing.
[0148] As an example, in an embodiment applied to speech signal preprocessing and denoising, the first speech feature sequence can be a Mel-frequency cepstral coefficient sequence obtained from the original audio data after frame segmentation, windowing, and feature extraction. The mask is a binary sequence generated by a speech activity detection algorithm to identify the presence of speech. In this scenario, the preset value is set to a value representing the presence of valid human voice in the current frame. During processing, the system iterates through and checks each first speech feature in the Mel-frequency cepstral coefficient sequence. When the mask value corresponding to a certain first speech feature is equal to the preset value, it indicates that the frame feature belongs to a valid segment containing human speech, rather than silence or pure background noise, and the system then identifies the first speech feature as the target speech feature. After the iteration is complete, the system reassembles all the identified target speech features in chronological order to obtain the filtered first speech feature sequence. In this way, the system successfully eliminates invalid silence frames and noise frames, providing cleaner input data for subsequent speech recognition models.
[0149] As an example, in an embodiment applied to high-precision lip-syncing for virtual digital humans, the first speech feature sequence can be a hybrid shape weight sequence predicted based on the input speech to control the facial mesh deformation of the virtual object, while the mask is a quality control sequence generated based on speech recognition confidence or phoneme clarity. In this scenario, the preset value represents a state of high confidence or available driving parameters. When the system executes the filtering logic, it compares the corresponding quality control mask with each first speech feature in the hybrid shape weight sequence in real time. When the value of the mask corresponding to the first speech feature is equal to the preset value, it means that the speech pronunciation at that moment is clear and the corresponding facial deformation parameters are accurate and reliable, and the system determines the first speech feature as the target speech feature. Subsequently, the system combines these confirmed target speech features to obtain the filtered first speech feature sequence. This process can effectively filter out abnormal driving parameters caused by unclear pronunciation, elision, or audio glitches, ensuring that the lip-syncing changes of the virtual digital human are smooth and natural when speaking, avoiding tremors or incorrect mouth movements.
[0150] Thus, by utilizing the matching relationship between the mask value and the preset value, target speech features with practical processing value can be accurately identified and extracted from the first speech feature sequence containing redundant, silent, or invalid information. This processing method not only effectively eliminates noise interference and non-key frames, significantly improving the data purity and effective information density of the feature sequence, but also provides a more concise and high-quality input basis for subsequent data processing or model inference by recombining the filtered target speech features. This reduces computational resource consumption while helping to improve the overall processing efficiency and accuracy of the output results.
[0151] In step 1033B, for each of the sub-noise features, the filtered first speech feature sequence and the sub-noise features are fused to obtain the first fused feature sequence.
[0152] In some embodiments, the above-mentioned fusion of the filtered first speech feature sequence and the sub-noise features to obtain the first fused feature sequence can be achieved by: fusing each first speech feature in the filtered first speech feature sequence with the filtered first speech feature sequence to obtain the fused features corresponding to each first speech feature; and combining the fused features to obtain the first fused feature sequence.
[0153] In some embodiments, the filtered first speech feature sequence is traversed to locate and acquire each first speech feature contained in the sequence. For any first speech feature in the sequence, the system treats it as a basic data unit to be processed, while simultaneously invoking sub-noise features as auxiliary information to be fused. Based on this, the system performs a fusion operation, combining the first speech feature with the sub-noise features. This fusion process can be manifested as feature vector concatenation, element-wise addition, or weighted operation based on specific weights, with the aim of enabling the current first speech feature to carry or associate with the background interference, environmental patterns, or channel characteristics represented by the sub-noise features. Through the above operations, the system calculates and obtains the fused features corresponding to the first speech feature, which expands the feature dimensions regarding the noise environment while retaining the original speech information.
[0154] In some embodiments, after traversing and fusing all feature units in the filtered first speech feature sequence, the system performs a sequence reconstruction operation. This process involves integrating the generated fusion features according to their temporal position or logical arrangement in the original sequence. By orderly combining these fusion features carrying noise information, a first fusion feature sequence is finally constructed. This sequence, as a composite feature stream that fuses speech content and noise attributes, can provide richer contextual information for subsequent neural network models, helping the model to learn the nonlinear relationship between speech and noise more accurately. As an example, the first speech feature sequence is filtered according to a mask sequence. For each first speech feature, if its corresponding mask value is 1, it is determined as a target speech feature. All target speech features are combined to form the filtered first speech feature sequence. This sequence only contains those speech features considered clean. The filtered first speech feature sequence is fused with noise features to obtain the first fusion feature sequence. Sub-noise features may be extracted from noise samples and are used to describe the characteristics of noise. The clean speech features and noise features are weighted and averaged or combined in other ways to generate a comprehensive feature sequence containing speech and noise information.
[0155] As an example, in an embodiment applied to training a noise-resistant speech recognition model, the filtered first speech feature sequence can be a sequence of Mel-spectrum feature frames extracted from clean speech data, while the sub-noise features are noise embedding vectors representing background noise patterns in a specific environment. In this scenario, to enable the model to learn speech feature changes under different noise environments, the system needs to construct training data containing noise information. Specifically, for each feature frame in the Mel-spectrum feature frame sequence, the system uses it as the first speech feature and concatenates or weights it with the noise embedding vector representing the current environmental interference, i.e., the sub-noise features. Through this fusion process, each generated fused feature not only retains the original speech spectrum structure but also carries background noise distribution information. Subsequently, the system rearranges and combines these fused features according to the original time step order to obtain the first fused feature sequence. This sequence is used as enhanced input data and fed into the neural network for training, effectively improving the robustness and accuracy of the speech recognition system in complex acoustic environments.
[0156] As an example, in an embodiment applied to a real-time speech enhancement and denoising system, the filtered first speech feature sequence can be a noisy speech spectral feature sequence retained after speech activity detection, while the sub-noise features are stationary noise statistical feature vectors estimated from non-speech segments. To assist the denoising network in more accurately removing noise components, the system performs feature fusion logic. Specifically, the system traverses the noisy speech spectral feature sequence, extracts each first speech feature, and performs channel-level concatenation and fusion with the sub-noise features characterizing the current noise distribution. This operation explicitly associates the speech features at each time point with corresponding noise prior information, thereby obtaining fused features corresponding to each first speech feature. Based on this, the system combines all fused features to generate a first fused feature sequence. This sequence is then input into the signal reconstruction network, which uses the embedded noise statistical features as a reference to more accurately subtract noise components from the mixed signal, thereby outputting a clean speech signal with a high signal-to-noise ratio. In this way, useful information in the speech signal can be effectively separated and preserved while reducing the impact of noise. This improves the quality and clarity of the speech signal and enhances the accuracy of speech recognition and understanding. The fused feature sequence not only contains clean speech features, but also retains some noise information.
[0157] In step 104, a driving parameter sequence matching the first speech feature sequence is determined based on the first fusion feature sequence.
[0158] In some embodiments, the driving parameter sequence includes multiple driving parameters, which are used to control at least one of the facial expressions and head postures of the virtual object. The driving parameter sequence is used to generate facial animation of the virtual object. The facial animation of the virtual object is a video of the virtual object's face. The meanings of facial video and facial animation are equivalent. This application embodiment uses facial animation as an example for explanation.
[0159] In some embodiments, driving parameters are used to control at least one of the facial expressions and head postures of the virtual object. For example, driving parameter A is used to control the facial expressions of the virtual object, driving parameter B is used to control the head posture of the virtual object, driving parameter A1 in driving parameter A is used to control the raising or lowering of the virtual object's eyebrows, driving parameter A2 in driving parameter A is used to control the opening and closing of the virtual object's lips, driving parameter A3 in driving parameter A is used to control the raising or lowering of the corners of the virtual object's mouth, and driving parameter A4 in driving parameter A is used to control the puffing or hollowing of the virtual object's cheeks.
[0160] In some embodiments, multiple driving parameters may control the facial expressions of the virtual object, multiple driving parameters may control the head posture of the virtual object, or one driving parameter may control the head posture of the virtual object and one driving parameter may control the facial expressions of the virtual object.
[0161] In some embodiments, a driving parameter sequence refers to a set of parameters arranged in chronological order, used to control the animation of virtual objects, particularly changes in facial expressions and head posture. Each driving parameter corresponds to a specific animation control point or deformation target of the virtual object, such as raising eyebrows, opening the mouth, or rotating the eyes. A driving parameter is a single element in the driving parameter sequence; it is a numerical value representing the state or position of a specific animation control point of the virtual object at a particular point in time.
[0162] As an example, one driving parameter might represent the degree to which eyebrows are raised, and another might represent the width of the mouth opening. Suppose there is a virtual character whose facial expressions and head posture need to be driven based on a speech signal. First, multiple first fusion feature sequences are obtained, containing both useful and noisy information from the speech signal. The driving parameter sequence might include the following parameters: Eyebrow raising parameter: controls the degree to which eyebrows are raised; Eye opening parameter: controls the degree to which eyes are opened; Mouth opening parameter: controls the degree to which the mouth is opened; Head rotation parameter: controls the angle at which the head turns left or right. These driving parameter sequences are sent to the virtual object's animation system, which adjusts the virtual object's facial expressions and head posture in real time based on these parameters to match the emotion and tone in the speech signal. For example, when the speech signal contains the emotion of anger, the driving parameter sequence might indicate expressions and postures such as furrowed eyebrows, a closed mouth, and a slightly forward-leaning head.
[0163] In some embodiments, referring to FIG6, FIG6 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application. Step 104 shown in FIG3 can be implemented by steps 1041A to 1042A shown in FIG6.
[0164] In step 1041A, according to the arrangement order of the fusion features in the first fusion feature sequence, the fusion features at the same position in the multiple first fusion feature sequences are fused to obtain a third fusion feature sequence.
[0165] In some embodiments, it is necessary to ensure that all first fusion feature sequences are synchronized in time or order. This means that features in each first fusion feature sequence are extracted at the same time point or in the same order so that they can be accurately matched and merged in subsequent fusion processes. For each time point or in-order position, features at the same position are extracted from all first fusion feature sequences. These features may come from different speech processing modules or algorithms, each capturing different aspects of the speech signal. Features at the same position can be fused using various methods, such as weighted averaging, maximum value selection, feature-level fusion (e.g., concatenating multiple feature vectors). Through the above fusion process, a new third fusion feature sequence is generated. This sequence integrates information from multiple first fusion feature sequences.
[0166] In some embodiments, the third fusion feature sequence refers to a single feature sequence obtained by fusing features at the same position in multiple first fusion feature sequences.
[0167] In some embodiments, the above-mentioned method of fusing fusion features at the same position in multiple first fusion feature sequences according to the arrangement order of each fusion feature in the first fusion feature sequence to obtain a third fusion feature sequence can be achieved by splicing fusion features at the same position in multiple first fusion feature sequences according to the arrangement order of each fusion feature in the first fusion feature sequence to obtain a third fusion feature sequence.
[0168] As an example, suppose there are three first fusion feature sequences, each containing three fusion features, representing facial expression features from three different video streams. These features can be numerical representations, such as eyebrow height, eye opening degree, and mouth opening degree. The three first fusion feature sequences are as follows: Sequence 1: [Eyebrow height 1, Eye opening degree 1, Mouth opening degree 1], Sequence 2: [Eyebrow height 2, Eye opening degree 2, Mouth opening degree 2]; Sequence 3: [Eyebrow height 3, Eye opening degree 3, Mouth opening degree 3]. Following the order of the fusion features in the first fusion feature sequences, fusion features at the same position in multiple first fusion feature sequences are fused. This means merging features at the same position in Sequence 1, Sequence 2, and Sequence 3. The eyebrow height feature is extracted from Sequence 1, Sequence 2, and Sequence 3. These three eyebrow height features are fused, for example, by taking the average or using a weighted average, to obtain a comprehensive eyebrow height feature. The eye opening degree feature is extracted from Sequence 1, Sequence 2, and Sequence 3. The three eye opening features are fused, for example, by averaging or using a weighted average, to obtain a comprehensive eye opening feature. Mouth opening features are extracted from sequences 1, 2, and 3. These three mouth opening features are fused, for example, by averaging or using a weighted average, to obtain a comprehensive mouth opening feature. A third fused feature sequence is obtained, which contains the fused eyebrow height, eye opening, and mouth opening features: Third fused feature sequence: [Comprehensive eyebrow height, comprehensive eye opening, comprehensive mouth opening].
[0169] In step 1042A, a driving parameter sequence that matches the first speech feature sequence is determined based on the third fusion feature sequence.
[0170] In some embodiments, the above determination can be achieved through a parameter determination model: by calling the parameter determination model, a driving parameter sequence matching the first speech feature sequence is determined based on the third fusion feature sequence.
[0171] As an example, suppose there is a virtual character whose facial expressions and head posture need to be driven by speech signals. First, multiple first fusion feature sequences are extracted using speech signal processing techniques. These sequences contain information from different aspects of the speech signal, such as acoustic features and semantic features. Next, fusion features at the same position in the multiple first fusion feature sequences are fused according to the order of their respective fusion features. For example, if there are two first fusion feature sequences, one extracted from acoustic features and the other from semantic features, the features from these two sequences are fused at each time point to obtain a third fusion feature sequence containing both acoustic and semantic information. Then, a machine learning model (such as a deep neural network) is used to learn the mapping relationship between the third fusion feature sequence and the driving parameter sequence. This model learns how to extract relevant information from the third fusion feature sequence and generate the corresponding driving parameter sequence using training data. Once the model is trained, a new third fusion feature sequence can be input into the model, and the model will output a defined driving parameter sequence. These driving parameter sequences contain parameters that control the virtual object's facial expressions and head posture, such as the degree of eyebrow raising, mouth opening, and eye rotation angle.
[0172] In this way, information from multiple first fusion feature sequences can be integrated to obtain a more comprehensive and accurate third fusion feature sequence. This sequence contains various information such as acoustics and semantics in the speech signal, and can better reflect the speaker's intention and emotion. Based on this third fusion feature sequence, a driving parameter sequence matching the first speech feature sequence can be determined, thereby achieving precise control over the virtual object's facial expressions and head posture.
[0173] In some embodiments, referring to FIG7, FIG7 is a schematic flowchart of the facial animation generation method provided in the present application embodiment. Step 104 shown in FIG3 can be implemented by steps 1041B to 1042B shown in FIG7.
[0174] In step 1041B, sub-driving parameter sequences are determined based on each of the first fusion feature sequences.
[0175] In some embodiments, the above-mentioned determination of sub-driving parameter sequences based on each of the first fusion feature sequences can be achieved in the following way: for each of the first fusion feature sequences, a parameter determination model is invoked, and the sub-driving parameter sequences corresponding to the first fusion feature sequences are determined based on the first fusion feature sequences.
[0176] In some embodiments, the parameter determination model includes a decoding layer. The above-mentioned parameter determination model, based on the first fused feature sequence, determines the sub-driving parameter sequence corresponding to the first fused feature sequence, which can be achieved by calling the decoding layer and determining the sub-driving parameter sequence corresponding to the first speech feature sequence based on the first fused feature sequence.
[0177] In some embodiments, each first fusion feature sequence may capture different aspects of the speech signal, such as acoustic features, semantic features, etc. By training the parameter determination model separately for each first fusion feature sequence, each model can focus more on specific types of information, thereby improving the accuracy of the determination. Different first fusion feature sequences may contain complementary information. By determining the sub-driving parameter sequences separately, it can be ensured that information from each aspect is fully considered, and a more comprehensive and accurate driving parameter sequence can be obtained when these sub-sequences are finally fused. Since the processing of each first fusion feature sequence is independent, parallel computing techniques can be used to accelerate the determination process and improve the efficiency of the system.
[0178] In some embodiments, a parameter inference process is performed for each independent first fused feature sequence. During this process, a pre-trained parameter determination model is invoked. This model constructs a nonlinear mapping relationship from a composite feature space containing speech and noise information to a target driving parameter space by learning from a large number of samples. When the first fused feature sequence is input into the parameter determination model, the computational unit within the model analyzes the input feature distribution, thereby outputting the sub-driving parameter sequence corresponding to the first fused feature sequence.
[0179] In some embodiments, particularly those involving the specific structure of the decoding layer, the logic for generating the aforementioned parameters is primarily implemented within the decoding layer. The first fused feature sequence is fed as input data to the decoding layer, which uses its network weights to perform dimensionality reduction and reconstruction on the high-dimensional fused features. This process aims to decouple and extract dynamic control information related to the original speech content from the fused features, while simultaneously utilizing the noise features fused within to adjust or optimize the robustness of parameter prediction. Through the operations of the decoding layer, a sub-driving parameter sequence corresponding to the first speech feature sequence in terms of time and content is finally derived from the first fused feature sequence, realizing the process of transforming acoustic features into specific parameter indicators for driving subsequent tasks. In step 1042B, sub-driving parameters at the same position in each of the sub-driving parameter sequences are combined to obtain a driving parameter sequence matching the first speech feature sequence.
[0180] As an example, consider a virtual character whose facial expressions and head posture need to be driven by speech signals. First, multiple first fusion feature sequences are extracted using speech signal processing techniques. These sequences contain information from different aspects of the speech signal, such as acoustic and semantic features. Based on each first fusion feature sequence, a parameter determination model is used to determine the resulting sub-driving parameter sequences. For example, if there are two first fusion feature sequences, one extracted from acoustic features and the other from semantic features, parameter determination models can be trained for each sequence to determine the corresponding sub-driving parameter sequences. Sub-driving parameters at the same positions in each sub-driving parameter sequence are combined to obtain a driving parameter sequence that matches the first speech feature sequence. For example, assuming that at a certain time point, the acoustic feature sub-driving parameter sequence is [0.8, 0.6, 0.4] and the semantic feature sub-driving parameter sequence is [0.7, 0.5, 0.3], then the corresponding parameters in these two sub-sequences can be combined to obtain the driving parameter sequence [0.75, 0.55, 0.35].
[0181] In this way, by utilizing the unique information of each first fused feature sequence, the corresponding sub-driving parameter sequence is determined. By combining the parameters at the same positions in these sub-sequences, a comprehensive and accurate driving parameter sequence can be obtained. This sequence can better match the original speech feature sequence, thereby achieving precise control over the facial expressions and head posture of the virtual object. This not only improves the accuracy of driving parameter determination but also enhances the expressive ability of the virtual character, bringing users a more realistic and immersive experience.
[0182] In some embodiments, referring to FIG8, FIG8 is a schematic flowchart of the facial animation generation method provided in the embodiments of this application. Step 104 shown in FIG3 can be implemented by steps 1041C to 1043C shown in FIG8.
[0183] In step 1041C, based on multiple first fusion feature sequences, the first driving parameters are determined to obtain the first driving parameter sequence.
[0184] In some embodiments, the determination of the first driving parameters based on multiple first fusion feature sequences to obtain the first driving parameter sequence can be achieved by calling a parameter determination model to determine the first driving parameters based on multiple first fusion feature sequences to obtain the first driving parameter sequence.
[0185] In some embodiments, the parameter determination model includes a decoding layer, and the step of invoking the parameter determination model to perform the first determination of driving parameters based on the plurality of first fused feature sequences to obtain the first driving parameter sequence may be implemented as follows: invoking the decoding layer of the parameter determination model to perform the first determination of driving parameters based on the plurality of first fused feature sequences to obtain the first driving parameter sequence.
[0186] In some embodiments, the parameter determination model, such as a deep neural network, has strong learning capability and can learn complex mapping relationships from a plurality of first fused feature sequences. Through training, the model can learn the association between different feature sequences, so as to determine driving parameters more accurately. One-time determination can reduce calculation steps and improve efficiency. The model directly learns the driving parameter sequence from the plurality of first fused feature sequences, avoiding the complexity of multiple determinations and combinations. The first driving parameter sequence obtained through one-time determination is globally consistent, because the model considers the mutual relationship of all first fused feature sequences during the learning process. This helps ensure the coherence and naturalness of the driving parameter sequence.
[0187] In step 1042C, based on the i-th driving parameter sequence, the (i+1)-th determination of driving parameters is performed to obtain the (i+1)-th driving parameter sequence.
[0188] In some embodiments, 1≤i<N, N is a positive integer greater than 1, the i-th driving parameter sequence includes a plurality of driving parameter combinations, and each driving parameter combination corresponds to one animation frame in the facial animation.
[0189] In some embodiments, the step of performing the (i+1)-th determination of driving parameters based on the i-th driving parameter sequence to obtain the (i+1)-th driving parameter sequence may be implemented as follows: invoking the parameter determination model to perform the (i+1)-th determination of driving parameters based on the i-th driving parameter sequence to obtain the (i+1)-th driving parameter sequence.
[0190] In some embodiments, the parameter determination model includes a decoding layer, and the step of invoking the parameter determination model to perform the (i+1)-th determination of driving parameters based on the i-th driving parameter sequence to obtain the (i+1)-th driving parameter sequence may be implemented as follows: invoking the decoding layer of the parameter determination model to perform the (i+1)-th determination of driving parameters based on the i-th driving parameter sequence to obtain the (i+1)-th driving parameter sequence.
[0191] In some embodiments, driving parameters are variables that control facial animation; they can represent eyebrow height, eye opening, mouth shape, head position, etc. A driving parameter combination is a set of multiple driving parameters that act simultaneously within an animation frame, collectively defining the virtual character's facial expression and head posture in that frame. An animation frame is a static image in an animation; by playing multiple animation frames consecutively, dynamic animation effects can be created. Each animation frame corresponds to a driving parameter combination. Based on the i-th driving parameter sequence, the (i+1)-th driving parameter is determined, resulting in the (i+1)-th driving parameter sequence. This process involves using a model or algorithm to determine the next driving parameter sequence based on the current one.
[0192] In some embodiments, step 1042C can be implemented as follows: for each of the driving parameter combinations in the i-th driving parameter sequence, feature extraction is performed on the driving parameter combination to obtain parameter features; the first speech feature sequence is fused with each of the parameter features to obtain multiple i-th parameter feature sequences; based on the multiple i-th parameter feature sequences, the driving parameters are determined for the (i+1)-th time to obtain the (i+1)-th driving parameter sequence.
[0193] In some embodiments, the above-mentioned determination of the driving parameters for the (i+1)th time based on the plurality of i-th parameter feature sequences to obtain the (i+1)-th driving parameter sequence can be achieved by calling a parameter determination model, determining the driving parameters for the (i+1)th time based on the plurality of i-th parameter feature sequences to obtain the (i+1)-th driving parameter sequence.
[0194] In some embodiments, the parameter determination model includes a decoding layer. The above-mentioned call to the parameter prediction model, based on multiple i-th parameter feature sequences, performs the determination of the i+1-th driving parameter to obtain the i+1-th driving parameter sequence, can be achieved as follows: call the decoding layer of the parameter prediction model, based on multiple i-th parameter feature sequences, perform the determination of the i+1-th driving parameter to obtain the i+1-th driving parameter sequence.
[0195] In some embodiments, feature extraction is performed on each combination of driving parameters in the i-th driving parameter sequence to obtain parameter features. These features may include the relationships between driving parameters, their changing trends, and their correspondence with the speech signal. The first speech feature sequence is fused with each parameter feature to obtain multiple i-th parameter feature sequences. This fusion can be a simple concatenation or implemented using a specific fusion algorithm (such as weighted averaging, feature-level fusion, etc.). Based on the multiple i-th parameter feature sequences, the driving parameters are determined for the (i+1)-th time to obtain the (i+1)-th driving parameter sequence. Typically, a machine learning model (such as a recurrent neural network, long short-term memory network, etc.) is used to learn the mapping relationship between the first fused feature sequence and the driving parameter sequence. It is iterative, meaning that each determination is based on the previous driving parameter sequence and speech feature sequence. In this way, a series of driving parameter sequences can be gradually generated, corresponding to multiple animation frames in the animation.
[0196] As an example, suppose there is a virtual character whose facial expressions and head posture need to be driven by speech signals. First, a first speech feature sequence is extracted using speech signal processing techniques. This sequence contains acoustic features, semantic features, and other information from the speech signal. Next, feature extraction is performed on each combination of driving parameters in the i-th driving parameter sequence to obtain parameter features. For example, if there are three combinations of driving parameters, each corresponding to one of three animation frames, parameter features can be extracted for each of these three combinations. These features may include the relationships and trends between the driving parameters. Then, the first speech feature sequence is fused with each parameter feature to obtain multiple i-th parameter feature sequences. For example, if there are three parameter features, three i-th parameter feature sequences can be obtained, each containing speech features and corresponding parameter features. Next, based on the multiple i-th parameter feature sequences, the (i+1)-th driving parameters are determined to obtain the (i+1)-th driving parameter sequence. This step typically involves using machine learning models (such as recurrent neural networks, long short-term memory networks, etc.) to learn the mapping relationship between the first fused feature sequence and the driving parameter sequence. By learning these relationships, the model can determine the next driving parameter sequence. Finally, the determined sequence of driving parameters is sent to the animation system of the virtual object. The system then adjusts the virtual character's facial expressions and head posture in real time based on these parameters to match the emotion and tone of voice in the speech signal. This allows for the creation of a natural and realistic virtual character animation.
[0197] Thus, by fusing speech feature sequences with driving parameter features, the information in the speech signal can be fully utilized, improving the accuracy of driving parameter determination. The iterative determination process optimizes each determination based on the previous result, contributing to improved robustness and accuracy of the entire system. It can dynamically adapt to changes in the speech signal and adjust the driving parameter sequence in real time, enabling facial animation to accurately reflect the speaker's expressions and emotions. By extracting and fusing features for each combination of driving parameters, the efficiency of driving parameter determination can be improved, reducing computational load. This can enhance the expressiveness of virtual characters, making them more natural and realistic, and providing a better user experience.
[0198] In step 1043C, i is traversed, and the Nth driving parameter sequence obtained by traversing i is determined as the driving parameter sequence that matches the first speech feature sequence.
[0199] In some embodiments, the process iterates through parameter i, and the Nth driving parameter sequence obtained from iterating through i is determined as the driving parameter sequence that matches the first speech feature sequence. This means that a series of driving parameter sequences are gradually generated through an iterative process. Specifically, starting from the first determination, the first driving parameter sequence is obtained; then, based on the first driving parameter sequence, the second determination is performed to obtain the second driving parameter sequence; and so on, until the Nth determination is performed to obtain the Nth driving parameter sequence. This process is similar to a loop, where each determination depends on the result of the previous one.
[0200] In some embodiments, a progressive generation or optimization mechanism based on multiple iterations is employed. This mechanism aims to gradually approximate or construct high-precision target parameters through staged calculations. The processing logic first relies on the multiple first fusion feature sequences generated above as the initial information input basis. The system performs initial parameter inference operations based on these first fusion feature sequences to calculate the driving parameter sequence corresponding to the initial stage. This initial driving parameter sequence serves as the starting point of the entire generation chain, providing the basic data form for subsequent fine-tuning. The system follows a recursive logic of deriving the next state from the previous state. In each iteration step, the driving parameter sequence generated in the current iteration round is used as a reference or input data, and the next stage of calculation is performed through a preset algorithm model to derive and generate the driving parameter sequence corresponding to the next iteration round. This process can be understood as a gradual correction, denoising, feature enhancement, or autoregressive generation of the parameter sequence, ensuring that the parameter sequence after each iteration is better than the previous result in terms of accuracy or expressiveness. The above recursive operation is strictly executed according to the preset total number of iterations, and the iteration process is completely traversed. When the number of iterations reaches a preset final threshold, the system stops looping and confirms the driving parameter sequence output by the last iteration as the final result. The final result is identified as the driving parameter sequence that precisely matches the first speech feature sequence in terms of timing, prosody, and content, and is used to support subsequent speech-driven tasks.
[0201] As an example, suppose there is a virtual character whose facial expressions and head posture need to be driven by speech signals. First, multiple first fusion feature sequences are extracted using speech signal processing techniques. These sequences contain information from different aspects of the speech signal, such as acoustic features and semantic features. Next, based on these first fusion feature sequences, the first driving parameters are determined, resulting in the first driving parameter sequence. This step typically involves using machine learning models (such as deep neural networks, support vector machines, etc.) to learn the mapping relationship between the first fusion feature sequences and the driving parameter sequences. By learning these relationships, the model can determine a driving parameter sequence containing parameters that control the virtual character's facial expressions and head posture. Then, based on the i-th driving parameter sequence, the (i+1)-th driving parameters are determined, resulting in the (i+1)-th driving parameter sequence. This step is iterative, meaning each determination is based on the previous driving parameter sequence. In this way, a series of driving parameter sequences can be gradually generated, corresponding to multiple animation frames in an animation. Finally, i is traversed, and the N-th driving parameter sequence obtained from traversing i is determined as the driving parameter sequence that matches the first speech feature sequence. This sequence contains all the necessary parameters to control the virtual character's facial expressions and head posture, and these parameters are consistent and coordinated as a whole.
[0202] As an example, in an embodiment applied to driving facial expressions in a virtual digital human, facial motion data is synthesized based on the generation logic of a diffusion model. In this scenario, multiple first fused feature sequences are composite features fused from input speech audio features and random Gaussian noise vectors at different time steps. The system first inputs these fused features containing noise information into a pre-trained denoising network, and performs preliminary inference through a reverse diffusion process to generate an initial noisy facial deformation coefficient sequence, i.e., the first driving parameter sequence. Subsequently, the system performs a cyclic iterative denoising operation, using the facial deformation coefficient sequence generated in the current iteration as a basis, and combining it with corresponding conditional features to calculate and derive the next iteration's facial deformation coefficient sequence with fewer noise components and clearer expression details, i.e., the (i+1)th driving parameter sequence. This process is repeated continuously until a preset number of iterations are completed. The final Nth driving parameter sequence is identified as a high-fidelity facial expression driving parameter sequence that is strictly aligned with the input speech. This sequence is transmitted to the rendering engine to drive the lip movements and facial muscle movements of the virtual digital human, achieving realistic speech animation synthesis.
[0203] As an example, in an embodiment applied to the training and inference of a high-quality speech synthesis acoustic model, to generate refined acoustic feature parameters, multiple first fused feature sequences are data fused from linguistic features obtained by text encoding and noise distribution features representing different sampling stages. The system initiates the acoustic parameter generation process based on these fused features, first predicting an initial Mel-spectrum sequence that can characterize the general outline of the speech but contains random perturbations; this is the first driving parameter sequence. Based on this, the system uses an iterative refinement algorithm to progressively reconstruct the spectrum. In each iteration step, the system corrects the frequency domain texture and restores the harmonic structure based on the Mel-spectrum sequence output from the previous round, thereby determining the next round Mel-spectrum sequence after quality improvement, i.e., the (i+1)th driving parameter sequence. After a complete traversal of all iteration steps i, the final Nth driving parameter sequence obtained by the system is the target acoustic parameter sequence. This driving parameter sequence is then input into the vocoder module to drive the vocoder to synthesize a speech signal with continuous waveform, clear sound quality, and rhythmic quality. Thus, the first driving parameter sequence is obtained by determining the first driving parameter sequence based on multiple first fused feature sequences. Then, the (i+1)th driving parameter sequence is determined based on the i-th driving parameter sequence, resulting in the (i+1)-th driving parameter sequence. The process iterates through i, and the N-th driving parameter sequence obtained from iterating through i is determined as the driving parameter sequence that matches the first speech feature sequence. This improves the accuracy of driving parameter determination. Iterating through i and determining the N-th driving parameter sequence that matches the first speech feature sequence is not merely a simple iterative determination, but a gradual denoising and optimization process. Through N iterations, noise and inaccuracies in the driving parameter sequence can be effectively reduced, making the final N-th driving parameter sequence more accurately reflect the original speech feature sequence.
[0204] In other embodiments, the first fusion feature sequence includes M fusion features, where M is an integer greater than 2; step 104 above can also be implemented as follows: based on the first fusion feature in the first fusion feature sequence, determine a first driving parameter that matches the first first speech feature in the first speech feature sequence; based on the second fusion feature in the first fusion feature sequence and the first driving parameter, determine a second driving parameter that matches the second first speech feature in the first speech feature sequence; based on the j-th fusion feature in the first fusion feature sequence and the first to (j-1)-th driving parameters, determine a j-th driving parameter that matches the j-th first speech feature in the first speech feature sequence, where j is an integer greater than 2; iterate through j until j equals M to obtain a sequence of driving parameters that match the first speech feature sequence.
[0205] In some embodiments, the processing logic begins execution from the start of the sequence. The first fusion feature at the beginning of the first fusion feature sequence is extracted as input, and a first driving parameter that matches the first first speech feature in the first speech feature sequence in terms of both timing and content is calculated and derived. This first driving parameter serves not only as the current output but also as the starting state information for subsequent calculations.
[0206] In some embodiments, the processing focus is shifted to the second time step of the sequence. At this stage, parameter determination no longer relies solely on the current input features, but incorporates the output from the previous time step as a condition. Combining the second fusion feature in the first fusion feature sequence with the first driving parameter determined in the preceding steps, both pieces of information are used to determine the second driving parameter that matches the second first speech feature in the first speech feature sequence. The logic of using historical output to assist the current prediction is extended to each subsequent time step. For any j-th position in the sequence, where j represents an integer greater than 2, the following operation is performed: The j-th fusion feature in the first fusion feature sequence is obtained, and all previously generated historical driving parameters are invoked, i.e., the set of parameters from the first driving parameter up to the j-th minus 1 driving parameter. These historical driving parameters are considered as contextual information capturing the dynamic evolution of speech, and are input into the computational model along with the current j-th fusion feature, thereby accurately determining the j-th driving parameter that matches the j-th first speech feature in the first speech feature sequence. The variable j is iterated in ascending order of time, and the steps of generating the current parameter based on the current feature and historical parameters are repeated until the value of variable j increases to equal the total length M of the sequence. After the entire traversal process is completed, all the driving parameters generated in sequence are integrated in chronological order to obtain a complete driving parameter sequence that matches the full length of the first speech feature sequence. As an example, in a specific embodiment applied to the precise driving of virtual digital human mouth shapes, the first fused feature sequence is specified as a composite audio feature sequence containing M temporal frames, where M represents the total number of frames of the driving data to be generated. In this scenario, the driving parameter sequence corresponds to the lip shape deformation coefficient sequence that controls the mouth and facial muscle movement of the virtual digital human. The generation process of this sequence follows a temporal autoregressive dependency logic to ensure that the generated lip shape animation has the continuity of physical movement in the time dimension. The processing flow starts at the first time step of the sequence. Based on the first fused feature in the first fused feature sequence, the speech pronunciation content and the corresponding potential lip shape at that moment are analyzed to infer and determine the first driving parameter controlling the initial state of the virtual digital human. This parameter defines the facial geometry at the moment of speech initiation.
[0207] Continuing from the previous example, we proceed to the second time step. To prevent abrupt changes in lip movements and maintain smoothness, we consider not only the second fusion feature in the first fusion feature sequence but also the first driving parameter generated in the previous time step. By combining the speech information at the current time step with the lip movement state at the previous time step, we calculate the second driving parameter that matches the second first speech feature, thus achieving a reasonable transition from the initial state to the second state. For any subsequent j-th time step, where j is an integer greater than 2, we perform a prediction operation based on full historical information. We call the current j-th fusion feature in the first fusion feature sequence and simultaneously obtain all previously generated historical parameters, i.e., the parameter set from the first driving parameter up to the j-1 driving parameter. These historical driving parameters are used as conditional information representing the lip movement trajectory and contextual dependencies, and are input into the prediction model along with the current fusion feature. Based on this, the model analyzes the most reasonable lip shape at the current time step, given all past motion states and current speech content, and then determines the j-th driving parameter that matches the j-th first speech feature. The variable j is iterated sequentially over time, and the steps of deriving the current state based on historical information are repeated continuously until the value of j increases to equal the total sequence length M. Once the entire iteration is complete, the sequentially generated driving parameters constitute a complete sequence, which is the driving parameter sequence matching the first speech feature sequence. This sequence is then transmitted to the rendering engine, driving the virtual digital human to present a speech effect that is perfectly synchronized with the speech rhythm and dynamically coherent.
[0208] In some embodiments, the above-mentioned determination of the first driving parameter matching the first first speech feature in the first speech feature sequence based on the first fusion feature in the first fusion feature sequence can be achieved by calling the decoding layer of the parameter determination model and determining the first driving parameter matching the first first speech feature in the first speech feature sequence based on the first fusion feature in the first fusion feature sequence.
[0209] In some embodiments, the above-mentioned determination of a second driving parameter matching the second first speech feature in the first speech feature sequence based on the second fusion feature in the first fusion feature sequence and the first driving parameter can be achieved by: calling the decoding layer of the parameter determination model, and determining a second driving parameter matching the second first speech feature in the first speech feature sequence based on the second fusion feature in the first fusion feature sequence and the first driving parameter.
[0210] In some embodiments, the j-th driving parameter that matches the j-th first speech feature in the first speech feature sequence is determined based on the j-th fusion feature in the first fusion feature sequence and the first to (j-1)-th driving parameters. This can be achieved by calling the decoding layer of the parameter determination model and determining the j-th driving parameter that matches the j-th first speech feature in the first speech feature sequence based on the j-th fusion feature in the first fusion feature sequence and the first to (j-1)-th driving parameters.
[0211] In some embodiments, for processing the start position of the sequence, the system performs calculations by calling parameters to determine the decoding layer configured within the model. Specifically, the process involves inputting the first fused feature in the first fused feature sequence into this decoding layer. The decoding layer parses and maps the input feature information to calculate and determine the first driving parameter that matches the first first speech feature in the first speech feature sequence. The first driving parameter determined in this step not only serves as the current output but also establishes the initial state basis for subsequent sequence generation.
[0212] In some embodiments, the subsequent parameter determination process for the second position also relies on calling the decoding layer of the parameter determination model. At this stage, the input information to the decoding layer not only includes the second fused feature in the first fused feature sequence but also explicitly includes the first driving parameter generated in the previous processing step. The decoding layer comprehensively processes the current feature input with the historical output from the previous moment to determine the second driving parameter that matches the second first speech feature in the first speech feature sequence, thereby ensuring the temporal continuity of parameter generation.
[0213] In some embodiments, a recursive processing mechanism based on historical information is extended to the parameter generation at any subsequent j-th time step. Specifically, the system continues to call the decoding layer of the parameter determination model, taking the j-th fusion feature in the first fusion feature sequence as the input at the current time, and simultaneously inputting the first driving parameter up to the j-th driving parameter accumulated in previous steps as historical context information. The decoding layer infers based on the joint distribution of the current feature and the historical sequence, thereby accurately determining the j-th driving parameter that matches the j-th first speech feature in the first speech feature sequence. In some embodiments, the driving parameter sequence is determined step by step by utilizing each fusion feature in the first fusion feature sequence and the previously determined driving parameters to determine the driving parameter that matches the corresponding first speech feature in the first speech feature sequence. Instead of determining the entire driving parameter sequence at once, each driving parameter is determined step by step. This helps reduce the complexity of the determination and improve the accuracy of the determination. When determining the j-th driving parameter, not only the j-th fusion feature is used, but also the previously determined first to j-1 driving parameters are used. Historical information can be utilized to improve the coherence and consistency of the determination. It is determined gradually, allowing it to better adapt to changes in the speech signal. When the speech signal changes, more accurate driving parameters can be determined by adjusting the current fusion features and previous driving parameters. It can be adjusted based on different first fusion feature sequences and first speech feature sequences, offering high flexibility.
[0214] As an example, suppose there is a virtual character whose facial expressions and head posture need to be driven by speech signals. First, a first speech feature sequence is extracted using speech signal processing techniques. This sequence contains acoustic features, semantic features, and other information from the speech signal. Next, feature extraction is performed on each combination of driving parameters in the i-th driving parameter sequence to obtain parameter features. Then, the first speech feature sequence is fused with each parameter feature to obtain multiple i-th parameter feature sequences. Each first fused feature sequence contains M fused features, where M is an integer of 2. Based on the first fused feature in the first fused feature sequence, a first parameter matching the first first speech feature in the first speech feature sequence is determined. Machine learning models (such as deep neural networks, support vector machines, etc.) are typically used to learn the mapping relationship between fused features and driving parameters. Based on the second fused feature in the first fused feature sequence and the first driving parameter, a second driving parameter matching the second first speech feature in the first speech feature sequence is determined. This step uses not only the second fused feature but also the previously determined first driving parameter to improve the coherence and consistency of the determination. Based on the j-th fusion feature and the first to (j-1)-th driving parameters in the first fusion feature sequence, the j-th driving parameter that matches the j-th first speech feature in the first speech feature sequence is determined. This step uses the j-th fusion feature and the previously determined first to (j-1)-th driving parameters to improve the accuracy of the determination. j is iterated until j equals M, resulting in a sequence of driving parameters that match the first speech feature sequence. This sequence contains all the necessary parameters for controlling the virtual character's facial expressions and head posture, and these parameters are consistent and coordinated overall.
[0215] Thus, by progressively determining the driving parameters, each driving parameter can be optimized step by step, improving the accuracy of the entire driving parameter sequence. When determining the j-th driving parameter, not only the j-th fusion feature is used, but also the previously determined driving parameters from 1 to (j-1). Historical information can be utilized, improving the coherence and consistency of the determination. Because it is determined progressively, it can better adapt to changes in the speech signal. When the speech signal changes, more accurate driving parameters can be determined by adjusting the current fusion feature and previous driving parameters. Adjustments can be made based on different first fusion feature sequences and first speech feature sequences, providing high flexibility.
[0216] In some embodiments, the driving parameter sequence is determined by a parameter determination model. Before fusing the sub-noise features with the first speech feature sequence to obtain the first fused feature sequence, the parameter determination model can be trained as follows: Acquire speech feature sequence samples, which include multiple first speech feature samples, the labels of which include sub-labels corresponding one-to-one with each first speech feature sample; invoke an initial parameter determination model to determine the driving parameter sequence corresponding to the speech feature sequence samples based on the speech feature sequence samples; determine a first loss value based on the labels of the speech feature sequence samples and the driving parameter sequence corresponding to the speech feature sequence samples; determine a second loss value based on the sub-labels of each first speech feature sample and each driving parameter in the driving parameter sequence corresponding to the speech feature sequence samples; update the model parameters of the initial parameter determination model based on the first loss value and the second loss value to obtain the parameter determination model.
[0217] In some embodiments, the label includes a sub-label for driving the head of the virtual object to rotate and translate, and a sub-label for driving the various facial parts of the virtual object to move.
[0218] In some embodiments, in the implementation scenario of constructing a parameter determination model for driving virtual digital lip shapes and facial expressions, the parameter determination model needs to be trained before performing the operation of fusing sub-noise features with a first speech feature sequence. This training process aims to enable the model to learn the mapping relationship from speech features to facial driving parameters.
[0219] In some embodiments, a sequence of speech feature samples for training is first acquired. This sequence comprises multiple first speech feature samples arranged in chronological order, such as consecutive speech spectrum frames. Simultaneously, sub-labels corresponding one-to-one with each first speech feature sample are acquired. These sub-labels constitute the ground truth of the training data, such as real facial blending deformation coefficients or skeletal keypoint data corresponding to each frame of speech acquired through facial capture technology.
[0220] In some embodiments, during the training execution phase, an initial parameter determination model that has not yet been fully trained is invoked, and the acquired speech feature sequence samples are input into the model. The model performs internal calculations and outputs a sequence of predicted driving parameters corresponding to the input samples. To comprehensively evaluate the model's predictive performance, a dual loss calculation mechanism is employed. On one hand, a first loss value is determined by comparing the overall label information of the speech feature sequence samples with the complete driving parameter sequence output by the model. This first loss value primarily measures the difference between the predicted sequence and the true sequence in terms of overall temporal structure and dynamic trends. On the other hand, the focus is on specific units within the sequence. A second loss value is determined by comparing the sub-labels of each first speech feature sample with the corresponding driving parameters in the driving parameter sequence one by one. This second loss value quantifies the accuracy of the prediction result for each frame.
[0221] In some embodiments, a total loss function is constructed based on the first and second loss values calculated above, and the model parameters of the initial parameter determination model are iteratively updated using the backpropagation algorithm. By continuously repeating the above forward prediction and parameter update steps until the model converges, a parameter determination model capable of accurately generating driving parameters is obtained. In some embodiments, the driving parameter sequence is implemented through a parameter determination model, which includes a self-attention layer, a cross-attention layer, and a decoding layer. The above-mentioned feature extraction of preset noise to obtain noise features can be achieved by calling the self-attention layer to extract noise features from the preset noise. The above-mentioned fusion of the sub-noise features with the first speech feature sequence to obtain a first fused feature sequence can be achieved by calling the cross-attention layer to fuse the sub-noise features with the first speech feature sequence to obtain a first fused feature sequence. The above-mentioned determination of a driving parameter sequence matching the first speech feature sequence based on multiple first fused feature sequences can be achieved by calling the decoding layer to determine a driving parameter sequence matching the first speech feature sequence based on multiple first fused feature sequences.
[0222] In some embodiments, the generation of the driving parameter sequence depends on the computational processing of a parameter determination model, which specifically includes a self-attention layer, a cross-attention layer, and a decoding layer in its network architecture. For the feature extraction process of the preset noise, the system implements this by calling the self-attention layer in the parameter determination model. The self-attention layer receives the preset noise as input data, calculates the correlation dependencies between elements within the input data, captures the inherent structural features of the data, and thus outputs the corresponding noise features.
[0223] In some embodiments, the process of fusing sub-noise features with a first speech feature sequence to generate a first fused feature sequence is accomplished by calling a cross-attention layer in the parameter determination model. The cross-attention layer receives sub-noise features and the first speech feature sequence as input, and uses an attention mechanism to align and interactively compute the two, thereby effectively fusing noise-related feature information into the speech feature representation to obtain a first fused feature sequence containing mixed information.
[0224] In some embodiments, the process of determining the final driving parameter sequence based on multiple first fused feature sequences is implemented by calling the decoding layer in the parameter determination model. The decoding layer parses the multiple input first fused feature sequences, maps the fused high-dimensional features to the target driving parameter space through decoding operations, and thus determines the driving parameter sequence that matches the first speech feature sequence in terms of time and content. As an example, referring to Figure 10, the parameter determination model shown in Figure 10 includes a self-attention layer, a cross-attention layer, and a decoding layer. The self-attention layer is called to extract noise features from preset noise. The cross-attention layer is called to fuse the sub-noise features with the first speech feature sequence to obtain the first fused feature sequence. The decoding layer is called to determine the driving parameter sequence that matches the first speech feature sequence based on the multiple first fused feature sequences.
[0225] In some embodiments, a self-attention layer is a mechanism that enables a model to focus on the relationships between elements at different positions in a sequence when processing sequential data. It allows the model to dynamically assign different weights to other elements in the sequence when processing an element, thereby capturing long-range dependencies. Self-attention layers are typically used in the encoder part to help the model understand the contextual information of the input sequence.
[0226] In some embodiments, a cross-attention layer is a mechanism that allows the decoder to focus on the hidden states of the encoder while generating the output. This mechanism enables the decoder to generate more relevant and accurate outputs based on information from the input sequence. Cross-attention layers are particularly important in sequence-to-sequence tasks such as machine translation and speech synthesis.
[0227] In some embodiments, the decoding layer is the part of the transformer model used to generate the output sequence. It typically consists of multiple decoder blocks, each containing a self-attention layer and a cross-attention layer. The task of the decoding layer is to progressively generate each element of the target sequence based on information provided by the encoder and previously generated output elements. During the generation process, the decoding layer can utilize the self-attention layer to process the contextual information of the generated output sequence and the cross-attention layer to reference information from the input sequence.
[0228] In some embodiments, a large number of speech feature sequence samples need to be collected. These samples include multiple first speech feature samples, each with a corresponding label. The label includes sub-labels that correspond one-to-one with each first speech feature sample. These sub-labels may be the ground truth values of the driving parameters or other information related to the speech features. An initial parameter set is used to determine the model. This model may be pre-trained or randomly initialized. The model structure can be designed according to the specific task; common structures include recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and convolutional neural networks (CNNs). Using the initial parameter set, the driving parameter sequence corresponding to the speech feature sequence samples is determined based on the speech feature sequence samples. This determination process may involve encoding the speech feature sequences and then decoding the driving parameter sequence. Based on the labels of the speech feature sequence samples and the determined driving parameter sequence, a first loss value is determined. This loss value reflects the difference between the determined driving parameter sequence and the ground truth; common loss functions include mean squared error (MSE) and cross-entropy loss. Based on the sub-labels of each first speech feature sample and each driving parameter in the determined driving parameter sequence, a second loss value is determined. This loss value may reflect a specific relationship between the driving parameters and the sub-labels; for example, certain driving parameters should be associated with specific speech features. Based on the first and second loss values, the model parameters are updated according to the initial parameters. Optimization algorithms, such as stochastic gradient descent (SGD) and Adam, are typically used to minimize the loss function and optimize the model parameters.
[0229] Thus, a parameter determination model comprising a self-attention layer, a cross-attention layer, and a decoding layer was adopted. This hierarchical network architecture significantly improved the accuracy and diversity of the generated driving parameters. Specifically, the self-attention layer extracts features from the preset noise, fully capturing the global dependencies and potential distribution features within the noise data, thereby providing richly expressive noise features for subsequent processing. By invoking the cross-attention layer to fuse the sub-noise features with the first speech feature sequence, deep interaction and temporal alignment of cross-modal features are achieved, ensuring that the generated intermediate features not only respond to the changing patterns of the speech signal but also retain the randomness and diversity introduced by the noise. Finally, the decoding layer parses the first fused feature sequence based on multiple fused features, accurately mapping a driving parameter sequence that highly matches the speech content. This attention-based feature processing method effectively enhances the model's ability to model long sequence data, ensuring that the lip movements and facial expressions of the virtual object during pronunciation are both accurate, synchronized, and vividly natural.
[0230] In some embodiments, the acquisition of speech feature sequence samples described above can be achieved as follows: extracting a video frame sequence and an audio frame sequence of the target object from a video containing the target object; performing feature extraction on the audio frame sequence to obtain the speech feature sequence samples; identifying driving parameters for each video frame in the video frame sequence to obtain driving parameters corresponding to each video frame; and constructing labels for the speech feature sequence samples based on the driving parameters corresponding to each video frame.
[0231] In some embodiments, the processing logic first parses video data containing the target object. A sequence of video frames containing visual information of the target object and a sequence of audio frames containing acoustic information of the target object are separated and extracted from the video. Based on this, feature extraction operations are performed on the extracted audio frame sequence to convert the original audio signal into a vectorized representation, thereby obtaining speech feature sequence samples. Simultaneously, image analysis and parameter recognition are performed on each video frame in the video frame sequence, determining the driving parameters corresponding to the content of each video frame. Finally, the driving parameters identified from each video frame are used as ground truth data to construct and generate labels corresponding to the speech feature sequence samples, thus completing the preparation of training data.
[0232] In some embodiments, the process of identifying driving parameters for each video frame in a video frame sequence to obtain the driving parameters corresponding to each video frame mainly involves feature analysis and parameterized reconstruction of facial images. Image data in the video frame sequence is read frame by frame, and the coordinates of facial feature points of the target object in the image are located using a facial keypoint detection algorithm. Subsequently, based on the detected facial feature point coordinates, a preset 3D facial statistical model is solved and fitted. This fitting operation adjusts the coefficients of the 3D facial statistical model through an optimization algorithm, so that the 3D face shape corresponding to the model, after being projected onto a 2D plane, matches the facial feature points in the video frame. After fitting, the coefficients representing changes in facial expression and morphology are extracted as the driving parameters corresponding to that video frame. By traversing all video frames in the video frame sequence and repeating the above identification and extraction steps, the driving parameters corresponding one-to-one with each video frame in the time dimension are finally determined.
[0233] In some embodiments, a video frame sequence including the target object and a corresponding audio frame sequence are extracted from a video containing the target object. This step ensures the synchronization of video and audio, which is crucial for subsequent feature extraction and driving parameter recognition. Feature extraction is performed on the audio frame sequence to obtain speech feature sequence samples. Driving parameters are recognized for each video frame in the video frame sequence to obtain the driving parameters corresponding to each video frame. This involves computer vision techniques, such as facial landmark detection and expression recognition, to extract driving parameters related to facial expressions and head posture. Based on the driving parameters corresponding to each video frame, labels are constructed for the speech feature sequence samples. This step associates the driving parameters with the corresponding speech feature sequence samples to form labels. Labels can be the ground truth values of the driving parameters or other information related to the driving parameters, such as emotion labels, intonation labels, etc. A dataset containing speech feature sequence samples and their labels can be obtained, and these datasets can be used to train the parameter determination model. The advantage of this method is that it can provide high-quality training data, ensuring that the model can learn the mapping relationship between speech features and driving parameters, thereby improving the model's determination accuracy.
[0234] As an example, video processing software (such as FFmpeg, OpenCV, etc.) is used to extract video frame sequences from a video. These frame sequences contain changes in the facial expressions and head posture of the target object. The corresponding audio frame sequences are then extracted. These audio frame sequences contain the speech information of the target object. Feature extraction is performed on the audio frame sequences to obtain speech feature sequence samples. For example, speech signal processing techniques can be used to extract features such as Mel-frequency cepstral coefficients (MFCC) and fundamental frequency (F0) for each audio frame. These features collectively constitute the speech feature sequence samples, with each sample corresponding to one or more audio frames. Driving parameters are identified for each video frame in the video frame sequence. For example, computer vision techniques (such as facial landmark detection, expression recognition, etc.) can be used to extract the driving parameters of the target object's facial expressions and head posture in each video frame. These driving parameters may include the position, shape, and motion parameters of parts such as the eyes, mouth, and eyebrows, as well as head posture parameters (such as pitch, yaw, roll). Based on the driving parameters corresponding to each video frame, labels are constructed for the speech feature sequence samples. For example, each speech feature sequence sample can correspond to one or more driving parameters, which together constitute a label. The label can be the truth value of the driving parameter, or other information related to the driving parameter, such as sentiment label, intonation label, etc.
[0235] Thus, by extracting video and audio frame sequences from videos containing the target object, and performing feature extraction and driving parameter identification respectively, high-quality speech feature sequence samples and their labels can be constructed. This not only ensures the synchronization of speech and video but also provides rich training data for the parameter determination model, helping the model learn the complex mapping relationship between speech features and driving parameters. This not only improves the model's determination accuracy but also makes the generated facial animations more natural and realistic, providing strong support for real-time driving of virtual characters and showing broad application prospects.
[0236] In some embodiments, the determination of the second loss value based on the sub-labels of each of the first speech feature samples and each driving parameter in the driving parameter sequence can be implemented as follows: For each of the first speech feature samples, obtain the target sub-labels adjacent to the sub-labels of the first speech feature samples in the labels, and obtain the target driving parameters adjacent to the driving parameters of the first speech feature samples in the driving parameter sequence samples; determine the third loss value based on the target sub-labels and the target driving parameters; obtain the first rate of change of the driving parameters corresponding to the first speech feature samples, and obtain the second rate of change of the target driving parameters, and determine the fourth loss value based on the first rate of change and the second rate of change; and perform a weighted summation of the third loss value and the fourth loss value to obtain the second loss value.
[0237] In some embodiments, for each first speech feature sample, a target sub-label adjacent to the sub-label of that sample is found in the label. This typically means considering temporal relationships, i.e., sub-labels of the previous and next frames. Simultaneously, a target driving parameter adjacent to the driving parameter of the first speech feature sample is found in the driving parameter sequence. This also considers temporal relationships. Based on the target sub-label and target driving parameter, a third loss value is determined, which can be calculated by comparing the differences between sub-labels and driving parameters of adjacent frames, for example, using mean squared error (MSE) or other appropriate loss functions. A first rate of change of the driving parameter corresponding to the first speech feature sample is obtained. This may involve calculating the derivative or difference of the driving parameter over time to measure its rate of change. A second rate of change of the target driving parameter is obtained, also involving calculating the rate of change of the target driving parameter over time. Based on the first and second rates of change, a fourth loss value is determined, calculated by comparing the difference between the rate of change of the driving parameter and the rate of change of the target driving parameter. The third and fourth loss values are weighted and summed to obtain the second loss value. The weights can be adjusted according to the importance of the specific task to ensure that the model can simultaneously focus on the relationship between adjacent frames and their rate of change during the optimization process.
[0238] In some embodiments, the first rate of change refers to the difference between the driving parameters of the current frame and the driving parameters of the previous frame, typically expressed as the time derivative or difference of the driving parameters. For each first speech feature sample, its corresponding driving parameter is obtained, and then the difference between this driving parameter and the driving parameter of the previous frame is calculated. This can be achieved through a simple subtraction operation, i.e., subtracting the driving parameter of the previous frame from the driving parameter of the current frame to obtain the first rate of change. The second rate of change refers to the difference between the target driving parameter and the target driving parameter of the previous frame, also expressed as the time derivative or difference of the target driving parameter. For each target driving parameter, its difference from the target driving parameter of the previous frame is obtained. This can also be achieved through a subtraction operation, i.e., subtracting the target driving parameter of the previous frame from the target driving parameter of the current frame to obtain the second rate of change.
[0239] In some embodiments, during the specific implementation of determining the second loss value based on the sub-labels and driving parameter sequences of each first speech feature sample, the computational logic mainly covers the supervision of adjacent time-series data and the constraint of the parameter change rate. For each first speech feature sample, the sub-label adjacent to it in the time dimension is first located in the label data, and this sub-label is determined as the target sub-label; at the same time, the driving parameter adjacent to the driving parameter corresponding to the first speech feature sample in the time dimension is located in the generated driving parameter sequence, and this driving parameter is determined as the target driving parameter.
[0240] In some embodiments, a third loss value representing the prediction accuracy at adjacent time steps is calculated based on the difference between the target sub-label and the target driving parameters. To further constrain the smoothness or consistency of parameter changes, the rate of change of the driving parameters corresponding to the first speech feature sample is calculated as a first rate of change, and the rate of change of the target driving parameters is calculated as a second rate of change. Based on the numerical relationship or degree of difference between the first and second rates of change, a fourth loss value is constructed and determined. A second loss value used for optimizing model training is obtained by performing a weighted summation operation on the third and fourth loss values, that is, by multiplying each by a preset weight coefficient and then adding them together.
[0241] In some embodiments, the determination of the first loss value based on the label of the speech feature sequence sample and the driving parameter sequence corresponding to the speech feature sequence sample can be achieved as follows: the norm of the difference between the label of the speech feature sequence sample and the driving parameter sequence corresponding to the speech feature sequence sample is determined as the first loss value.
[0242] As an example, the expression for the first loss value above can be:
[0243] Among them, L simple Used to indicate the first loss value, The label is used to indicate the speech feature sequence sample, and X0 is used to indicate the driving parameter sequence corresponding to the speech feature sequence sample.
[0244] As an example, the expression for the second loss value above can be: L2 = λ vel L vel +λ smooth L smooth (3)
[0245] Where L2 is used to indicate the second loss value, L vel Used to indicate the third loss value, L smooth Used to indicate the fourth loss value, λ vel The weight λ used to indicate the third loss value smooth Used to indicate the weight of the fourth loss value.
[0246] As an example, the expression for the third loss value mentioned above can be:
[0247] Among them, L vel Used to indicate the third loss value, X t+1 and X t Indicate the target driving parameters at time t+1 and time t, respectively. and These indicate the target sub-labels at time t+1 and time t, respectively.
[0248] As an example, the expression for the fourth loss value above can be:
[0249] Among them, L smooth Used to indicate the fourth loss value, The second rate of change of the target driving parameters are indicated respectively. Used to indicate the first rate of change of the driving parameter corresponding to the first speech feature sample.
[0250] This allows for an effective measurement of the difference between the determined driving parameter sequence and the true values, especially when considering the relationships between adjacent frames and their rates of change. This not only helps improve the determination accuracy of the model but also ensures the temporal smoothness and naturalness of the generated driving parameter sequence, thereby enhancing the realism and fluidity of facial animation.
[0251] In step 105, facial animation of the virtual object is generated based on the driving parameter sequence.
[0252] In some embodiments, in the technical solution for animation production or real-time driving of virtual objects, the process of generating facial animation of the virtual object based on the driving parameter sequence involves the technical processing of converting abstract parameter signals into visualized three-dimensional model deformation or two-dimensional image sequences. The core of this process lies in establishing a mapping relationship between the driving parameters and the facial structure of the virtual object. The driving parameter sequence consists of several driving parameter frames arranged in chronological order, each containing a set of values describing the facial state at a specific moment. These sets of values are input into a rendering engine or animation compositing module to control the facial bone rotation angle, mesh vertex displacement, or fusion weight of the expression base of the virtual object. By parsing the data in the driving parameter sequence frame by frame and applying it to a pre-built virtual object model, the facial features of the virtual object undergo continuous deformation that conforms to physical laws or preset rules. Subsequently, the deformed model state of each frame is rendered and output, thereby obtaining a continuously playing facial animation video stream.
[0253] In some embodiments, in a lower-level implementation based on blendshape technology, each driving parameter in the driving parameter sequence is specifically represented as a set of weight coefficients. These weight coefficients correspond to multiple preset basic facial expression shapes in the virtual object model, such as upturned corners of the mouth, closed eyelids, and raised eyebrows. When generating facial animation, the driving parameter sequence is traversed, and for each time step, the corresponding weight coefficients are obtained. The weight coefficients are linearly weighted and superimposed with the vertex data of the corresponding basic facial expression shapes to calculate the final mesh shape of the virtual object's face at that moment. To avoid jumps or jitter between adjacent frames, the weight coefficients of adjacent time steps can be smoothed. If the change in weight coefficients between two adjacent time steps exceeds a preset change threshold, interpolation is performed between them to generate transition frame parameters, thereby ensuring the visual smoothness of the generated facial animation.
[0254] In some embodiments, in another lower-level implementation based on skeletal animation, the driving parameters in the driving parameter sequence are represented as rotation quaternions or translation vectors of facial bone nodes. The virtual object's facial model is bound to a sophisticated skeletal system, and mesh vertices are associated with corresponding bone nodes through a skinning algorithm. During animation generation, the driving parameter sequence is parsed to determine the target pose of each facial bone in three-dimensional space. The bone matrix is updated according to the target pose, causing deformation of the associated mesh skin. Furthermore, a physical simulation mechanism can be introduced; if the head movement speed indicated by the driving parameters is detected to be greater than a preset speed threshold, the inertial sway amplitude of the hair or accessories is automatically calculated and superimposed into the final animation effect to enhance the realism of the animation.
[0255] In some embodiments, in applications involving intelligent Q&A with virtual customer service, speech is synthesized based on the generated reply text, and a sequence of driving parameters matching the phonemes and prosody of the speech is calculated simultaneously. In this scenario, the driving parameter sequence is used to drive the facial model of the virtual customer service representative, ensuring that its lip movements strictly follow the rhythm of the speech. For example, when the driving parameter value corresponding to the "a" sound is greater than a preset threshold, the virtual object's mandible rotates downwards and the corners of its mouth stretch to the sides. As the speech content plays, the driving parameter sequence is continuously applied to generate a facial animation of the virtual customer service representative speaking naturally, which is then presented to the user through a display terminal, providing an interactive experience with visual feedback.
[0256] In some embodiments, in applications involving immersive virtual social interaction, users input voice through a microphone, and an algorithm generates a sequence of driving parameters containing emotional features in real time. In this scenario, the driving parameter sequence includes not only lip-sync parameters but also facial expression parameters representing emotions. If the parameter value indicating pleasant emotions in the driving parameter sequence remains above a preset intensity threshold, a facial animation of a virtual avatar with a smile and expressive eyes is generated based on the driving parameter sequence. Specific processing includes increasing the height of the mesh vertices in the cheekbone area and adjusting the texture map corresponding to the orbicularis oculi muscle to present a smile. Finally, the generated facial animation is transmitted in real time to other user terminals on the social platform, enabling the virtual avatar to replace the real user in face-to-face communication rich in emotional expression in virtual space. In some embodiments, the first voice feature sequence includes multiple first voice features, and the driving parameter sequence includes driving parameter combinations corresponding one-to-one with the first voice features. The driving parameter combinations include first driving parameters and second driving parameters. Referring to Figure 9, Figure 9 is a flowchart illustrating the facial animation generation method provided in this application embodiment. Step 105 shown in Figure 3 can be implemented through steps 1051 to 1052 shown in Figure 9.
[0257] In step 1051, for each combination of driving parameters in the driving parameter sequence, the head of the virtual object is driven to move based on the first driving parameter.
[0258] In some embodiments, the first speech feature sequence refers to a sequence containing multiple speech features, which may be extracted from audio signals and used to describe speech characteristics such as pitch, volume, and speech rate. The driving parameter sequence is a sequence of parameters corresponding to these speech features, used to control the animation of a virtual object (such as a virtual character). The first driving parameter and the second driving parameter may control different parts of the virtual object or different animation effects. For example, the first driving parameter may control the head movement of the virtual object, while the second driving parameter may control facial expressions or other body parts. The statement "based on the first driving parameter, the head of the virtual object is driven to move" means that for each combination of driving parameters in the driving parameter sequence, the first driving parameter is used to control the head movement of the virtual object. The driving parameters are mapped to the animation skeleton or control points of the virtual object to achieve head rotation, tilting, or other types of movement to mimic the speaker's head movements.
[0259] As an example, suppose there is a virtual character whose head movement needs to be driven by a speech. First, a first speech feature sequence is extracted from the speech, which may include pitch, volume, speech rate, etc. Then, a parameter-determining model is used to determine the corresponding driving parameter sequence based on these speech features. Each combination of driving parameters in the driving parameter sequence includes two parameters: a first driving parameter and a second driving parameter. For example, the first driving parameter might represent the head rotation angle, while the second driving parameter might represent the head tilt angle. For each combination of driving parameters in the driving parameter sequence, the virtual object's head is moved based on the first driving parameter. For example, if the first driving parameter represents the head rotation angle, then this parameter can be used to control the virtual character's head to rotate left or right. Meanwhile, the second driving parameter might be used to control the movement of other parts, such as facial expressions or body posture.
[0260] In some embodiments, the first driving parameters refer to parameters used to control the movement of the virtual object's head. These parameters may include the head's rotation angle (e.g., pitch, yaw, roll), tilt amplitude, nodding frequency, etc. By adjusting these parameters, the natural movements of a speaker's head, such as nodding, shaking, and tilting, can be simulated. The second driving parameters refer to parameters used to control the movement of various parts of the virtual object's face. These parameters may include the raising and lowering of eyebrows, the opening and closing of eyes, the opening and closing of the mouth, and the raising or lowering of the corners of the mouth, etc. By adjusting these parameters, a wide range of facial expressions of a speaker, such as smiling, surprise, anger, and sadness, can be simulated.
[0261] As an example, suppose there is a virtual character whose head and facial movements need to be driven by a speech. A first speech feature sequence is extracted from this speech, which may include pitch, volume, and speech rate. Then, a parameter-determining model is used to determine the corresponding driving parameter sequence based on these speech features. Each combination of driving parameters in the driving parameter sequence includes two parameters: a first driving parameter and a second driving parameter. For example, the first driving parameter might represent the head rotation angle, while the second driving parameter might represent facial expression parameters, such as the amplitude of a smile or the degree of eye opening. For each combination of driving parameters in the driving parameter sequence, the virtual object's head is driven to move based on the first driving parameter. For example, if the first driving parameter represents the head rotation angle, then this parameter can be used to control the virtual character's head to rotate left or right. Simultaneously, the various parts of the virtual object's face are driven to move based on the second driving parameter. For example, if the second driving parameter represents the amplitude of a smile, then this parameter can be used to control the degree to which the corners of the virtual character's mouth turn up; if the second driving parameter represents the degree of eye opening, then this parameter can be used to control the degree to which the virtual character's eyes widen or close.
[0262] In step 1052, based on the arrangement order of each of the driving parameter combinations in the driving parameter sequence, the animation frames corresponding to each driving parameter combination are spliced together to obtain the facial animation of the virtual object.
[0263] In some embodiments, the corresponding animation frames are arranged chronologically according to the order in which the driving parameter combinations appear in the driving parameter sequence. This means that the animation frame corresponding to the first driving parameter combination will be the first frame in the sequence, the animation frame corresponding to the second driving parameter combination will be the second frame, and so on. These sequentially arranged animation frames are then stitched together to form a continuous animation sequence. This process is similar to playing a series of static images in sequence to produce a dynamic effect. Through this stitching process, a complete facial animation of a virtual object is obtained. This animation can accurately reflect the head movement and facial expression changes of the virtual object under voice input, making the animation performance of the virtual character more natural and vivid.
[0264] As an example, suppose there is a virtual character whose facial animation needs to be generated based on a dialogue. First, the speech feature sequence is extracted from the dialogue, and then a parameter determination model is used to determine the driving parameter sequence based on these features. Each driving parameter combination in the driving parameter sequence includes two parameters: a first driving parameter controlling head movement and a second driving parameter controlling facial expression. Next, for each driving parameter combination, the virtual object's head is moved based on the first driving parameter, and the various parts of the virtual object's face are moved based on the second driving parameter, resulting in animation frames corresponding to each driving parameter combination. Now, these animation frames are concatenated based on the order of the driving parameter combinations in the driving parameter sequence. For example, suppose there are three driving parameter combinations in the driving parameter sequence, corresponding to three animation frames: frame 1, frame 2, and frame 3. Frames 1, 2, and 3 are concatenated sequentially according to the order in which the driving parameter combinations appear, forming a continuous animation sequence.
[0265] In some embodiments, the facial animation can be a facial animation including preset audio. The above-mentioned method of splicing the animation frames corresponding to the combination of each driving parameter to obtain the facial animation of the virtual object can also be implemented in the following way: splicing the animation frames corresponding to the combination of each driving parameter to obtain an animation frame sequence, and splicing the animation frame sequence and preset audio to obtain the facial animation.
[0266] In some embodiments, the facial animation described above may also be a facial animation that does not include preset audio. The above method of splicing the video frames corresponding to the combination of each driving parameter to obtain the facial animation of the virtual object can also be implemented in the following way: splicing the animation frames corresponding to the combination of each driving parameter to obtain an animation frame sequence, and determining the animation frame sequence as the facial animation of the virtual object.
[0267] In some embodiments, voice-over recording is performed simultaneously with animation frame production. Voice actors record corresponding dialogue and sound effects based on the animation's plot and character personalities. The produced animation frames are then synchronized with the voice-over. This step requires ensuring that the lip movements, facial expressions, etc., of the animation frames match the syllables and intonation of the voice-over to achieve a natural visual and auditory effect. The synchronized animation frames and voice-over are then spliced together to generate the animated video.
[0268] As an example, in applications involving virtual news broadcasting or online program recording, the processing of each combination of driving parameters in the driving parameter sequence aims to give the virtual anchor realistic behavioral performance. In this scenario, the system parses the driving parameter sequence and extracts a first driving parameter specifically used to characterize the head's spatial posture from each combination. This first driving parameter specifically includes the rotation angle and displacement data of the head in three-dimensional space. The system maps this first driving parameter to the skeletal nodes or mesh vertices of the virtual anchor's head model, thereby driving the virtual object's head to perform movements such as pitch, tilt, or turn to match changes in voice intonation. Simultaneously, based on the temporal arrangement of each combination of driving parameters in the driving parameter sequence, the system renders the virtual object's state at each moment to generate corresponding animation frames. The system then continuously splices and encodes all the generated animation frames on the timeline according to this arrangement, thereby synthesizing a virtual object facial animation video stream with smooth head movements and lip expressions.
[0269] As an example, in applications involving intelligent human-computer interaction or virtual educational assistants, the system also executes the aforementioned driving and generation logic to enhance the immersive experience of user interaction. The system acquires a sequence of driving parameters containing emotional and semantic features and iterates through each combination of driving parameters. For each combination, the system adjusts the virtual assistant's head orientation and posture in real time based on the first driving parameter, ensuring its head movements respond to user input or its own speech content. After adjusting the posture of each frame, the system serializes and combines the rendered individual animation frames according to the predetermined order of the driving parameter combinations in the sequence. Through this ordered splicing process, the system constructs a continuously playing image sequence, ultimately obtaining the virtual object's facial animation, enabling the virtual character to exhibit natural head behaviors such as nodding, shaking its head, or gazing when interacting. This allows for precise control of the virtual object's head and facial movements, ensuring close synchronization with the input speech feature sequence. Driving the virtual object's head movement based on the first driving parameter can simulate the natural head movements of a speaker, such as nodding and shaking its head. By driving the movement of various parts of a virtual object's face using the second driving parameter, a rich array of facial expressions, such as smiling, surprise, and anger, can be simulated. By stitching these animation frames together according to the sequence of driving parameters, a continuous and natural facial animation can be obtained. This not only improves the realism and expressiveness of virtual character animation but also provides powerful technical support for film production, game development, and virtual reality applications. By precisely controlling the head and facial movements of virtual objects, a more immersive experience can be created, allowing viewers to perceive more realistic character performances.
[0270] Thus, noise features are obtained by extracting features from preset noise, and a first speech feature sequence is obtained by extracting features from preset speech. Based on the various parts of the virtual object, the noise features are decomposed into multiple sub-noise features. For each sub-noise feature, it is fused with the first speech feature sequence to obtain a first fused feature sequence. Based on multiple first fused feature sequences, a driving parameter sequence matching the first speech feature sequence is determined. Based on the driving parameter sequence, the facial animation of the virtual object is generated. In this way, by decomposing the noise features into multiple sub-noise features based on the various parts of the virtual object's head (such as eyes, mouth, nose, etc.) and corresponding the noise features with the various parts of the virtual object's head, subsequent fusion and control become more precise. By fusing the sub-noise features with the first speech feature sequence to obtain the first fused feature sequence, the information of speech and noise is combined, providing more comprehensive information for determining the driving parameters. By determining a driving parameter sequence that matches the first speech feature sequence based on multiple first fusion feature sequences, and since the noise features correspond to the various parts of the virtual object's head, the driving parameter sequence determined based on multiple first fusion feature sequences can achieve precise control of each part. By generating the facial animation of the virtual object based on the driving parameter sequence, the driving parameter sequence can accurately control the various parts of the virtual object's head, facial expressions, and head posture, thereby effectively improving the adaptation between the virtual object's facial animation and speech.
[0271] The following will describe an exemplary application of the embodiments of this application in a real-world facial animation application scenario.
[0272] This application proposes an audio-driven head-eye motion generation technology based on a diffusion model, solving the problems of lacking vivid rhythmic eyebrow movements and overall head motion in the field of 3D face driving. Meanwhile, most current 3D facial digital human methods can only generate single 3D lip shapes and templated expressions, lacking prediction of head posture. Furthermore, most predict specific vertex positions of the 3D face rather than animation curve coefficients that are not limited to the character's identity, resulting in mostly rigid and limited 3D character outputs that cannot be integrated into modern art production pipelines. The method in this application, however, utilizes the powerful generation capabilities of the Diffusion Model to generate harmonious head-eye movements that conform to audio conditions, eliminating the complex process of manual adjustment or customization of head-eye motion. It can also seamlessly integrate with current art production pipelines, enabling the implementation of 3D digital human related business services. This application's embodiments have the following characteristics: A Transformer network is used as the denoising network for the Diffusion Model. Real-time detection of 2D single-person facial speaking videos generates 3D Blendshape weights (driving parameters, eye-tracking driving parameters, head rotation and translation parameters, etc.) as training data. This enables the use of a diffusion model to generate audio-driven head pose, eye movements, and facial driving parameters. This approach establishes a long-term correlation between speech and head-eye movements, resulting in consistent and reasonable driving results for the head, eyes, and face. This application's embodiments utilize a mutual attention mechanism between the head, eyes, and eyebrows, and a hierarchical mutual attention mechanism with audio rhythm for information interaction. Modeling is performed using an independent encoder to achieve multi-level learning. The learning capability of the diffusion model is used to generate virtual human head and facial driving parameters in a 3D scene from audio-driven data. The output is a coordinated rhythm of the head, eyebrows, and eyes, along with basic lip-sync animation curves that match the audio, resulting in a vivid overall 3D head-driven animation. This application embodiment does not require manual intervention in preprocessing or postprocessing. It can generate coordinated head movements and facial driving coefficients based on the corresponding speech, and can be seamlessly integrated into the entire art production pipeline. The corresponding animation curves can be directly imported into the virtual human scene in 3D application scenarios (such as Maya, UE) for animation driving and secondary editing, thereby improving the overall efficiency of head animation production.
[0273] This application's embodiments are applied in the field of virtual face driving. Based on voice generation, these embodiments can easily be extended to 3D applications and seamlessly integrate with subsequent art editing operations. Furthermore, because these embodiments can generate overall head and eye movement parameters consistent with the audio, the work of manually simulating random head and eye movements is significantly reduced, making them applicable to voice-driven 3D face tasks. Additionally, since these embodiments directly generate ARKit's Blendshape Weight, the generated parameters can be extended to any virtual human face and seamlessly integrated into the current art production pipeline, enabling direct implementation in 3D digital human-related business projects.
[0274] To address the current lack of vivid overall head animation in the 3D field, this application proposes a method for generating 3D face driving parameters using a Diffusion Transformer Model driven by multiple conditions, including audio. This application requires no additional human intervention, customizing head and eye movement schemes for each character, enabling end-to-end voice-driven generation of the entire head animation. Through this end-to-end training, this application captures the coordinated relationships between head movements, eye movements, and facial expressions, making the overall 3D virtual human performance more lifelike. Overall process: This application builds a data processing framework (sample generation and label construction) based on the open-source product MediaPipe, transforming 2D talking head videos into 3D driving parameters. Using this product, this application obtains parameters driving head movements in 3D head scenes as training data, including BlendShapes driving lip movements, eye movements, and facial expressions, as well as head rotation and translation parameters driving head movements.
[0275] After constructing training data that pairs voice and 3D driving parameters, this embodiment of the application uses a Diffusion Transformer Model to train voice-driven 3D head motion. Specifically, the steps are as follows:
[0276] First, training data for the 3D driving parameters required for training is constructed. For the purpose of this application's embodiments, namely generating coordinated and vivid 3D head movements, existing 3D head datasets cannot meet the current needs. They are limited by acquisition equipment and difficulty, and the content of these datasets is often small and homogeneous, lacking individual characteristics and flexibility.
[0277] Based on the above problems, this application embodiment uses a large amount of 2D human head speaking video data as the data source, and uses the open-source real-time detection and acquisition model Mediapipe as the main tool for converting 2D data into 3D data in this application embodiment.
[0278] MediaPipe can generate eye-tracking parameters and head rotation / translation coefficients corresponding to the face in the current 2D image based on the current video frame. Through this process, this embodiment can obtain 3D data from rich 2D information and extract audio from the data, thereby constructing matching data for audio-3D driving parameters for subsequent training. Although the 3D data extracted by MediaPipe has certain problems, such as jitter and incomplete expressions, this embodiment can eliminate these problems through post-processing steps, such as smoothing. By generating 3D Blendshape weights (driving parameters) through real-time detection of 2D single-person face speaking videos as training data (labels), it is possible to use 3D animation parameters collected from a large number of in-the-wild 2D datasets for model training. Unlike the limited and low-precision 3D datasets currently available, the dynamic facial and head movements in a large number of 2D videos provide excellent audio-facial motion data matching and a strong sense of rhythm for the model's basic data.
[0279] Secondly, multiple conditions are added to the network training through various condition addition methods to achieve more controllable generation. Since the embodiments of this application are based on existing 2D datasets for data construction, other conditional parameters in the original 2D datasets can also be applied to the 3D digital human driving field. Audio conditions are encoded using Wav2Vec2, and an attention mechanism is calculated based on cross attention and the initial 3D driving parameters to obtain the relationship between each 3D driving parameter and the audio conditions.
[0280] Then, different types of attention maps are introduced, and relative position encoding is added to guide the transformer network in generating the final driving coefficients. Although the Transformer can globally model the temporal information of long sequences, it does not have the positional modeling capabilities of recurrent neural networks such as RNNs. Therefore, to build a time series, relative position encoding such as ALiBi is introduced to perform a relative position-aware operation. This allows the network to perceive that for each frame, the conditions between adjacent positions and the driving parameters are most relevant, with their importance decreasing sequentially with the extension of relative time. This enables the network to fully consider the impact of temporal issues on the generation results of the current frame when modeling the attention mechanism of various conditions and driving parameters, resulting in more reliable and continuous head movement results.
[0281] Overall, the embodiments of this application first utilize a large amount of 2D data to convert 2D face videos into 3D driving parameters and head rotation parameters based on the MediaPipe open-source model. Then, the Diffusion Transformer network is used to generate head and eye movement driving parameters under audio conditions, and self-attention and cross-attention mechanisms are implemented. The driving parameters are used as Q, and the audio is used as K and V. Furthermore, the powerful generative capabilities of the diffusion model are used to generate diverse results based on audio.
[0282] In some embodiments, referring to Figure 10, Figure 10 is a schematic diagram of the principle of the facial animation generation method provided in this application embodiment. Given a speech feature sequence A 0:T (Obtained by encoding the speech shown in Figure 10), A 0:T Generate the corresponding 3D facial and eye animation parameter sequences and head motion parameter animation sequence X. 0:T (That is, the sequence of driving parameters shown in Figure 10). X 0:T
[0283] In some embodiments, referring to Figure 10, noise feature injection: by using the priority rule of X[1-7]>X[2], noise (noise time step) at different time steps is injected into the input sequence to enhance the robustness of the model to dynamic changes. Time series data (such as driving parameter sequences) are encoded in rounds to capture the dependencies in the time dimension. The long-term dependencies within the input sequence (such as motion features, head features, etc.) are extracted from the attention layer, for example, the association between head movement and overall action at different time steps is analyzed. Cross attention layer: Multimodal information (such as static domain B_depth and dynamic domain E_depth, E_1 / E_2) is fused to align different features (such as style features and target variable force layer) and enhance cross-level interaction. Dynamic-static feature fusion: Static parameters (such as static domain) and dynamic parameters (such as V, E_depth) are combined through multi-layer cross attention to solve the non-steady-state problem in the time series. Decoding layer generation: Based on the fused features, the target parameter sequence (such as mouth features, head parameter detection results) is decoded and generated, while the randomness of the output is controlled by the noise time step. The final output is a predicted sequence containing motion features, style features, and physical parameters (such as V), and the parameters are updated through time stepping. The prediction results are gradually optimized through multi-round encoding (round-encoding) to ensure temporal consistency and physical rationality. Dynamic noise injection is used to balance determinism and randomness through noise time steps and priority rules (X[1-7]>X[2]). Hierarchical attention combines self-attention and cross-attention to take into account both intra-sequence and cross-domain feature alignment. Multimodal fusion is used to model the hybrid of static parameters (B_depth) and dynamic parameters (E_1 / E_2), which is suitable for complex temporal prediction tasks (such as human motion generation).
[0284] In some embodiments, for speech representation, this application uses the self-supervised pre-trained speech model Wav2Vec2 to extract speech features, which performs excellently in facial animation generation tasks. To ensure that the sampling rate of speech features is consistent with that of facial motion, this application introduces a resampling layer after the temporal convolutional layer of Wav2vec2, so that the speech feature sequence and the facial animation sequence are aligned in time steps.
[0285] In some embodiments, in the 3D face representation of this application, a 51-dimensional BlendShape weight vector and a 3-dimensional head rotation vector are used to describe facial expressions and head pose. BlendShape weight vector: set as w∈R 51 Each element w i The values for the i-th BlendShape coefficient range from [0, 1]. These coefficients control subtle changes in facial features, such as raised eyebrows and upturned corners of the mouth, and can generate a wide range of expressions through linear combination.
[0286] In some embodiments, the head rotation vector: head rotation is represented using Euler angles, as... These correspond to rotation angles (in radians or degrees) around the x, y, and z axes, respectively. This three-dimensional vector is used to describe changes in head posture, such as nodding or shaking.
[0287] Therefore, a complete 3D face representation can be expressed as: p = [w; r] ∈ R 51 (6)
[0288] Where [·; ·] denotes the concatenation operation of vectors.
[0289] In some embodiments, the denoising network of this application consists of a pre-trained Wav2Vec2 encoder and a Transformer decoder. The Wav2Vec2 encoder extracts speech features A, and the Transformer decoder extracts noise observations X. n Iterative generation of predicted facial movements To ensure proper alignment of voice and facial movements, embodiments of this application introduce an alignment mask between mutual attention, where facial movement features at the target location only focus on the corresponding voice feature a. t Meanwhile, to enhance the model's learning of local features, this embodiment uses alignment bias as a memory mask in the Transformer's cross-attention layer, focusing only on audio features corresponding to adjacent frames. To handle sequences of arbitrary length, this embodiment employs a windowing strategy. The speech feature sequence is divided into segments of length T. w The window (filled if insufficient).
[0290] In some embodiments, for diffusion step n, the denoising network receives the current noisy facial motion parameters. speech features Outputs clean facial, eye, and head movement parameters.
[0291] In some embodiments, referring to Figure 11, which is a schematic diagram of the principle of the facial animation generation method provided in this application embodiment, Figure 11 shows the structure of the hierarchical audio-driven blend module (HABM) (i.e., the cross-attention layer described above). This module can effectively capture the implicit relationship between audio features and different facial regions (such as lips, eyes, eyebrows, and head posture), thereby generating coherent and natural 3D facial animation.
[0292] By employing a multi-layered structure, this embodiment decomposes facial motion into several layered components, ensuring accurate synchronization between audio-driven facial movements and head posture. The layered audio-driven fusion module of this embodiment divides the face into several key regions: lips, eyes, eyebrows, and head posture. By decomposing facial movements into these regions, the model can learn fine-grained control over the dynamics of each part. This structure enables the model to handle varying degrees of complexity in facial motion, from subtle lip movements to large head rotations. Let the input audio features be A = {a1, a2, ..., a...} T}, where T is the number of frames in the input audio sequence. For each time step, this embodiment of the application uses a feature extractor to obtain features of different parts of the decomposed face for the BlendShape parameter B of the noisy lips, eye movements, eyebrow movements, and head pose:
[0293] The objective of this application is to fuse audio features and facial expression features at the feature level by making full use of cross-attention mechanisms:
[0294] Among them, Q f =W Q f t K a =W K A, V a =W V A is a query, key, and value matrix of facial and audio features after linear transformation, and d is a scaling factor for the feature dimensions.
[0295] This hierarchical modular design allows for independent learning of different facial regions while also taking into account their interrelationships, resulting in more vivid and diverse facial animations.
[0296] In some embodiments, referring to Figure 11, the cross-attention layer shown in Figure 11 is used for multimodal feature fusion. Its core is to align and integrate information from different modalities (such as noise features, time-step encoding, and language sequence encoding) through an attention mechanism. The following is a detailed analysis of its feature fusion process: Noise Encoding (X): Contains noise features Et, noise feature vectors, and keys (K), used to capture randomness or uncertainty in the input data. Time-Step Encoding (Y): Contains time-step features Et, noise feature vectors, and values (V), used to model dependencies in the time dimension. Language Sequence Encoding: Provides language-related contextual information for aligning other modalities (such as noise and time steps). The core of the cross-attention layer is to achieve feature alignment and fusion through a query, key, and value mechanism: Query Generation: Generates a query vector from one modality (such as language sequence encoding) to "ask" for relevant information from other modalities. Key and Value Extraction: Extracts keys (K) and values (V) from another modality (such as noise encoding or time-step encoding). The key is used to calculate the similarity with the query, and the value is used for weighted fusion. Attention weight calculation: The similarity between the query and the key is calculated using a dot product or scaled dot product to generate attention weights. For example, the query of a language sequence is weighted with the key of the noise encoding to determine which noise features are relevant to the language context. The values (V) are weighted and summed using the attention weights to generate the fused feature representation. Through cross-attention, the features of the noise encoding (X) and the time-step encoding (Y) are aligned to capture the dynamic changes of noise in the time dimension. The language sequence encoding is aligned with the noise or time-step encoding to ensure consistency between the language context and dynamic features. For example, in a speech-driven animation task, the language sequence may drive noise features (such as mouth movements) and time-step features (such as movement rhythm). The fused features can be used for decoding or processing, such as generating target variables (such as animation parameters, speech features, etc.). The feature alignment and fusion effect are gradually optimized through multiple rounds of cross-attention iterations. Efficient alignment of noise, time steps, and language sequences is achieved through cross-attention. The attention mechanism dynamically adjusts feature weights according to the input, enhancing model flexibility. The language sequence encoding provides semantic context, guiding the feature fusion of other modalities.
[0297] In some embodiments, during the generation process, the results X under a given audio A condition are iteratively processed. 0 Sampling is performed. Specifically, in this embodiment of the application, the estimated clean samples are:
[0298] Then noise was reintroduced to obtain X. n-1This process is repeated n = N, N-1, ..., 1. At each sampling step t, embodiments of this application replace the predicted noise with a linear combination of conditional and unconditional estimates, thereby controlling the speech and style consistency of the generated results. Specifically, embodiments of this application set the style conditions to empty during training with a certain probability to enhance the model's generalization ability.
[0299] In some embodiments, to train the model of this application to generate realistic 3D facial animations, this application designs a series of loss functions designed to constrain the model's output from different levels. These loss functions include simple loss, geometric loss (vertex loss, velocity loss, and smoothing loss), and loss specific to head motion. The definition and function of each loss function will be described in detail below.
[0300] In some embodiments, Simple Loss (L) simple First, a simple loss method was employed to measure the error between the predicted and actual samples. Specifically, the simple loss method calculated the predicted facial motion parameter sequence. The squared L2 norm between the actual sequence X0 and the true sequence X0:
[0301] in, X0 represents the facial motion parameters generated by the model, while X0 is the set of real facial motion parameters and head parameters.
[0302] In some embodiments, velocity loss L vel The velocity loss aims to improve the temporal consistency of generated facial animations, preventing discontinuities or jitter in facial movements. It is defined as the squared L2 norm between the velocities of the real and predicted mesh sequences (differences in vertex positions between adjacent frames):
[0303] Among them, X t and Let represent the sets of real and predicted facial motion parameters and head parameters for frame t, respectively. By minimizing velocity loss, embodiments of this application can constrain the smooth transition of generated facial motions over time, maintaining the continuity of the action.
[0304] In some embodiments, the smoothing loss (L) smooth The smoothing loss is used to penalize large accelerations at predicted mesh vertices, improving the smoothness of the animation. It is defined as the squared L2 norm of the difference in velocity (i.e., acceleration) between adjacent vertices in the predicted mesh sequence:
[0305] Where M represents the acceleration before and after frame t. By minimizing the smoothing loss, unnatural rapid changes in the generated facial movements can be prevented, ensuring the smoothness of the movements.
[0306] In some embodiments, the total loss function, which combines the above loss terms, is defined as follows in this application embodiment: L = L simple +λ vel L vel ++λ smooth L smooth (14)
[0307] By jointly optimizing these loss functions, the model in this application embodiment can generate 3D facial animations with hyperspatial and temporal consistency, which can accurately capture the details of facial expressions and ensure the continuity and naturalness of the animation.
[0308] In some embodiments, the voice-driven 3D virtual human full-face animation method implemented based on the embodiments of this application can generate rich head and eye movements. This solution is the first in the 3D field to model the entire head movement animation and present vivid animation effects. Compared with contemporaneous work, for any audio, the results of the embodiments of this application can generate sufficiently vivid micro-expressions in the eyes and head movements, and can be seamlessly applied to 3D application editing software, such as Maya and Blender.
[0309] In some embodiments, referring to FIG12, FIG12 is a schematic diagram of the effect of facial animation provided in the embodiments of this application. FIG12 shows the process of generating a three-dimensional facial animation from the input of the original video data in the embodiments of this application. The facial animation generation method provided in the embodiments of this application processes the original data to obtain a two-dimensional video, and then converts the two-dimensional video to obtain a three-dimensional facial animation. The three-dimensional facial animation has a variety of mixed shapes, including mixed shape A, mixed shape B, mixed shape C, mixed shape D, mixed shape E, and mixed shape F.
[0310] In some embodiments, referring to Figure 13, which is a schematic diagram of the facial animation effect provided by the embodiments of this application, demonstrating the performance of the embodiments of this application on a Chinese dataset. By introducing joint training of eye movement, eyebrow movement, and head movement during the training process, the animation generated by the embodiments of this application is more vivid and realistic. By training the joint eye movement, eyebrow movement, and head movement, the generated expressions are more natural and vivid. Figure 13 shows the animation frames corresponding to voice A (animation frames of eye movement, eyebrow movement, and head movement and animation frames of mouth shape only), voice B (animation frames of eye movement, eyebrow movement, and head movement and animation frames of mouth shape only), voice C (animation frames of eye movement, eyebrow movement, and head movement and animation frames of mouth shape only), voice D (animation frames of eye movement, eyebrow movement, and head movement and animation frames of mouth shape only), and voice E (animation frames of eye movement, eyebrow movement, and head movement and animation frames of mouth shape only).
[0311] In some embodiments, referring to Figure 14, which is a schematic diagram of the facial animation effect provided by the present application embodiment, Figure 14 shows the comparison results of the rendering visualization of the present application and the prior art in Unreal Engine. Even when tested on an unseen English dataset, the present application significantly outperforms the prior art that uses BlendShape weight prediction in terms of lip shape accuracy. Furthermore, compared to related technologies 1 and 2, the present application makes the digital human's expressions more human-like due to the synchronous coordination of eye movement, eyebrow movement, and head movement, enhancing the overall expressiveness and realism. The embodiments of the present application maintain excellent performance on unseen distributed datasets, demonstrating the model's generalization ability. Compared to other methods for generating facial driving parameters, the embodiments of the present application show good performance in terms of vividness and lip shape accuracy. This consistent performance across different datasets proves the effectiveness of jointly training eye movement, eyebrow movement, and head movement.
[0312] In some embodiments, the Diffusion Transformer-based model generates all data in parallel. Therefore, the previous frame training data is not included as a condition to guide the generation of the result during training. This results in the inability to learn the complete data distribution based on the existing conditions when the data distribution is complex, thus failing to generate vivid results. Subsequently, by using a fusion of autoregressive and diffusion models, the previous frame information can be added to guide the generation of the current frame result, which can bring more accurate generation results.
[0313] It is understood that in the embodiments of this application, data related to noise and so on are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0314] The following description further illustrates the exemplary structure of the facial animation generation device 455 provided in this application embodiment as a software module. In some embodiments, as shown in FIG2, the software module stored in the facial animation generation device 455 in the memory 450 may include: a feature extraction module 4551, configured to extract features from preset noise to obtain noise features, and to extract features from preset speech to obtain a first speech feature sequence, the first speech feature sequence including multiple first speech features; a decomposition module 4552, configured to decompose the noise features into multiple sub-noise features, the multiple sub-noise features respectively corresponding to a part of the head of a virtual object; a fusion module 4553, configured to fuse the multiple sub-noise features with the first speech feature sequence to obtain a first fused feature sequence; and a determination module 4554, configured to determine a driving parameter sequence matching the first speech feature sequence based on the first fused feature sequence, the driving parameter sequence including multiple driving parameters, the driving parameters being used to control at least one of the facial expressions and head postures of the virtual object, and to generate the facial animation of the virtual object based on the driving parameter sequence.
[0315] In some embodiments, the fusion module is further configured to fuse the first speech feature sequence and the noise feature to obtain a second speech feature sequence; and to fuse the plurality of sub-noise features with the second speech feature sequence to obtain a plurality of first fused feature sequences.
[0316] In some embodiments, the fusion module is further configured to concatenate the first speech feature and the noise feature for each first speech feature in the first speech feature sequence to obtain a concatenated feature; and to arrange the concatenated features according to the order of the first speech features in the first speech feature sequence to obtain the second speech feature sequence.
[0317] In some embodiments, the fusion module is further configured to perform two different linear transformations on the second speech feature sequence to obtain a third speech feature sequence and a fourth speech feature sequence; and to perform the following processing on each of the sub-noise features: fusing the third speech feature sequence and the sub-noise feature to obtain a second fused feature sequence; and fusing the fourth speech feature sequence and the second fused feature sequence to obtain the first fused feature sequence corresponding to the sub-noise feature.
[0318] In some embodiments, the first speech feature sequence includes a plurality of first speech features, and the third speech feature sequence includes second speech features that correspond one-to-one with the first speech features; the fusion module is further configured to select a target speech feature from the third speech feature sequence, the target speech feature being associated with the sub-noise feature; construct a fifth speech feature sequence based on the target speech feature and the second speech feature, wherein the second speech feature and the target speech feature are adjacent in position in the third speech feature sequence; and fuse the fifth speech feature sequence and the sub-noise feature to obtain the second fused feature sequence.
[0319] In some embodiments, the fusion module is further configured to obtain a mask that corresponds one-to-one with each of the first speech features in the first speech feature sequence; based on the mask, filter the first speech features in the first speech feature sequence to obtain a filtered first speech feature sequence; and fuse the filtered first speech feature sequence and the sub-noise features for each of the sub-noise features to obtain the first fused feature sequence.
[0320] In some embodiments, the fusion module is further configured to perform the following processing on each of the first speech features in the first speech feature sequence: determining the correlation between the first speech feature and the sub-noise feature; and determining the mask corresponding to the first speech feature based on the correlation.
[0321] In some embodiments, the fusion module is further configured to: determine the first speech feature as a third speech feature and determine that the third speech feature is associated with the sub-noise feature when the first speech feature and the sub-noise feature correspond to the same part of the virtual object; determine that the first speech feature is associated with the sub-noise feature when the first speech feature is adjacent to the third speech feature in the first speech feature sequence; and determine that the first speech feature is not associated with the sub-noise feature when the first speech feature and the sub-noise feature correspond to different parts of the virtual object and the first speech feature and the third speech feature are not adjacent.
[0322] In some embodiments, the feature extraction module is further configured to extract features from the preset noise to obtain initial noise features, and to extract features from the head data of the virtual object to obtain head features of the virtual object; and to fuse the initial noise features and the head features to obtain the noise features.
[0323] In some embodiments, the aforementioned determining module is further configured to fuse fusion features at the same position in the plurality of first fused feature sequences according to the arrangement order of each fused feature in the first fused feature sequences, to obtain a third fused feature sequence; and determine the driving parameter sequence matching the first speech feature sequence based on the third fused feature sequence.
[0324] In some embodiments, the aforementioned determining module is further configured to respectively determine and obtain sub-driving parameter sequences based on each of the first fused feature sequences, and combine sub-driving parameters at the same position in each of the sub-driving parameter sequences to obtain the driving parameter sequence matching the first speech feature sequence.
[0325] In some embodiments, the aforementioned determining module is further configured to perform the first determination of driving parameters based on the plurality of first fused feature sequences to obtain a first driving parameter sequence; perform the (i+1)th determination of driving parameters based on the i-th driving parameter sequence to obtain an (i+1)th driving parameter sequence, where 1≤i<N and N is a positive integer greater than 1; traverse i, and determine the N-th driving parameter sequence obtained by traversing i as the driving parameter sequence matching the first speech feature sequence.
[0326] In some embodiments, the i-th driving parameter sequence includes a plurality of driving parameter combinations, and each driving parameter combination corresponds to one animation frame in the facial animation; the aforementioned determining module is further configured to, for each driving parameter combination in the i-th driving parameter sequence, perform feature extraction on the driving parameter combination to obtain parameter features; fuse the first speech feature sequence with each of the parameter features respectively to obtain a plurality of i-th parameter feature sequences; and perform the (i+1)th determination of driving parameters based on the plurality of i-th parameter feature sequences to obtain the (i+1)th driving parameter sequence.
[0327] In some embodiments, the first speech feature sequence includes a plurality of first speech features, the driving parameter sequence includes driving parameter combinations in one-to-one correspondence with the first speech features, and each driving parameter combination includes a first driving parameter and a second driving parameter; the aforementioned determining module is further configured to, for each driving parameter combination in the driving parameter sequence, drive the head of the virtual object to move based on the first driving parameter, and drive each part of the face of the virtual object to move based on the second driving parameter, to obtain an animation frame corresponding to the driving parameter combination; and splice the animation frames corresponding to each driving parameter combination according to the arrangement order of each driving parameter combination in the driving parameter sequence, to obtain the facial animation of the virtual object.
[0328] In some embodiments, the driving parameter sequence is determined by a parameter determination model. The facial animation generation device further includes: a training module configured to acquire speech feature sequence samples, the speech feature sequence samples including a plurality of first speech feature samples, the labels of the speech feature sequence samples including sub-labels corresponding one-to-one with each of the first speech feature samples; calling an initial parameter determination model to determine a driving parameter sequence corresponding to the speech feature sequence samples based on the speech feature sequence samples; determining a first loss value based on the labels of the speech feature sequence samples and the driving parameter sequence corresponding to the speech feature sequence samples; determining a second loss value based on the sub-labels of each of the first speech feature samples and each driving parameter in the driving parameter sequence corresponding to the speech feature sequence samples; and updating the model parameters of the initial parameter determination model based on the first loss value and the second loss value to obtain the parameter determination model.
[0329] In some embodiments, the training module is further configured to extract a video frame sequence including the target object and an audio frame sequence including the target object from a video including the target object; perform feature extraction on the audio frame sequence to obtain the speech feature sequence sample; identify the driving parameters of each video frame in the video frame sequence to obtain the driving parameters corresponding to each video frame; and construct a label for the speech feature sequence sample based on the driving parameters corresponding to each video frame.
[0330] In some embodiments, the training module is further configured to, for each first speech feature sample, obtain a target sub-label adjacent to the sub-label of the first speech feature sample in the label, and obtain a target driving parameter adjacent to the driving parameter of the first speech feature sample in the driving parameter sequence; determine a third loss value based on the target sub-label and the target driving parameter; obtain a first rate of change of the driving parameter corresponding to the first speech feature sample, and obtain a second rate of change of the target driving parameter, and determine a fourth loss value based on the first rate of change and the second rate of change; and perform a weighted summation of the third loss value and the fourth loss value to obtain the second loss value.
[0331] In some embodiments, the first fusion feature sequence includes M fusion features, where M is an integer greater than 2; the determining module is further configured to: determine a first driving parameter matching the first first speech feature in the first speech feature sequence based on the first fusion feature in the first fusion feature sequence; determine a second driving parameter matching the second first speech feature in the first speech feature sequence based on the second fusion feature in the first fusion feature sequence and the first driving parameter; determine a j-th driving parameter matching the j-th first speech feature in the first speech feature sequence based on the j-th fusion feature in the first fusion feature sequence and the first to (j-1)-th driving parameters, where j is an integer greater than 2; and iterate through j until j equals M to obtain a sequence of driving parameters matching the first speech feature sequence.
[0332] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the facial animation generation method described above in this application.
[0333] This application provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor will execute the facial animation generation method provided in this application, such as the facial animation generation method shown in FIG3.
[0334] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of electronic devices including one or any combination of the above-mentioned memories.
[0335] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0336] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0337] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0338] In summary, the embodiments of this application have the following beneficial effects:
[0339] (1) Noise features are obtained by extracting features from preset noise, and a first speech feature sequence is obtained by extracting features from preset speech. The noise features are decomposed into multiple sub-noise features. Since the sub-noise features correspond to different parts of the virtual object's head, the subsequent fusion and control have higher granularity and are therefore more accurate. By fusing the sub-noise features with the first speech feature sequence, a first fused feature sequence is obtained, thereby combining the information of speech and noise and providing more comprehensive information for determining the driving parameters. Based on the first fused feature sequence, a driving parameter sequence matching the first speech feature sequence is determined. Since the noise features correspond to different parts of the virtual object's head, the driving parameter sequence determined based on the first fused feature sequence can achieve precise control of the virtual object's facial expressions and head posture. Through the driving parameter sequence, the virtual object's facial animation is generated, thereby enabling the driving parameter sequence to accurately control the virtual object's facial expressions and head posture, thus effectively improving the adaptation of the virtual object's facial animation and speech.
[0340] (2) By extracting features from preset noise, the initial characteristics of the noise can be obtained, which helps to gain a deeper understanding of the nature of the noise and provides a basis for subsequent noise processing and suppression. Extracting features from the head data of virtual objects can not only capture the motion state of virtual objects in the virtual environment, but also infer the user's intentions and focus through head posture and direction. Fusing these two features can generate more comprehensive and accurate noise features, which can improve the realism of sound in the virtual environment, allowing users to experience more realistic sound effects in an immersive experience, and improving the clarity and intelligibility of speech in noisy environments, thereby improving the quality of communication and interaction.
[0341] (3) By decomposing the noise features of the virtual object's head into sub-noise features corresponding to each part, refined noise processing can be achieved, thereby significantly improving the visual realism and detail of the virtual image. Decomposition allows for personalized noise reduction processing for different areas of the head. For example, preset noise in the eye area can be softened to maintain the vividness of the eyes, while preset noise in the mouth area can be sharpened to enhance the clarity of the mouth shape. When the number of noise features cannot be evenly distributed, adding row features with zero feature elements can ensure that each part has a corresponding sub-noise feature, avoiding inconsistencies caused by uneven noise distribution and improving the consistency and stability of the processing results.
[0342] (4) By concatenating each first speech feature in the first speech feature sequence with its corresponding noise feature and arranging them in the original order, the resulting second speech feature sequence can more comprehensively reflect the characteristics of the speech signal. This helps improve the accuracy and robustness of speech recognition, especially in noisy environments. By combining the information of speech features and noise features, the influence of noise can be better suppressed, enhancing the intelligibility and clarity of speech.
[0343] (5) By performing two different linear transformations on the second speech feature sequence, the third and fourth speech feature sequences can be obtained. These two sequences reveal different aspects of the original speech features, enhancing the expressiveness and discriminative power of the features. The third speech feature sequence is fused with the sub-noise features to obtain the second fused feature sequence. This fusion effectively combines noise information with speech features, helping to better identify and suppress noise in subsequent processing. The fourth speech feature sequence is fused with the second fused feature sequence to form the first fused feature sequence, which not only contains rich speech information but also integrates noise characteristics. This feature sequence helps improve the robustness of the speech recognition system, reduce noise interference, and improve recognition accuracy and speech enhancement quality.
[0344] (6) By selecting target speech features associated with sub-noise features from the third speech feature sequence and constructing a fifth speech feature sequence by combining it with the adjacent second speech features, and then fusing the sequence with the sub-noise features to obtain the second fused feature sequence, this process can significantly improve the processing effect of speech signals. It not only enhances the discriminativeness of speech features and improves the accuracy of speech recognition, but also effectively reduces the impact of environmental noise on speech signals by fusing noise features, thereby improving the quality and robustness of speech signals.
[0345] (7) The first speech feature sequence and the noise feature are fused to obtain the second speech feature sequence. This process can effectively combine the information of speech and noise, and improve the robustness of speech features. Subsequently, the sub-noise features are fused with the second speech feature sequence to obtain the first fused feature sequence, which enhances the adaptability of speech features to noise, and significantly improves speech recognition and speech enhancement performance in noisy environments.
[0346] (8) It can effectively filter and adjust the noise components in the speech feature sequence, thereby improving the quality and clarity of the speech signal. The correlation between each first speech feature and the sub-noise feature is evaluated separately to ensure that only those features affected by noise are labeled and processed, while other useful features remain unchanged.
[0347] (9) It can effectively identify and separate speech features affected by noise, thereby improving the quality and clarity of speech signals. It can ensure that only those speech features associated with noise sources are labeled and processed, while other irrelevant features remain unchanged, thereby improving the accuracy and efficiency of speech processing.
[0348] (10) It can effectively separate and retain useful information in speech signals while reducing the impact of noise. It improves the quality and clarity of speech signals and enhances the accuracy of speech recognition and understanding. The fused feature sequence not only contains clean speech features but also retains a certain amount of noise information.
[0349] (11) Information from multiple first fusion feature sequences can be integrated to obtain a more comprehensive and accurate third fusion feature sequence. This sequence contains various information such as acoustics and semantics in the speech signal, which can better reflect the speaker's intention and emotion. Based on this third fusion feature sequence, a driving parameter sequence matching the first speech feature sequence can be determined, thereby achieving precise control over the facial expressions and head posture of the virtual object.
[0350] (12) The corresponding sub-driving parameter sequences are determined by utilizing the unique information of each first fusion feature sequence. By combining the parameters at the same position in these sub-sequences, a comprehensive and accurate driving parameter sequence can be obtained. This sequence can better match the original speech feature sequence, thereby achieving precise control over the facial expressions and head posture of the virtual object. This not only improves the accuracy of the driving parameter determination but also enhances the expressive ability of the virtual character, bringing users a more realistic and immersive experience.
[0351] (13) By fusing speech feature sequences with driving parameter features, the information in the speech signal can be fully utilized, improving the accuracy of driving parameter determination. The iterative determination process optimizes each determination based on the previous result, which helps improve the robustness and accuracy of the entire system. It can dynamically adapt to changes in the speech signal and adjust the driving parameter sequence in real time, enabling facial animation to accurately reflect the speaker's expression and emotions. By extracting and fusing features for each combination of driving parameters, the efficiency of driving parameter determination can be improved, and the computational load can be reduced. It can enhance the expressive ability of virtual characters, making them more natural and realistic, and bringing a better user experience.
[0352] (14) The first driving parameter is determined based on multiple first fusion feature sequences to obtain the first driving parameter sequence. Then, the driving parameter is determined based on the i-th driving parameter sequence to obtain the i+1-th driving parameter sequence. The i-th driving parameter sequence is iterated over and the N-th driving parameter sequence obtained by iterating over i is determined as the driving parameter sequence that matches the first speech feature sequence. This improves the accuracy of the driving parameter determination. The process of iterating over i and determining the N-th driving parameter sequence that matches the first speech feature sequence is not just a simple iterative determination, but also a gradual denoising and optimization process. Through N iterations, the noise and inaccuracy in the driving parameter sequence can be effectively reduced, so that the final N-th driving parameter sequence more accurately reflects the original speech feature sequence.
[0353] (15) By progressively determining the driving parameters, each driving parameter can be optimized step by step, improving the accuracy of the entire driving parameter sequence. When determining the j-th driving parameter, not only the j-th fusion feature is used, but also the previously determined driving parameters from the 1st to the (j-1)th. Historical information can be utilized to improve the coherence and consistency of the determination. It is determined progressively, which can better adapt to changes in the speech signal. When the speech signal changes, more accurate driving parameters can be determined by adjusting the current fusion feature and the previous driving parameters. It can be adjusted according to different first fusion feature sequences and first speech feature sequences, providing high flexibility.
[0354] (16) By extracting video and audio frame sequences from videos containing target objects and performing feature extraction and driving parameter identification respectively, high-quality speech feature sequence samples and their labels can be constructed. This not only ensures the synchronization of speech and video but also provides rich training data for the parameter determination model, helping the model learn the complex mapping relationship between speech features and driving parameters. This not only improves the model's determination accuracy but also makes the generated facial animations more natural and realistic, providing strong support for the real-time driving of virtual characters and showing broad application prospects.
[0355] (17) It can effectively measure the difference between the determined driving parameter sequence and the true value, especially when considering the relationship between adjacent frames and their rate of change. It not only helps to improve the determination accuracy of the model, but also ensures the temporal smoothness and naturalness of the generated driving parameter sequence, thereby improving the realism and fluidity of facial animation.
[0356] (18) It can precisely control the head and facial movements of virtual objects, making them closely synchronized with the input speech feature sequence. Driving the virtual object's head based on the first driving parameter can simulate the natural head movements of a speaker, such as nodding and shaking. Driving the various parts of the virtual object's face based on the second driving parameter can simulate the rich facial expressions of a speaker, such as smiling, surprise, and anger. By splicing these animation frames according to the order in the driving parameter sequence, a continuous and natural facial animation can be obtained. This not only improves the realism and expressiveness of virtual character animation but also provides strong technical support for film production, game development, virtual reality, and other scenarios. By precisely controlling the head and facial movements of virtual objects, a more immersive experience can be created, allowing viewers to feel a more realistic character performance.
[0357] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for generating facial animation, the method comprising: Noise features are obtained by extracting features from preset noise, and a first speech feature sequence is obtained by extracting features from preset speech. The noise feature is decomposed into multiple sub-noise features, each of which corresponds to a part of the head of the virtual object. The plurality of sub-noise features are fused with the first speech feature sequence to obtain a first fused feature sequence; Based on the first fused feature sequence, a driving parameter sequence matching the first speech feature sequence is determined. The driving parameter sequence includes multiple driving parameters, which are used to control at least one of the facial expressions and head postures of the virtual object. Based on the driving parameter sequence, the facial animation of the virtual object is generated.
2. The method according to claim 1, wherein, The step of fusing the plurality of sub-noise features with the first speech feature sequence to obtain a first fused feature sequence includes: The first speech feature sequence and the noise features are fused to obtain the second speech feature sequence; The multiple sub-noise features are fused with the second speech feature sequence to obtain multiple first fused feature sequences.
3. The method according to claim 2, wherein, The step of fusing the first speech feature sequence and the noise features to obtain the second speech feature sequence includes: For each first speech feature in the first speech feature sequence, the first speech feature and the noise feature are concatenated to obtain the concatenated feature; The concatenated features are arranged according to their order in the first speech feature sequence to obtain the second speech feature sequence.
4. The method according to claim 2, wherein, The step of fusing the plurality of sub-noise features with the second speech feature sequence to obtain a plurality of first fused feature sequences includes: The second speech feature sequence is subjected to two different linear transformations to obtain the third and fourth speech feature sequences. The following processing is performed on each of the aforementioned sub-noise features: The third speech feature sequence and the sub-noise features are fused to obtain the second fused feature sequence; The fourth speech feature sequence and the second fusion feature sequence are fused to obtain the first fusion feature sequence corresponding to the sub-noise feature.
5. The method according to claim 4, wherein, The step of fusing the fourth speech feature sequence and the second fusion feature sequence to obtain the first fusion feature sequence corresponding to the sub-noise feature includes: For each fourth speech feature in the fourth speech feature sequence, the fourth speech feature and the second fusion feature at the same position in the second fusion feature sequence are concatenated to obtain the concatenated feature corresponding to the fourth speech feature. According to the arrangement order of the fourth speech features in the fourth speech feature sequence, the splicing features corresponding to each fourth speech feature are arranged to obtain the first fused feature sequence.
6. The method according to claim 4, wherein, The first speech feature sequence includes multiple first speech features, and the third speech feature sequence includes second speech features that correspond one-to-one with the first speech features; The step of fusing the third speech feature sequence and the sub-noise features to obtain the second fused feature sequence includes: Select a target speech feature from the third speech feature sequence, wherein the target speech feature is associated with the sub-noise feature; Based on the target speech feature and the second speech feature, a fifth speech feature sequence is constructed, wherein the second speech feature and the target speech feature are adjacent in the third speech feature sequence. The fifth speech feature sequence and the sub-noise feature are fused to obtain the second fused feature sequence.
7. The method according to any one of claims 1 to 5, wherein, The step of fusing the plurality of sub-noise features with the first speech feature sequence to obtain a first fused feature sequence includes: Obtain the mask that corresponds one-to-one with each of the first speech features in the first speech feature sequence; Based on the mask, the first speech features in the first speech feature sequence are filtered to obtain the filtered first speech feature sequence. For each of the sub-noise features, the filtered first speech feature sequence and the sub-noise features are fused to obtain the first fused feature sequence.
8. The method according to claim 7, wherein, The step of obtaining the mask that corresponds one-to-one with each of the first speech features in the first speech feature sequence includes: For each of the first speech features in the first speech feature sequence, the following processing is performed: Determine the correlation between the first speech feature and the sub-noise feature; Based on the correlation, the mask corresponding to the first speech feature is determined.
9. The method according to claim 8, wherein, Determining the correlation between the first speech feature and the sub-noise feature includes: When the first speech feature and the sub-noise feature correspond to the same part of the virtual object, the first speech feature is determined as the third speech feature, and the third speech feature is determined to be associated with the sub-noise feature; When the first speech feature is adjacent to the third speech feature in the first speech feature sequence, it is determined that the first speech feature is associated with the sub-noise feature; When the first speech feature and the sub-noise feature correspond to different parts of the virtual object, and the first speech feature and the third speech feature are not adjacent, it is determined that the first speech feature and the sub-noise feature are not related.
10. The method according to claim 7, wherein, The step of filtering the first speech features in the first speech feature sequence based on the mask to obtain the filtered first speech feature sequence includes: For each of the first speech features in the first speech feature sequence, when the value of the mask corresponding to the first speech feature is equal to a preset value, the first speech feature is determined as the target speech feature; The target speech features are combined to obtain the first speech feature sequence.
11. The method according to claim 6, wherein, The step of selecting target speech features from the third speech feature sequence includes: The following processing is performed on each of the third speech features in the third speech feature sequence: The portion of the first speech feature corresponding to the virtual object and the portion of the sub-noise feature corresponding to the virtual object are compared to obtain a comparison result; When the comparison result indicates that the third speech feature and the sub-noise feature correspond to the same part of the virtual object, the third speech feature is determined as the target speech feature.
12. The method according to any one of claims 1 to 11, wherein, The step of extracting noise features from preset noise includes: Feature extraction is performed on the preset noise to obtain the initial noise features, and feature extraction is performed on the head data of the virtual object to obtain the head features of the virtual object; The initial noise features and the head features are fused together to obtain the noise features.
13. The method according to claim 12, wherein, The step of fusing the initial noise features and the head features to obtain the noise features includes: The initial noise features are determined as key vectors, and the header features are determined as value vectors and query vectors; or, the initial noise features are determined as the value vectors and query vectors, and the header features are determined as key vectors. The attention model is invoked to fuse the key vector, the value vector, and the query vector to obtain the noise features.
14. The method according to any one of claims 1 to 11, wherein, The step of determining the driving parameter sequence that matches the first speech feature sequence based on the first fused feature sequence includes: According to the arrangement order of each fusion feature in the first fusion feature sequence, the fusion features at the same position in multiple first fusion feature sequences are fused to obtain a third fusion feature sequence; Based on the third fusion feature sequence, a driving parameter sequence matching the first speech feature sequence is determined.
15. The method according to any one of claims 1 to 11, wherein, The step of determining the driving parameter sequence that matches the first speech feature sequence based on the first fused feature sequence includes: Based on each of the first fusion feature sequences, the sub-driving parameter sequences are determined respectively; The sub-driving parameters at the same position in each of the sub-driving parameter sequences are combined to obtain a driving parameter sequence that matches the first speech feature sequence.
16. The method according to any one of claims 1 to 11, wherein, The first fusion feature sequence includes M fusion features, where M is an integer greater than 2; the step of determining the driving parameter sequence matching the first speech feature sequence based on the first fusion feature sequence includes: Based on the first fusion feature in the first fusion feature sequence, a first driving parameter that matches the first first speech feature in the first speech feature sequence is determined; Based on the second fusion feature in the first fusion feature sequence and the first driving parameter, a second driving parameter that matches the second first speech feature in the first speech feature sequence is determined. Based on the j-th fusion feature in the first fusion feature sequence and the first to (j-1)-th driving parameters, determine the j-th driving parameter that matches the j-th first speech feature in the first speech feature sequence, where j is an integer greater than 2; The process involves iterating through j until j equals M, thereby obtaining a sequence of driving parameters that matches the first speech feature sequence.
17. The method according to any one of claims 1 to 8, wherein, The step of determining the driving parameter sequence that matches the first speech feature sequence based on the first fused feature sequence includes: Based on multiple first fusion feature sequences, the first driving parameters are determined to obtain the first driving parameter sequence. Determine the (i+1)-th driving parameter based on the i-th driving parameter sequence to obtain the (i+1)-th driving parameter sequence, where 1≤i<N, and N is a positive integer greater than 1; Traverse i, and determine the N-th driving parameter sequence obtained by traversing i as the driving parameter sequence matching the first speech feature sequence.
18. The method according to claim 17, wherein, The i-th driving parameter sequence comprises a plurality of driving parameter combinations, and each driving parameter combination corresponds to one video frame in the virtual object video; The step of determining the (i+1)-th driving parameter based on the i-th driving parameter sequence to obtain the (i+1)-th driving parameter sequence comprises: For each driving parameter combination in the i-th driving parameter sequence, perform feature extraction on the driving parameter combination to obtain parameter features; Fuse the first speech feature sequence with each of the parameter features respectively to obtain a plurality of i-th parameter feature sequences; Determine the (i+1)-th driving parameter based on the plurality of i-th parameter feature sequences to obtain the (i+1)-th driving parameter sequence.
19. The method according to any one of claims 1 to 18, wherein, The first speech feature sequence comprises a plurality of first speech features, the driving parameter sequence comprises driving parameter combinations in one-to-one correspondence with the first speech features, and the driving parameter combination comprises a first driving parameter and a second driving parameter; The step of generating the facial animation of the virtual object based on the driving parameter sequence comprises: For each driving parameter combination in the driving parameter sequence, drive the head of the virtual object to move based on the first driving parameter, and drive each part of the face of the virtual object to move based on the second driving parameter, so as to obtain an animation frame corresponding to the driving parameter combination; Splice the animation frames corresponding to each driving parameter combination based on the arrangement order of each driving parameter combination in the driving parameter sequence to obtain the facial animation of the virtual object.
20. The method according to any one of claims 1 to 19, wherein, The driving parameter sequence is determined and obtained by a parameter determination model, and before fusing the plurality of sub-noise features with the first speech feature sequence to obtain a first fused feature sequence, the method further comprises: Acquiring a speech feature sequence sample, wherein the speech feature sequence sample comprises a plurality of first speech feature samples, and labels of the speech feature sequence sample comprise sub-labels corresponding to each of the first speech feature samples in one-to-one correspondence; Calling an initial parameter determination model, and determining a driving parameter sequence corresponding to the speech feature sequence sample based on the speech feature sequence sample; Determining a first loss value based on the label of the speech feature sequence sample and the driving parameter sequence corresponding to the speech feature sequence sample; Determining a second loss value based on the sub-label of each first speech feature sample and each driving parameter in the driving parameter sequence corresponding to the speech feature sequence sample; Updating model parameters of the initial parameter determination model based on the first loss value and the second loss value to obtain the parameter determination model.
21. The method according to claim 20, wherein, The step of acquiring the speech feature sequence sample comprises: Extracting a video frame sequence comprising the target object and an audio frame sequence of the target object from a video comprising the target object; Feature extraction is performed on the audio frame sequence to obtain the speech feature sequence sample, and the driving parameters of each video frame in the video frame sequence are identified to obtain the driving parameters corresponding to each video frame. Based on the driving parameters corresponding to each video frame, labels are constructed for the speech feature sequence samples.
22. The method according to claim 20, wherein, The determination of the second loss value based on the sub-labels of each of the first speech feature samples and each driving parameter in the driving parameter sequence corresponding to the speech feature sequence samples includes: For each of the first speech feature samples, a target sub-label adjacent to the sub-label of the first speech feature sample is obtained from the label, and a target driving parameter adjacent to the driving parameter of the first speech feature sample is obtained from the driving parameter sequence corresponding to the speech feature sequence sample. Based on the target sub-label and the target driving parameters, a third loss value is determined; Obtain the first rate of change of the driving parameter corresponding to the first speech feature sample, and obtain the second rate of change of the target driving parameter. Based on the first rate of change and the second rate of change, determine the fourth loss value. The second loss value is obtained by weighted summation of the third loss value and the fourth loss value.
23. An apparatus for generating facial animation, the apparatus comprising: The feature extraction module is configured to extract noise features from preset noise and extract features from preset speech to obtain a first speech feature sequence. The decomposition module is configured to decompose the noise feature into multiple sub-noise features, each of which corresponds to a part of the head of the virtual object. The fusion module is configured to fuse the plurality of sub-noise features with the first speech feature sequence to obtain a first fused feature sequence; The determination module is configured to determine a driving parameter sequence that matches the first speech feature sequence based on the first fused feature sequence. The driving parameter sequence includes multiple driving parameters, which are used to control at least one of the facial expressions and head postures of the virtual object. Based on the driving parameter sequence, the module generates a facial animation of the virtual object.
24. An electronic device, the electronic device comprising: Memory, configured to store computer-executable instructions or computer programs; When a processor is configured to execute computer-executable instructions or computer programs stored in the memory, it implements the facial animation generation method according to any one of claims 1 to 22.
25. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program, when executed by a processor, implement the method for generating facial animation according to any one of claims 1 to 22.
26. A computer program product comprising a computer program or computer-executable instructions, wherein the computer program or computer-executable instructions, when executed by a processor, implement the method for generating facial animation as described in any one of claims 1 to 22.