Virtual image generation method and device, equipment, medium and program product

By employing speech rate adaptive feature extraction and emotion recognition technology in the voice-driven virtual digital human model, the facial and body animations of the virtual avatar were optimized, solving the problems of speech rate adaptability and temporal correlation, improving the synchronization accuracy and naturalness of the virtual avatar, and enhancing the user experience.

CN120998233APending Publication Date: 2025-11-21LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511235170.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing voice-driven virtual digital human models lack dynamic adaptability to speech rate, resulting in low lip-reading accuracy in fast/slow speech scenarios, stiff facial expression transitions, unsmooth animations, and difficulty in capturing the temporal correlation of speech, which affects user experience and the realism of interaction.

Method used

By using a feature extraction sampling window with a sliding step size determined by speech rate, features are extracted from speech data. Combined with emotion recognition and a facial driving parameter generation model, a facial driving parameter sequence is output. The facial and limb models of the virtual avatar are then driven by a limb movement sequence to generate the target virtual avatar.

Benefits of technology

It improves the synchronization accuracy and stability of digital avatars, reduces response latency, enhances the smoothness and naturalness of virtual avatar animations, and improves user experience and the realism of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998233A_ABST
    Figure CN120998233A_ABST
Patent Text Reader

Abstract

The invention provides a virtual image generation method and device, equipment, a medium and a program product. The virtual image generation method comprises the following steps: performing feature extraction on voice data based on a feature extraction sampling window for determining a sliding step length by using a speech speed to obtain a voice feature sequence; an emotion category sequence and a voice feature sequence which are obtained through emotion recognition of the voice data are input into a face driving parameter generation model, a face driving parameter sequence is output, and the face driving parameter generation model is obtained based on training of the sample voice data; sample labels of the sample voice data comprise face driving parameter labels and phoneme labels; determining a limb action sequence based on the emotion category sequence, wherein the emotion category sequence, the face driving parameter sequence and the limb action sequence correspond in time sequence; and driving a face model of the target virtual image by using the face driving parameter sequence, and driving a limb model of the target virtual image by using the limb action sequence to generate the target virtual image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of virtual avatars and virtual digital human technologies, and more specifically, to a method, apparatus, device, medium, and program product for generating virtual avatars. Background Technology

[0002] With the rapid development of artificial intelligence technology, the generation and application of virtual digital humans (virtual avatars) have gradually become a new research hotspot. Virtual digital humans (Digital Human / Meta Human) can be understood as digitally created avatars that closely resemble human figures, and can be used in various fields such as metaverse interaction, games, movies, education, live streaming, and customer service. As an example, the driving methods for virtual digital humans can include voice-driven systems. Voice-driven technology can use deep learning to convert voice input into facial expression data for virtual digital humans, thereby completing the facial animation synthesis of the virtual digital human.

[0003] However, related voice-driven virtual digital human solutions suffer from problems such as response delays and unnatural driving effects, resulting in poor user experience and interaction realism. Summary of the Invention

[0004] One aspect of this disclosure provides a method for generating a virtual avatar, comprising: extracting features from speech data based on a feature extraction sampling window that uses speech rate to determine the sliding step size, thereby obtaining a speech feature sequence; inputting an emotion category sequence obtained by emotion recognition of the speech data and the speech feature sequence into a face-driven parameter generation model, and outputting a face-driven parameter sequence, wherein the face-driven parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face-driven parameter labels and phoneme labels; determining a body movement sequence based on the emotion category sequence, wherein the emotion category sequence, the face-driven parameter sequence, and the body movement sequence correspond temporally; using the face-driven parameter sequence to drive the facial model of the target virtual avatar, and using the body movement sequence to drive the body model of the target virtual avatar, thereby generating the target virtual avatar.

[0005] Another aspect of this disclosure provides a virtual avatar generation apparatus, including a feature extraction module for extracting features from speech data based on a feature extraction sampling window that uses speech rate to determine the sliding step size, thereby obtaining a speech feature sequence; a first processing module for inputting an emotion category sequence and a speech feature sequence obtained by emotion recognition of the speech data into a face-driving parameter generation model, and outputting a face-driving parameter sequence, wherein the face-driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face-driving parameter labels and phoneme labels; a second processing module for determining a limb movement sequence based on the emotion category sequence, wherein the emotion category sequence, the face-driving parameter sequence, and the limb movement sequence correspond temporally; and a driving module for driving the facial model of the target virtual avatar using the face-driving parameter sequence and driving the limb model of the target virtual avatar using the limb movement sequence, thereby generating the target virtual avatar.

[0006] Another aspect of this disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to perform the following operations: extracting features from speech data based on a feature extraction sampling window that uses speech rate to determine the sliding step size, to obtain a speech feature sequence; inputting an emotion category sequence obtained by emotion recognition of the speech data and the speech feature sequence into a face-driven parameter generation model, and outputting a face-driven parameter sequence, wherein the face-driven parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face-driven parameter labels and phoneme labels; determining a body movement sequence based on the emotion category sequence, wherein the emotion category sequence, the face-driven parameter sequence, and the body movement sequence correspond temporally; and using the face-driven parameter sequence to drive the facial model of a target virtual avatar, and using the body movement sequence to drive the body model of the target virtual avatar, to generate a target virtual avatar.

[0007] Another aspect of this disclosure provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, perform the following operations: extracting features from speech data based on a feature extraction sampling window that uses speech rate to determine the sliding step size, to obtain a speech feature sequence; inputting the emotion category sequence and the speech feature sequence obtained from emotion recognition of the speech data into a face-driven parameter generation model, and outputting a face-driven parameter sequence, wherein the face-driven parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face-driven parameter labels and phoneme labels; determining a body movement sequence based on the emotion category sequence, wherein the emotion category sequence, the face-driven parameter sequence, and the body movement sequence correspond temporally; using the face-driven parameter sequence to drive the facial model of a target virtual avatar, and using the body movement sequence to drive the body model of the target virtual avatar, to generate a target virtual avatar.

[0008] Another aspect of this disclosure provides a computer program, including computer-executable instructions, which, when executed by a processor, perform the following operations: extracting features from speech data based on a feature extraction sampling window that uses speech rate to determine the sliding step size, to obtain a speech feature sequence; inputting an emotion category sequence obtained from emotion recognition of the speech data and the speech feature sequence into a face-driven parameter generation model, and outputting a face-driven parameter sequence, wherein the face-driven parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face-driven parameter labels and phoneme labels; determining a body movement sequence based on the emotion category sequence, wherein the emotion category sequence, the face-driven parameter sequence, and the body movement sequence correspond temporally; using the face-driven parameter sequence to drive the facial model of a target virtual avatar, and using the body movement sequence to drive the body model of the target virtual avatar, to generate a target virtual avatar. Attached Figure Description

[0009] To gain a more complete understanding of this disclosure and its advantages, reference will now be made to the following description taken in conjunction with the accompanying drawings, wherein:

[0010] Figure 1 The illustrations illustrate application scenarios of the virtual avatar generation method, apparatus, device, medium, and program product according to embodiments of the present disclosure;

[0011] Figure 2 A flowchart illustrating a virtual avatar generation method according to an embodiment of the present disclosure is shown schematically.

[0012] Figure 3 This diagram schematically illustrates the technical architecture of a virtual avatar generation method according to an embodiment of the present disclosure.

[0013] Figure 4An exemplary diagram illustrating the CMVN normalization comparison effect according to an embodiment of the present disclosure is shown.

[0014] Figure 5 A flowchart illustrating a virtual avatar generation method according to an embodiment of the present disclosure is shown schematically;

[0015] Figure 6 A schematic diagram illustrating the structure of a virtual avatar generation apparatus according to an embodiment of the present disclosure is shown; and

[0016] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a virtual avatar generation method according to an embodiment of the present disclosure. Detailed Implementation

[0017] Embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0020] The accompanying drawings show some block diagrams and / or flowcharts. It should be understood that some blocks or combinations thereof in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when executed by the processor, these instructions can create means for implementing the functions / operations described in these block diagrams and / or flowcharts.

[0021] Therefore, the technology disclosed herein can be implemented in hardware and / or software (including firmware, microcode, etc.). Additionally, the technology disclosed herein can take the form of a computer program product stored on a computer-readable medium, which can be used by or in conjunction with an instruction execution system. In the context of this disclosure, a computer-readable medium can be any medium capable of containing, storing, transmitting, propagating, or transmitting instructions. For example, a computer-readable medium can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, apparatuses, or propagation media. Specific examples of computer-readable media include: magnetic storage devices, such as magnetic tape or hard disk drives (HDDs); optical storage devices, such as optical discs (CD-ROMs); memories, such as random access memory (RAM) or flash memory; and / or wired / wireless communication links.

[0022] In the publicly available technical solutions, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0023] This disclosure relates to the fields of artificial intelligence and virtual digital human technology, and can be applied, for example, to the field of virtual digital human driving technology. For ease of understanding, some of the technical terms mentioned herein are described below:

[0024] A digital human is a virtual image that exists in digital space in digital form, has an anthropomorphic or real human appearance, behavior and characteristics, and is presented through display devices. It can also be called a digital virtual human, virtual digital human, etc.

[0025] MFCC (Mel-scale Frequency Cepstral Coefficients) is an important acoustic feature extraction method in the field of speech recognition. It transforms speech signals into a form suitable for machine learning algorithms by simulating the auditory characteristics of the human ear. MFCC simulates the auditory characteristics of the human ear, converting the speech signal into spectral features through a series of mathematical transformations. The human ear perceives different frequencies of sound differently, and the Mel frequency scale is designed based on this perceptual characteristic. MFCC uses a Mel filter bank to decompose the spectrum into multiple frequency bands, and the energy of each frequency band is the MFCC coefficient.

[0026] According to one embodiment of this disclosure, a virtual digital human may include a driveable 3D digital human model generated using modeling methods, which is skeletally rigged and weighted. The 3D model of the virtual digital human may be created, for example, through modeling software, 3D scanning, artificial intelligence generation, etc. The 3D digital human model may consist of, for example, polygonal meshes (defining geometry), material maps (providing surface texture), and a skeletal system. For example, limbs and faces of the 3D digital human model may be driven separately using skeletal rig and BlendShape (blending shape animation).

[0027] In one embodiment, a skeletal system (i.e., skeletal rigging of the digital human) can be created inside the body of a 3D digital human model. Mesh vertices can be associated with skeletal joints by assigning weights to each joint on the model's surface, thus determining which bones influence each joint and the degree of influence. Skeletal animation, a computer animation technique widely used in 3D games, animated films, and other fields, is based on controlling bones to drive the movement of a 3D model. Skeletal animation technology divides the 3D model into two parts: the skin used to draw the model and the bones used to control the movement. The bones are composed of a series of bone nodes, each representing a bone, which are interconnected through parent-child relationships. By controlling the position and rotation of these bone nodes, the skin can be driven to deform accordingly, thereby achieving animation effects. Under bone control, the vertices of the skin mesh can be dynamically calculated through vertex blending, while the movement of a bone is relative to its parent bone and driven by animation keyframe data. A skeletal animation typically includes bone hierarchy data, mesh data, mesh skin info, and bone animation (keyframe) data.

[0028] BlendShape is a technique for 3D modeling and animation, primarily used to generate facial expressions. It allows for the creation of various combinations of facial expressions by linearly weighting the vertex positions of a mesh. Specifically, BlendShape transforms a "base mesh" into a "target shape" through vertex interpolation, typically used for creating facial animation and assisting in skeletal animation. The main working principle of BlendShape is: reading the vertex indices and position information of the "target shape," and then, based on weights, interpolating each vertex of the "base mesh" towards the "target shape," thus deforming the "base mesh." BlendShape-based solutions allow the final facial pose to be a linear combination of multiple facial expressions, known as a "BlendShape Target." Each "Target" can be a complete expression or a "delta" micro-expression, such as raising one eyebrow. In one embodiment, BlendShape can be used in conjunction with skeletal animation to provide more precise expression control and more lifelike virtual avatars.

[0029] In one embodiment, voice can be used to drive the face of a 3D digital human model. Voice-driven technology can be understood as using voice as the driving source and the head of the 3D model as the driving target, employing a deep learning model to generate 3D facial animation (including animations of lip movements, eyebrow changes, etc.). As an example, a pre-trained deep neural network model (e.g., the Audio2Face network) can be used to convert audio input into facial expression data (e.g., BlendShape weights) to achieve end-to-end voice-to-expression driving. This voice-driven technology can achieve lip-sync and expression generation by extracting features from the voice data and mapping them to the digital human's facial driving parameters (such as BlendShape weights).

[0030] However, existing voice-driven virtual digital human solutions suffer from the following technical problems: the models lack dynamic adaptability to speech rate, resulting in low lip-sync accuracy in both fast and slow speech scenarios; the models struggle to capture the temporal relationships of speech, leading to stiff facial expressions and unsmooth animations; and the recognition accuracy for certain special pronunciations (such as plosives and fricatives) is low. This results in high model response latency and unnatural driving effects, thus impacting user experience and the realism of interaction.

[0031] In view of this, the present disclosure provides a method, apparatus, device, medium, and program product for generating virtual avatars. The virtual avatar generation method includes: extracting features from speech data based on a feature extraction sampling window that uses speech rate to determine the sliding step size, obtaining a speech feature sequence; inputting an emotion category sequence obtained from emotion recognition of the speech data and the speech feature sequence into a face-driving parameter generation model, outputting a face-driving parameter sequence, wherein the face-driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face-driving parameter labels and phoneme labels; determining a limb movement sequence based on the emotion category sequence, wherein the emotion category sequence, the face-driving parameter sequence, and the limb movement sequence correspond temporally; using the face-driving parameter sequence to drive the facial model of the target virtual avatar, and using the limb movement sequence to drive the limb model of the target virtual avatar, thereby generating the target virtual avatar.

[0032] Figure 1 The illustrations illustrate application scenarios of the virtual avatar generation method, apparatus, device, medium, and program product according to embodiments of the present disclosure.

[0033] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0034] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0035] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0036] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0037] It should be noted that the virtual avatar generation method provided in this embodiment can generally be executed by server 105. Correspondingly, the virtual avatar generation device provided in this embodiment can generally be located in server 105. The virtual avatar generation method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the virtual avatar generation device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0039] Figure 2 A flowchart illustrating a virtual avatar generation method according to an embodiment of the present disclosure is shown schematically.

[0040] like Figure 2 As shown, the method includes operations S210~S240.

[0041] In operation S210, based on the feature extraction sampling window that uses speech rate to determine the sliding step size, feature extraction is performed on the speech data to obtain a speech feature sequence.

[0042] In operation S220, the emotion category sequence and speech feature sequence obtained by emotion recognition of speech data are input into the face driving parameter generation model, and the face driving parameter sequence is output. The face driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face driving parameter labels and phoneme labels.

[0043] In operation S230, a body movement sequence is determined based on the emotion category sequence, wherein the emotion category sequence, the facial driving parameter sequence, and the body movement sequence correspond in time.

[0044] In operation S240, the facial model of the target virtual image is driven by the facial driving parameter sequence, and the limb model of the target virtual image is driven by the limb movement sequence to generate the target virtual image.

[0045] According to embodiments of this disclosure, a virtual avatar may include a virtual digital human, and the target virtual avatar may be presented in the form of an animated video of the virtual digital human. Voice data may, for example, be audio recordings of the target virtual avatar to be generated. For instance, if the target virtual avatar to be generated is a digital human giving a speech, the voice data may be audio recordings of the speech content, and the target virtual avatar may be presented as an animated video of the digital human giving a speech according to the speech content (the visuals of the animated video are synchronized with the audio of the speech content).

[0046] According to embodiments of this disclosure, features can be extracted from speech data based on a feature extraction sampling window to obtain a speech feature sequence. The speech feature sequence may include multiple speech features arranged in temporal order. In one embodiment, the speech features may be MFCC features, and the feature extraction sampling window may be called an MFCC window, which is used to process a longer speech signal into multiple short time frames to facilitate feature extraction from the short time frames. Exemplarily, the MFCC extraction process can be as follows:

[0047] 1. Pre-emphasis: In order to compensate for the high-frequency attenuation in the speech signal, the speech signal is pre-emphasized.

[0048] 2. Framing: Dividing a continuous speech signal into multiple short frames (speech frames), with frames overlapping by up to 50%.

[0049] 3. Windowing: Apply a window function (such as a Hamming window) to each frame to reduce the impact of frame boundaries.

[0050] 4. Fast Fourier Transform (FFT): Perform FFT on the signal after the window function to obtain the spectral information.

[0051] 5. Mel Filter Bank: This filter passes the spectrum through a set of Mel filters. The Mel frequency scale is similar to the frequency perceived by the human ear. The output of each filter is the energy of the signal in that frequency band.

[0052] 6. Logarithmic operation: Take the logarithm of the filter output to obtain the energy.

[0053] 7. Discrete Cosine Transform (DCT): Perform DCT on the filter output to obtain the MFCC coefficients.

[0054] 8. Retain low-order coefficients: Usually only the first L MFCC coefficients are retained as feature vectors.

[0055] In one embodiment, the feature extraction sampling window is used for the framing operation in the extraction process described above. Each short frame corresponds to a time window, which is the feature extraction sampling window. The sliding step size is the frame shift, which can be understood as the time interval between consecutive windows. The window length of the feature extraction sampling window can be set to 20-30ms (e.g., 25ms). The sliding step size of the feature extraction sampling window (i.e., the frame shift for MFCC feature extraction) can be determined based on the speech rate, so that the speech feature sequence can achieve dynamic adaptive sampling based on the speech rate.

[0056] In one embodiment, the sliding step size can be inversely proportional to the speech rate. That is, the faster the speech rate, the smaller the sliding step size of the feature extraction sampling window and the larger the frame sampling density. This allows for denser sampling in fast-paced scenarios to ensure feature integrity and capture transient / fine-grained features, which is beneficial for improving the synchronization accuracy of digital lip reading and reducing response latency. Conversely, the slower the speech rate, the larger the sliding step size of the feature extraction sampling window and the smaller the frame sampling density. This allows for sparser sampling in slow-paced scenarios to improve the stability of lip reading recognition and reduce redundant frames, which is beneficial for improving the stability, continuity, and fluency of digital lip reading and reducing computational load.

[0057] According to one embodiment of this disclosure, a Large Language Model (LLM) can be used to perform emotion recognition on speech data to obtain an emotion category sequence. The emotion category sequence may include multiple emotion category labels arranged in chronological order. As an example, the emotion category labels can be preset, such as including but not limited to "happy, neutral, disgusted, sad, angry, grief, surprised," etc. For example, the content of the speech data is "I am really happy today," where "I" corresponds to "neutral," "today" corresponds to "neutral," "really" corresponds to "neutral," "very" corresponds to "happy," and "happy" corresponds to "happy." The emotion category sequence can be [neutral, neutral, neutral, happy, happy].

[0058] In one embodiment, an emotion category sequence and a speech feature sequence can be input into a facial driving parameter generation model to output a facial driving parameter sequence. This sequence can include multiple facial driving parameters arranged in temporal order. For example, facial driving parameters may include the weights of a BlendShape. Optionally, to reduce the complexity of the facial model while maintaining its naturalness, symmetrical processing can be applied to some of the facial motion units (BlendShape meshes) to ensure the naturalness of the driving effect. For example, the final output might be the weight parameters of 17 facial motion units as driving data.

[0059] In one embodiment, a sequence of limb movements can be determined based on a sequence of emotion categories. For example, corresponding limb drive parameters can be pre-configured for different emotion category labels, and these limb drive parameters characterize limb movements. Limb drive parameters may include, for example, transformation matrices (rotation, translation, scaling) of skeletal joints.

[0060] It is understandable that emotion category sequences can influence the facial expressions and body movements of the target virtual avatar through facial driving parameter sequences and body movement sequences, thereby optimizing the overall emotional expression effect of the target virtual avatar.

[0061] According to one embodiment of this disclosure, the face-driven parameter generation model is trained based on sample speech data. The sample labels of the sample speech data include face-driven parameter labels and phoneme labels. A phoneme is the smallest unit of speech defined based on the natural properties of speech. From an acoustic perspective, a phoneme is the smallest unit of speech defined from the perspective of sound quality. From a physiological perspective, one articulation action forms one phoneme. Exemplarily, the phoneme labels can select labels for specific phonemes that are difficult to identify or easily confused (target phonemes may include, but are not limited to, bilabial consonants ( / b / ), labiodental consonants ( / p / ), and bilabial nasal consonants ( / m / ), etc.) to help the face-driven parameter generation model focus on specific phonemes and improve the model's recognition accuracy and lip-sync accuracy for these specific phonemes.

[0062] It should be noted that the emotion category sequence, facial driving parameter sequence, and body movement sequence can be unified onto the same timeline based on the timestamp, so that they are aligned with each other in time sequence, thus ensuring that the final generated target virtual image is synchronized with the audio and video.

[0063] In one embodiment, the facial model of the target virtual avatar can be driven by a sequence of facial driving parameters, and the limb model of the target virtual avatar can be driven by a sequence of limb movements to generate the target virtual avatar.

[0064] According to embodiments of this disclosure, by extracting features from speech data using a feature extraction sampling window with a sliding step size determined by speech rate, dynamic adaptive sampling based on speech rate can be achieved for speech feature sequences. The sliding step size of the feature extraction sampling window (i.e., the frame shift for MFCC feature extraction) can be inversely proportional to the speech rate. This allows for denser sampling in scenarios with fast speech rates to ensure feature integrity and capture transient / fine-grained features, which is beneficial for improving the synchronization accuracy of digital lip reading and reducing response latency. Conversely, it allows for sparser sampling in scenarios with slow speech rates to improve the stability of lip reading recognition and reduce redundant frames, which is beneficial for improving the stability, continuity, and fluency of digital lip reading and reducing computational load. The emotion category sequence and speech feature sequence can be input into a facial driving parameter generation model, which outputs a facial driving parameter sequence. A body movement sequence can be determined based on the emotion category sequence. Therefore, the emotion category sequence can influence the facial expressions and body movements of the target virtual avatar through the facial driving parameter sequence and the body movement sequence, thereby optimizing the overall emotional expression effect of the target virtual avatar. The sample data labels of the face-driven parameter generation model include specific phoneme labels, which can help the face-driven parameter generation model focus on specific phonemes and improve the recognition accuracy and lip-sync accuracy of the face-driven parameter generation model for the above-mentioned specific phonemes.

[0065] Figure 3 The diagram illustrates the technical architecture of a virtual avatar generation method according to an embodiment of the present disclosure.

[0066] like Figure 3 As shown, for the input speech data 301, a large model can be used to perform emotion recognition on the audio data 301 to obtain an emotion category sequence 302. Features can be extracted from the speech data 301 based on a feature extraction sampling window whose sliding step size is determined using speech rate, resulting in a speech feature sequence 303.

[0067] The emotion category sequence 302 and the speech feature sequence 303 can be input into the facial driving parameter generation model 304, and the output is a facial driving parameter sequence 305. A body movement sequence 306 can be determined based on the emotion category sequence 302.

[0068] The facial model of the target virtual avatar can be driven using the facial driving parameter sequence 305, and the limb model of the target virtual avatar can be driven using the limb motion sequence 306 to generate the target virtual avatar. For example, the display method of the target virtual avatar may include holographic display, 2D display, 3D display, etc.

[0069] According to embodiments of this disclosure, feature extraction of speech data based on a feature extraction sampling window that uses speech rate to determine the sliding step size, to obtain a speech feature sequence, includes: performing detrending processing on the speech data to obtain intermediate speech data; segmenting the intermediate speech data based on a preset segmentation window and a preset segmentation step size to obtain multiple initial speech frame data arranged in time sequence, each initial speech frame data having the same frame length; determining multiple initial speech frame data sequences based on the multiple initial speech frame data, each initial speech frame data sequence including multiple initial speech frame data; and performing speech recognition on any initial speech frame data in the initial speech frame data sequence using a speech recognition model to obtain a speech feature sequence that matches the initial speech frame data. The speech frame data sequence is associated with speech frame text, and the number of characters corresponding to the speech frame text is determined; the number of characters is used to represent the speech rate; the sliding step size of the feature extraction sampling window for the initial speech frame data sequence is determined based on the number of characters, resulting in the sliding step size corresponding to each of the multiple initial speech frame data sequences; wherein, the sliding step size is inversely proportional to the number of characters; the window size of the feature extraction sampling window is less than the window size of the preset segmentation window; feature extraction is performed on the multiple initial speech frame data based on the feature extraction sampling window and the sliding step size, resulting in the speech features corresponding to each of the multiple initial speech frame data sequences; and a speech feature sequence is determined based on the multiple speech features corresponding to each of the multiple initial speech frame data sequences.

[0070] According to one embodiment of this disclosure, the speech data can first undergo detrending processing to reduce the impact of changes in speech volume, resulting in intermediate speech data. The intermediate speech data can then be segmented based on a preset segmentation window and a preset segmentation step size to obtain multiple initial speech frames arranged in time sequence, each with the same frame length (e.g., 520ms). Those skilled in the art can reasonably set the window length and sliding step size (i.e., the preset segmentation step size) of the preset segmentation window according to actual needs or application scenarios; no limitations are imposed here.

[0071] In one embodiment, 10400ms of speech data can be segmented to obtain 20 initial speech frames of 520ms each. These 20 initial speech frames can be arranged in temporal order, or they can be divided into multiple initial speech frame sequences in temporal order. For example, the 20 initial speech frames can be arranged in temporal order as VF1, VF2, ..., VF19, VF20. Alternatively, four initial speech frame sequences can be determined by grouping them into sets of five: [VF1, VF2, VF3, VF4, VF5], [VF6, VF7, VF8, VF9, VF10], [VF11, VF12, VF13, VF14, VF15], and [VF16, VF17, VF18, VF19, VF20].

[0072] For example, a speech recognition model (such as an ASR (Automatic Speech Recognition) model) can be used to perform speech recognition on any initial speech frame data in the initial speech frame data series to obtain the speech frame text. The number of characters n corresponding to this speech frame text can then be determined. This number of characters n can determine the sliding step s of the feature extraction sampling window (MFCC window) used for the initial speech frame data sequence. For instance, if the number of characters corresponding to the four initial speech frame data sequences mentioned above are n1, n2, n3, and n4 respectively, for computational efficiency and cost considerations, the sliding step s1 of the MFCC window corresponding to VF1, VF2, VF3, VF4, and VF5 can be determined based on n1, and the sliding step s2 of the MFCC window corresponding to VF6, VF7, VF8, VF9, and VF10 can be determined based on n2, and so on. It is understandable that the frame length of each initial speech frame is 520ms. The larger the number of characters n, the faster the speech rate, and the smaller the number of characters n, the slower the speech rate.

[0073] In one embodiment, the window length of the MFCC window can be set to, for example, 25ms. The sliding step size s of the MFCC window for the initial speech frame data sequence can be determined based on the number of characters n, resulting in the sliding step size corresponding to each of the multiple initial speech frame data sequences. For example, the sliding step size of the MFCC window can be dynamically adjusted based on the number of characters n (representing speech rate) corresponding to the 520ms initial speech frame data, thereby achieving dynamic adjustment of the MFCC sampling frequency. If the audio bit rate is b and the algorithm output frame rate is f, the sliding step size s between consecutive MFCC windows can be determined, for example, in the following way:

[0074]

[0075] According to embodiments of this disclosure, by determining the sliding step size s of the feature extraction sampling window (MFCC window) for the initial speech frame data sequence based on the number of characters n, the MFCC can have a better dynamic adjustment effect for sentences with rapidly changing speech rates. For sentences with fast speech rates, the sliding step size of the MFCC window will be smaller to increase the dynamic effect of recognition and also improve the dynamic recognition effect of mouth shapes. For sentences with slow speech rates, the sliding step size of the MFCC window will be larger to increase the stability of recognition and make the stability and continuity of mouth shapes better.

[0076] In one embodiment, features can be extracted from the 20 initial speech frames based on the MFCC window and the corresponding sliding step size s, respectively, to obtain the MFCC features corresponding to each of the 20 initial speech frames. Optionally, the extracted MFCC features can also be normalized using the CMVN (Cepstral Mean Variance Normalization) algorithm to obtain speech features. Multiple speech features can be arranged in temporal order to obtain a speech feature sequence.

[0077] In one embodiment, the main idea of ​​CMVN is to normalize the cepstral coefficients by subtracting their mean from each coefficient and dividing by the standard deviation. This eliminates scale and offset differences between different speech samples, making them comparable in the cepstral space. The purpose of CMVN is to normalize the cepstral coefficients of an audio signal to reduce variations between different speakers and recording conditions, and to improve the comparability and robustness of features.

[0078] For example, each MFCC window extracts 64×13 features. Then, the velocity and acceleration of these 64×13 signals are calculated, resulting in 64×39 signals. These 64×39 signals are then normalized to -1 to 1 using the CMVN algorithm. The processing effect is as follows: Figure 4 As shown. An example of the CMVN normalization formula is shown below:

[0079]

[0080]

[0081]

[0082] Where n is the characteristic number, Let be the mean of the j-th feature. Let be the standard deviation of the j-th feature. This is the result of normalization.

[0083] According to embodiments of this disclosure, the emotion category sequence is obtained by performing emotion recognition on semantic text sequences and / or speech data based on a large model to obtain the emotion category sequence; wherein, the semantic text sequence is determined by performing semantic recognition on speech data based on a speech recognition model.

[0084] According to one embodiment of this disclosure, a speech recognition model can be used to perform semantic recognition on speech data to obtain semantic text. The semantic text may include multiple text segments (such as words), and these multiple text segments can be arranged in chronological order to determine a semantic text sequence. A large model can be used to perform emotion recognition on the semantic text sequence to obtain multiple emotion category labels, and an emotion category sequence can be determined based on these multiple emotion category labels.

[0085] According to another embodiment of this disclosure, a large model (such as a large speech model) can be used to perform speech emotion recognition (SER) on speech data to obtain multiple emotion category labels, and an emotion category sequence can be determined based on the multiple emotion category labels. Speech emotion recognition (SER) refers to the technique of identifying the speaker's emotional state through the acoustic features of a speech segment (these features are independent of the speech's content and language information).

[0086] According to another embodiment of this disclosure, a speech recognition model can be used to perform semantic recognition on speech data to obtain semantic text. The semantic text may include multiple text segments (such as words), and these segments can be arranged in chronological order to determine a semantic text sequence. A large speech model can be used to perform speech emotion recognition on the speech data to obtain a first emotion category sequence. The large model can then be used to perform emotion recognition on the semantic text sequence to obtain a second emotion category sequence. The first and second emotion category sequences can be aligned according to their timestamps and input together into the large model to output the final emotion category sequence.

[0087] According to embodiments of this disclosure, the face-driving parameter generation model includes a speech-driven face network and a long short-term memory network. The speech-driven face network includes a frequency analysis network, an articulation network, and an output network. The model inputs an emotion category sequence and a speech feature sequence obtained from emotion recognition of speech data into the face-driving parameter generation model, and outputs a face-driving parameter sequence by: encoding the emotion category sequence to obtain an emotion vector sequence; extracting features from the speech feature sequence based on the frequency analysis network to obtain an intermediate speech feature sequence; extracting features from the intermediate speech feature sequence and the emotion vector sequence based on the articulation network to obtain a speech feature map sequence; mapping the speech feature map sequence to a first face-driving parameter sequence based on the output network; inputting the speech feature sequence into the long short-term memory network to output a second face-driving parameter sequence; and fusing the first and second face-driving parameter sequences to obtain the face-driving parameter sequence.

[0088] Figure 5 A flowchart illustrating a virtual avatar generation method according to an embodiment of the present disclosure is shown schematically.

[0089] like Figure 5 As shown, the face-driven parameter generation model may include a voice-driven face network M510 (e.g., optionally an Audio2Face network) and a long short-term memory network M520 (LSTM). The voice-driven face network M510 may include a formal analysis network, an articulation network, and an output network.

[0090] In one embodiment, emotion recognition can be performed on speech data 501 to obtain an emotion category sequence 502. Features can be extracted from speech data 501 using a feature extraction sampling window that determines the sliding step size based on speech rate to obtain a speech feature sequence 503. The emotion category sequence 502 can be encoded to obtain an emotion vector sequence 504, for example, using one-hot encoding. Features can be extracted from speech feature sequence 503 using a frequency analysis network to obtain an intermediate speech feature sequence. Features can be extracted from the intermediate speech feature sequence and emotion vector sequence 504 using a pronunciation network to obtain a speech feature map sequence. The speech feature map sequence can be mapped to a first facial driving parameter sequence 505 using an output network.

[0091] For example, the speech feature sequence 503 can be input into a Long Short-Term Memory (LSTM) network M520, which outputs a second facial driving parameter sequence 506. LSTM can optimize the continuity of the model output. LSTM selectively retains or forgets information through input, forget, and output gates, maintaining the flow of information in long sequences, thus enabling the network to remember the input audio features. The LSTM structure allows the network to capture and understand complex dependencies in long sequences, thereby improving lip-sync coherence.

[0092] For example, the first facial driving parameter sequence 505 and the second facial driving parameter sequence 506 can be fused to obtain the facial driving parameter sequence 507. The fusion method can be selected as weighted summation, for example.

[0093] For example, a body movement sequence 508 can be determined based on an emotion category sequence 502. The facial model of the target virtual avatar can be driven using a facial driving parameter sequence 507, and the body movement sequence 508 can be used to drive the body model of the target virtual avatar, generating the target virtual avatar 509 (the target virtual avatar in...). Figure 5 (Presented in the form of multiple video frames).

[0094] According to embodiments of this disclosure, the face-driven parameter generation model is obtained through the following operations: performing emotion recognition on sample speech data to obtain a sample emotion category sequence; performing feature extraction on the sample speech data based on a feature extraction sampling window that uses the sample speech rate to determine the sliding step size to obtain a sample speech feature sequence; and training a pre-trained face-driven parameter generation model based on the sample speech feature sequence, the sample emotion category sequence, and the sample labels of the sample speech data to obtain the face-driven parameter generation model, wherein the sample labels include face-driven parameter labels and phoneme labels.

[0095] According to one embodiment of this disclosure, the sample labels for sample speech data may include face-driven parameter labels and phoneme labels (such as labels for specific phonemes mentioned above). The structure of the pre-trained face-driven parameter generation model can be referred to the face-driven parameter generation model described above, and will not be repeated here.

[0096] For example, a large model can be used to perform emotion recognition on sample speech data to obtain a sequence of sample emotion categories. The method for determining the emotion category sequence described above will not be repeated here. Features can be extracted from the sample speech data using a feature extraction sampling window that determines the sliding step size based on the sample speech rate, resulting in a sequence of sample speech features. The method for determining the speech feature sequence described above will not be repeated here. A pre-trained face-driven parameter generation model can be trained based on the sample speech feature sequence, the sample emotion category sequence, and the sample labels of the sample speech data to obtain a face-driven parameter generation model.

[0097] According to embodiments of this disclosure, training a pre-trained face-driven parameter generation model based on sample speech feature sequences, sample emotion category sequences, and sample labels of sample speech data to obtain the face-driven parameter generation model includes: inputting the sample speech feature sequences and sample emotion category sequences into the pre-trained face-driven parameter generation model to obtain a face-driven parameter prediction sequence; determining a first loss function value based on a first preset loss function term, according to the face-driven parameter prediction sequence and face-driven parameter labels; determining a second loss function value based on a second preset loss function term, according to preset weight hyperparameters and the first loss function value, wherein the preset weight hyperparameters are used to characterize the degree of attention the pre-trained face-driven parameter generation model pays to sample speech data with phoneme labels; determining a target loss function value based on the first loss function value and the second loss function value; and adjusting the model parameters of the pre-trained face-driven parameter generation model according to the target loss function value until a predetermined termination condition is met to obtain the face-driven parameter generation model.

[0098] According to one embodiment of this disclosure, sample speech feature sequences and sample emotion category sequences can be input into a pre-trained face-driven parameter generation model to obtain a face-driven parameter prediction sequence. The face-driven parameter prediction sequence can be obtained by referring to the method for determining the face-driven parameter sequence described above, and will not be repeated here. As an example, for specific phonemes such as / b / , / p / , and / m / , preset weight hyperparameters can be added to the loss function equation during training to enable the network to focus on learning the mouth shape of specific phonemes, thereby improving the model's accuracy for specific pronunciation mouth shapes.

[0099] As an example, the loss function equation is shown below:

[0100]

[0101] in, The target loss function value; The first preset loss function value represents the predicted sequence of facial driving parameters ( ) and facial driving parameter labels ( The mean square error (MSE) of ). To preset the weight hyperparameters, This is the value of the second loss function.

[0102] The preset weight hyperparameters are used to characterize the degree of attention the pre-trained face-driven parameter generation model pays to sample speech data with phoneme labels. When the model is backpropagating, the samples with phoneme labels will generate a larger gradient backpropagation, so that the network can focus on learning the mouth shape of specific phonemes.

[0103] In one embodiment, the model parameters of the pre-trained face-driving parameter generation model can be adjusted according to the target loss function value until a predetermined termination condition is met (a predetermined number of training rounds or convergence of the loss function) to obtain the face-driving parameter generation model.

[0104] Figure 6 A schematic block diagram of a virtual avatar generation apparatus according to an embodiment of the present disclosure is shown.

[0105] like Figure 6 As shown, the virtual avatar generation device 600 of this embodiment includes a feature extraction module 610, a first processing module 620, a second processing module 630, and a driving module 640.

[0106] The feature extraction module 610 is used to extract features from speech data based on the feature extraction sampling window whose sliding step size is determined by the speech rate, and obtain a speech feature sequence.

[0107] The first processing module 620 is used to input the emotion category sequence and speech feature sequence obtained by emotion recognition of speech data into the face driving parameter generation model and output the face driving parameter sequence. The face driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face driving parameter labels and phoneme labels.

[0108] The second processing module 630 is used to determine the body movement sequence based on the emotion category sequence, wherein the emotion category sequence, the facial driving parameter sequence, and the body movement sequence correspond in time.

[0109] The driving module 640 is used to drive the facial model of the target virtual image using a facial driving parameter sequence and to drive the limb model of the target virtual image using a limb movement sequence, thereby generating the target virtual image.

[0110] According to embodiments of this disclosure, the emotion category sequence is obtained by performing emotion recognition on semantic text sequences and / or speech data based on a large model to obtain the emotion category sequence; wherein, the semantic text sequence is determined by performing semantic recognition on speech data based on a speech recognition model.

[0111] According to embodiments of this disclosure, the feature extraction module 610 may include a first processing submodule, a second processing submodule, a third processing submodule, a fourth processing submodule, a fifth processing submodule, a sixth processing submodule, and a seventh processing submodule.

[0112] The first processing submodule is used to perform detrending processing on the speech data to obtain intermediate speech data.

[0113] The second processing submodule is used to segment the intermediate speech data based on a preset segmentation window and a preset segmentation step size to obtain multiple initial speech frame data arranged in time sequence, with each initial speech frame data having the same frame length.

[0114] The third processing submodule is used to determine multiple initial speech frame data sequences based on multiple initial speech frame data, each of which includes multiple initial speech frame data.

[0115] The fourth processing submodule is used to perform speech recognition on any initial speech frame data in the initial speech frame data sequence through a speech recognition model, obtain the speech frame text associated with the initial speech frame data sequence, and determine the number of characters corresponding to the speech frame text; the number of characters is used to represent the speech rate.

[0116] The fifth processing submodule is used to determine the sliding step size of the feature extraction sampling window for the initial speech frame data sequence based on the number of characters, and to obtain the sliding step size corresponding to each of the multiple initial speech frame data sequences; wherein, the sliding step size is inversely proportional to the number of characters; the window length of the feature extraction sampling window is less than the window length of the preset segmentation window.

[0117] The sixth processing submodule is used to extract features from multiple initial speech frame data based on the feature extraction sampling window and the sliding step size, so as to obtain the speech features corresponding to each of the multiple initial speech frame data.

[0118] The seventh processing submodule is used to determine the speech feature sequence based on the multiple speech features corresponding to each of the multiple initial speech frame data sequences.

[0119] According to embodiments of this disclosure, the facial driving parameter generation model includes a voice-driven facial network and a long short-term memory network. The voice-driven facial network includes a frequency analysis network, an articulation network, and an output network. The first processing module includes an eighth, ninth, tenth, eleventh, twelfth, and thirteenth processing submodules.

[0120] The eighth processing submodule is used to encode the emotion category sequence to obtain an emotion vector sequence.

[0121] The ninth processing submodule is used to extract features from the speech feature sequence based on the frequency analysis network to obtain the intermediate speech feature sequence.

[0122] The tenth processing submodule is used to extract features from the intermediate speech feature sequence and emotion vector sequence based on the pronunciation network to obtain a speech feature map sequence.

[0123] The eleventh processing submodule is used to map the speech feature map sequence into the first face driving parameter sequence based on the output network.

[0124] The twelfth processing submodule is used to input the speech feature sequence into the long short-term memory network and output the second face driving parameter sequence.

[0125] The thirteenth processing submodule is used to fuse the first facial driving parameter sequence and the second facial driving parameter sequence to obtain the facial driving parameter sequence.

[0126] According to embodiments of this disclosure, any plurality of modules among the feature extraction module 610, the first processing module 620, the second processing module 630, and the driving module 640 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the feature extraction module 610, the first processing module 620, the second processing module 630, and the driving module 640 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the feature extraction module 610, the first processing module 620, the second processing module 630, and the driving module 640 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0127] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a virtual avatar generation method according to an embodiment of the present disclosure.

[0128] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0129] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0130] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.

[0131] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method provided according to the embodiments of this disclosure.

[0132] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.

[0133] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the virtual avatar generation method provided in the embodiments of this disclosure.

[0134] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0135] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0136] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0137] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0138] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0139] Although this disclosure has been shown and described with reference to specific exemplary embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made to this disclosure without departing from the spirit and scope of the disclosure as defined by the appended claims and their equivalents. Therefore, the scope of this disclosure should not be limited to the above embodiments, but should be defined not only by the appended claims, but also by their equivalents.

Claims

1. A method for generating a virtual avatar, comprising: Based on the feature extraction sampling window that uses speech rate to determine the sliding step size, feature extraction is performed on speech data to obtain a speech feature sequence; The emotion category sequence obtained by emotion recognition of the speech data and the speech feature sequence are input into the face driving parameter generation model, and the face driving parameter sequence is output. The face driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face driving parameter labels and phoneme labels. A body movement sequence is determined based on the emotion category sequence, wherein the emotion category sequence, the facial driving parameter sequence, and the body movement sequence correspond in time. The facial model of the target virtual image is driven by the facial driving parameter sequence, and the limb model of the target virtual image is driven by the limb movement sequence to generate the target virtual image.

2. The method according to claim 1, wherein, The emotion category sequence is obtained using the following operation: Emotion recognition is performed on the semantic text sequence and / or the speech data based on a large model to obtain the emotion category sequence; wherein the semantic text sequence is determined based on the semantic recognition of the speech data using a speech recognition model.

3. The method according to claim 1, wherein, The feature extraction sampling window, which uses speech rate to determine the sliding step size, is used to extract features from the speech data, resulting in a speech feature sequence including: The speech data is detrended to obtain intermediate speech data; The intermediate speech data is segmented based on a preset segmentation window and a preset segmentation step size to obtain multiple initial speech frame data arranged in time sequence, and each initial speech frame data has the same frame length. Multiple initial speech frame data sequences are determined based on multiple initial speech frame data, each of the initial speech frame data sequences including multiple initial speech frame data; The speech recognition model is used to perform speech recognition on any initial speech frame data in the initial speech frame data sequence to obtain the speech frame text associated with the initial speech frame data sequence, and the number of characters corresponding to the speech frame text is determined; the number of characters is used to characterize the speech rate. The sliding step size for the feature extraction sampling window used for the initial speech frame data sequence is determined based on the number of characters, thereby obtaining the sliding step size corresponding to each of the multiple initial speech frame data sequences; wherein, the sliding step size is inversely proportional to the number of characters; the window length of the feature extraction sampling window is smaller than the window length of the preset segmentation window; Based on the feature extraction sampling window and the sliding step size, feature extraction is performed on multiple initial speech frame data to obtain the speech features corresponding to each of the multiple initial speech frame data; and The speech feature sequence is determined based on multiple speech features corresponding to each of the multiple initial speech frame data sequences.

4. The method according to claim 1, wherein, The facial driving parameter generation model includes a speech-driven facial network and a long short-term memory network. The speech-driven facial network includes a frequency analysis network, an articulation network, and an output network. The emotion category sequence and the speech feature sequence obtained by emotion recognition of the speech data are input into the facial driving parameter generation model, and the output facial driving parameter sequence includes: The emotion category sequence is encoded to obtain an emotion vector sequence; Based on the frequency analysis network, feature extraction is performed on the speech feature sequence to obtain an intermediate speech feature sequence; Based on the pronunciation network, feature extraction is performed on the intermediate speech feature sequence and the emotion vector sequence to obtain a speech feature map sequence; Based on the output network, the speech feature map sequence is mapped into a first face driving parameter sequence; The speech feature sequence is input into the Long Short-Term Memory network, which outputs a second face driving parameter sequence; and The first facial driving parameter sequence and the second facial driving parameter sequence are fused to obtain the facial driving parameter sequence.

5. The method according to any one of claims 1-4, wherein, The facial driving parameter generation model is obtained through the following operations: Emotion recognition is performed on the sample speech data to obtain a sequence of sample emotion categories; Based on the feature extraction sampling window that uses the sample speech rate to determine the sliding step size, feature extraction is performed on the sample speech data to obtain the sample speech feature sequence. Based on the sample speech feature sequence, the sample emotion category sequence, and the sample labels of the sample speech data, the pre-trained face-driven parameter generation model is trained to obtain the face-driven parameter generation model, wherein the sample labels include face-driven parameter labels and phoneme labels.

6. The method according to claim 5, wherein, The process of training a pre-trained facial driving parameter generation model based on the sample speech feature sequence, the sample emotion category sequence, and the sample labels of the sample speech data to obtain the facial driving parameter generation model includes: The sample speech feature sequence and the sample emotion category sequence are input into the pre-trained face driving parameter generation model to obtain the face driving parameter prediction sequence. Based on the first preset loss function term, the value of the first loss function is determined according to the predicted sequence of the facial driving parameters and the label of the facial driving parameters; Based on the second preset loss function term, the second loss function value is determined according to the preset weight hyperparameter and the first loss function value, wherein the preset weight hyperparameter is used to characterize the degree of attention the pre-trained face-driven parameter generation model pays to the sample speech data with the phoneme label; Based on the first loss function value and the second loss function value, determine the target loss function value; and The model parameters of the pre-trained facial driving parameter generation model are adjusted according to the target loss function value until a predetermined termination condition is met, thus obtaining the facial driving parameter generation model.

7. A virtual avatar generation device, comprising: The feature extraction module is used to extract features from speech data based on the feature extraction sampling window whose sliding step size is determined by the speech rate, and obtain speech feature sequences. The first processing module is used to input the emotion category sequence obtained by emotion recognition of the speech data and the speech feature sequence into the face driving parameter generation model, and output the face driving parameter sequence. The face driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face driving parameter labels and phoneme labels. The second processing module is used to determine the body movement sequence based on the emotion category sequence, wherein the emotion category sequence, the facial driving parameter sequence, and the body movement sequence correspond in time. The driving module is used to drive the facial model of the target virtual image using the facial driving parameter sequence and to drive the limb model of the target virtual image using the limb movement sequence, thereby generating the target virtual image.

8. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic is that the one or more processors execute the one or more computer programs to perform the following operations: Based on the feature extraction sampling window that uses speech rate to determine the sliding step size, feature extraction is performed on speech data to obtain a speech feature sequence; The emotion category sequence obtained by emotion recognition of the speech data and the speech feature sequence are input into the face driving parameter generation model, and the face driving parameter sequence is output. The face driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face driving parameter labels and phoneme labels. A body movement sequence is determined based on the emotion category sequence, wherein the emotion category sequence, the facial driving parameter sequence, and the body movement sequence correspond in time. The facial model of the target virtual image is driven by the facial driving parameter sequence, and the limb model of the target virtual image is driven by the limb movement sequence to generate the target virtual image.

9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by the processor, they perform the following operations: Based on the feature extraction sampling window that uses speech rate to determine the sliding step size, feature extraction is performed on speech data to obtain a speech feature sequence; The emotion category sequence obtained by emotion recognition of the speech data and the speech feature sequence are input into the face driving parameter generation model, and the face driving parameter sequence is output. The face driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face driving parameter labels and phoneme labels. A body movement sequence is determined based on the emotion category sequence, wherein the emotion category sequence, the facial driving parameter sequence, and the body movement sequence correspond in time. The facial model of the target virtual image is driven by the facial driving parameter sequence, and the limb model of the target virtual image is driven by the limb movement sequence to generate the target virtual image.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they perform the following operations: Based on the feature extraction sampling window that uses speech rate to determine the sliding step size, feature extraction is performed on speech data to obtain a speech feature sequence; The emotion category sequence obtained by emotion recognition of the speech data and the speech feature sequence are input into the face driving parameter generation model, and the face driving parameter sequence is output. The face driving parameter generation model is trained based on sample speech data, and the sample labels of the sample speech data include face driving parameter labels and phoneme labels. A body movement sequence is determined based on the emotion category sequence, wherein the emotion category sequence, the facial driving parameter sequence, and the body movement sequence correspond in time. The facial model of the target virtual image is driven by the facial driving parameter sequence, and the limb model of the target virtual image is driven by the limb movement sequence to generate the target virtual image.