Video emotion virtual image real-time generation method based on audio driving and mixed shape
Through the video emotional virtual image generation method based on audio-driven and mixed shapes, the 3D Gaussian mixed shapes and cross-attention mechanism is used to solve the problems of low generation efficiency and insufficient emotional intensity editing ability in the prior art, and efficient and rich emotional virtual image generation is achieved.
Patent Information
- Application Number
- CN202511004068.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-21
AI Technical Summary
When generating audio-driven video emotional virtual images, the prior art lacks the ability to edit emotional intensity, and the generation efficiency is low, making it difficult to adapt to the generalization between different emotional expressions.
Using a real-time generation method of video emotional virtual images based on audio-driven and mixed shapes, a neutral emotional space is built through 3D Gaussian mixed shapes, combining emotions, face grids and audio-driven deformation, virtual images are rendered using 3D Gaussian sputtering technology, and a cross-attention mechanism is introduced to control Gaussian attribute deformation.
Real-time generation of virtual character avatars of different audio, different emotions and different intensities is realized, improving the generation efficiency and richness of emotional expression, and making it more adaptable.
Smart Images

Figure CN120510259A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human portrait generation in computer vision, and in particular to a real-time generation method of a video emotion virtual image based on audio drive and mixed shapes. Background Art
[0002] The importance of animated 3D virtual human heads has increased significantly in virtual reality and related applications, where audio-driven character animation plays a key role in various fields such as human-computer interaction, digital humans, filmmaking, and virtual video conferencing.
[0003] With the advent of generative adversarial networks (GANs), early methods either directly learned a mapping from audio to video frames or used intermediate representations (e.g., facial landmarks) to connect audio input with video output. However, these methods primarily focused on addressing lip syncing and video quality issues, with limited research on generating emotionally rich videos. Recently, research on emotion-driven virtual head generation has gained traction. For example, some methods use one-shot encoded emotion labels as the source of emotional data, while EDTalk decouples lip movements, head pose, and emotion to adapt to new emotional expressions. However, these methods lack the ability to edit emotional intensity. With the introduction of the multi-intensity emotion dataset, MEAD, several new methods have emerged that enable editing emotional intensity. For example, EAT constructs an emotion deformation network and an emotion adaptation module to predict changes in 3D landmarks to generate emotionally expressive faces. However, these methods rely on the accuracy of external detectors and suffer from poor generalization across different emotional expressions. To address these limitations, EMOdiffhead leverages the DECA method to extract facial expression vectors, combines them with audio input, and guides a diffusion model to generate videos with accurate lip sync and rich emotional expressions. However, due to the inherent nature of diffusion models, these methods face challenges related to generation efficiency, which is crucial for real-world applications.
[0004] In summary, in order to better develop downstream applications of virtual reality, human-computer interaction, and video portraits, a real-time generation method of audio-driven video emotional virtual images is urgently needed. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for real-time generation of video emotion virtual images based on audio drive and mixed shapes, so as to realize real-time generation of virtual character avatars with different audio, different emotions and different intensities.
[0006] To achieve the object of the present invention, the present invention provides a method for real-time generation of a video emotional avatar based on audio drive and mixed shapes, comprising the following steps:
[0007] (1) Obtain data: a video of a person speaking, a basic face model, and a set of expression blend shapes;
[0008] (2) Each expression blend shape is represented by position, opacity, rotation, and scale attributes; the expression parameters corresponding to each frame of the video are fitted using 3D Gaussian sputtering technology for training, and the neutral emotion Gaussian model of the face is obtained after training;
[0009] (3) Using the A2ET model to obtain accurate emotional expression parameters based on the input source face image, audio, and emotion label;
[0010] (4) Using the LBS function, the expression parameters obtained in step (3) are converted into the corresponding emotional face mesh;
[0011] (5) The emotional face mesh, emotion category, audio, and emotion intensity are input into the spatial-audio-emotional attention module to predict the offset of each Gaussian attribute;
[0012] (6) Add the offset obtained in step (5) to the neutral emotion Gaussian model of the face in step (2) to obtain the Gaussian attributes of the specified emotion, audio, and emotion intensity, and then use the 3D Gaussian sputtering technology to render the virtual emotional portrait.
[0013] An electronic device includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the method for real-time generation of a video emotion virtual image based on audio drive and mixed shapes is implemented.
[0014] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned method for real-time generation of a video emotion virtual image based on audio drive and mixed shapes.
[0015] A computer program product includes a computer program, which, when executed by a processor, implements the above-mentioned method for real-time generation of a video emotion virtual image based on audio drive and mixed shapes.
[0016] Compared with the existing technology, the significant progress of the present invention lies in: (1) the present invention proposes an audio-driven 3D Gaussian point cloud framework for real-time 3D emotional speech synthesis; (2) 3D Gaussian points are used to construct the state space of neutral and emotional facial expressions respectively, so that the model can adapt to the mapping between different audio and emotional inputs; (3) a cross-attention mechanism is combined to integrate spatial, audio and emotional features to jointly control the deformation of Gaussian attributes.
[0017] In order to more clearly illustrate the functional characteristics and structural parameters of the present invention, further description is given below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0019] Figure 1 It is the overall flow chart of the present invention.
[0020] Figure 2 This is the overall framework diagram of the real-time generation method of emotional virtual images proposed in the present invention.
[0021] Figure 3 This is a graph showing the experimental results of the real-time generation of the emotional virtual image of the present invention.
[0022] Figure 4 This is a graph comparing the experimental results of the present invention with other methods. DETAILED DESCRIPTION
[0023] The present invention provides a real-time generation method for video emotion virtual images based on audio driving and mixed shapes, comprising: first, using 3D Gaussian mixed shapes to construct a neutral emotion space of a specific identity, and synchronizing it with emotion, face mesh and audio-driven deformation; then, encoding 3D Gaussian attributes into a shared implicit feature representation, and fusing them with audio features, mesh displacement coding features, emotion and intensity coding; then, inputting these features into a spatial-audio-emotion attention module to predict the offset of each Gaussian attribute, and finally obtaining a virtual image through Gaussian sputtering rendering; through the above steps, the present invention realizes the real-time generation of virtual character avatars with different audio, different emotions and different intensities.
[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0025] Combine Figure 1 、 Figure 2 A method for real-time generation of video emotion virtual images based on audio drive and blend shapes includes the following steps:
[0026] (1) Obtain data, obtain a video of a person speaking , basic face model , Expression blend shape collection ;
[0027] (2) Each expression blend shape is composed of position , Opacity , Rotate , Zoom Attribute representation: Use 3D Gaussian sputtering technology to fit the expression parameters corresponding to each frame of the video for training, and obtain the neutral emotion Gaussian model of the face after training. ;
[0028] (3) Using the A2ET model to input the source face image , audio , emotion tags To obtain accurate emotional expression parameters ;
[0029] (4) Use the LBS function to convert the expression parameters obtained in step (3) Convert to the corresponding emotional face mesh ;
[0030] (5) Emotional face mesh , emotion categories , audio and emotion intensity T are input into the spatial-audio-emotional attention module to predict the offset of each Gaussian attribute ;
[0031] (6) The offset obtained in step (5) Same as step (2) face neutral emotion Gaussian model Add up the Gaussian attributes of the specified emotion, audio, and emotion intensity, and then use 3D Gaussian sputtering technology to render the virtual emotional portrait .
[0032] Furthermore, in step (1), the basic face model and the Expression blend shape set The acquisition method is obtained through principal component analysis based on the FLAME model. Each model is represented as a set of 3D Gaussian distributions with multiple basic properties, including position , Opacity , Rotate , Zoom ; and The deviation between them is defined as the difference in their Gaussian properties, expressed as ; Arbitrary expressions with neutral emotions on the face The model can be expressed as:
[0033]
[0034] In any expression The face model below; get a neutral emotional expression , through the 3D Gaussian rendering function , we can get the corresponding neutral emotion grid , which is expressed as follows:
[0035]
[0036] Furthermore, the step (2) uses 3D Gaussian sputtering technology to fit the expression parameters corresponding to each frame of the video for training, and after training, a Gaussian model of neutral emotion of the face is obtained. During the training rendering process, the 3D Gaussian distribution is projected onto the 2D plane via the sputtering method. This projection involves a new covariance matrix , which is defined in the camera coordinate system as: ,in is the specified view transformation matrix, is the Jacobian matrix corresponding to the affine approximation of the projective transformation, Represents the 3D Gaussian original covariance matrix. The entire training process is essentially an optimization process of the covariance matrix.
[0037] Furthermore, the step (3) uses the A2ET model to input the source face image , audio , emotion tags To obtain accurate expression parameters , which is expressed as follows:
[0038]
[0039] The A2ET model is an audio-to-emoji converter trained on the massive VoxCeleb2 dataset.
[0040] Furthermore, the step (4) converts the acquired expression parameters into corresponding emotional face meshes using the LBS function, which is expressed as follows:
[0041]
[0042] in, and represents the standard skinning function and joint regressor in FLAME, and and A linear combination of blend shapes representing the generated pose and expression, using animation coefficients and , and the basis vectors of the pose blend shape and the basis vectors of the emotion blend shapes Represents skin weight, which defines the weight of each vertex affected by different joints. It is used to smoothly interpolate vertex positions to achieve natural deformation effects.
[0043] Furthermore, the offset of each Gaussian attribute predicted in step (5) is The deformation network used is composed of several small MLP regressors. For the t-th frame, the final output embedding from the crisscross attention module is The offsets mapped to each property are as follows:
[0044]
[0045] Represents the offset of Gaussian rotation, scale, color and opacity respectively, MLP regressors for Gaussian rotation, scale, color, and opacity, respectively.
[0046] Furthermore, in the spatial audio emotion attention module in step (5), the calculation of the attention score is formulated by the following equation:
[0047]
[0048] in, express The index of Corresponding to the calculated attention score. Indicates the number of features of each Gaussian. In the attention mechanism, and Represents query and key respectively.
[0049] Furthermore, the present invention uses RGB loss, VGG loss and flame loss in the entire training process:
[0050]
[0051] in and are the predicted image and the real image, represents the RGB loss function.
[0052]
[0053] in represents the feature outputs obtained from the first four layers of the pre-trained VGG network, Represents the VGG loss function.
[0054]
[0055] in represents the flame loss function, Indicates the pseudo-true value of the expression parameter, represents the pseudo-true value of the posture parameter, Represents the pseudo-true values of skin weights, which are determined based on the nearest flame vertex. represents the predicted expression parameters, represents the predicted posture parameters, represents the predicted skin weight, Represent the coefficients of expression posture and skin weight respectively, Represents the number of Gaussian points. The final loss is expressed as a combination of D-SSIM terms, which represents the inverse of the structural similarity index and is used to measure the structural difference between two images:
[0056]
[0057] The RGB loss function weight coefficient , flame loss function weight coefficient , VGG loss function weight coefficient , D-SSIM loss function weight coefficient .
[0058] Furthermore, the step (5) converts the emotional face grid , emotion categories , audio and emotion intensity T are input into the spatial-audio-emotional attention module to predict the offset of each Gaussian attribute :
[0059]
[0060] ,
[0061]
[0062]
[0063] The spatial audio emotion attention module consists of multiple cross-attention layers and feedforward layers Each layer is connected by a skip connection, where is the offset between the neutral emotion grid and the emotional grid, and then uses a multi-resolution hash encoding function Processing to improve efficiency, represents the initial spatial features of the nth frame, Indicates the The output features of the layer, Represent emotion category coding and intensity coding, Indicates the Intermediate features of the layer cross attention output; For text encoder, is the number of network layers. In order to predict the offset of each Gaussian attribute , we use a set of multi-layer perceptron (MLP) regressors , as shown below:
[0064]
[0065] Indicates the The output features of the layer, Represents the Gaussian rotation, scale, color and opacity offsets respectively.
[0066] Furthermore, the step (6) converts the offset obtained in step (5) Same as step (2) face neutral emotion Gaussian model Add up the Gaussian attributes of the specified emotion, audio, and emotion intensity, and then use 3D Gaussian sputtering technology to render the virtual emotional portrait , the Gaussian attributes in the final sentiment space are:
[0067]
[0068] , , , Represent the rotation, scale, color and opacity attributes under neutral mood, Represent the final rotation, scale, color and opacity attributes under the specified emotion category, Represents the offset of rotation, scale, color and opacity attributes under the specified emotion category; when all Gaussian parameters in the emotion deformation space are obtained, these parameters are used for rendering; N points are overlapped on a pixel in depth order to form the mixed color of the pixel , as shown below:
[0069] ,
[0070] in Indicates the influence of each Gaussian point on the pixel, represents the transmission term, Represents the Gaussian point set of the pixel, Indicates the color of each Gaussian point under the specified emotion category.
[0071] like Figure 3 As shown, the present invention has been tested under different emotions and intensities. It can be found that under 7 different emotions, the present invention can generate expressions that match the emotions. In addition, there are obvious expression changes between different intensities of the same emotion. Figure 4 As shown, the present invention is compared with different methods, and it can be found that the generated results of the present invention are closer to the true value and have better effects.
[0072] The above describes the specific implementation examples of the present invention in detail. It should be noted that the present invention is not limited to the above specific implementation examples, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A method for real-time generation of video emotional virtual images based on audio drive and mixed shapes, characterized in that: The following steps are involved: (1) Obtain data: a video of a person speaking, a basic face model, and a set of expression blend shapes; (2) Each expression blend shape is represented by position, opacity, rotation, and scale attributes; the expression parameters corresponding to each frame of the video are fitted using 3D Gaussian sputtering technology for training, and the neutral emotion Gaussian model of the face is obtained after training; (3) Using the A2ET model to obtain accurate emotional expression parameters based on the input source face image, audio, and emotion label; (4) Using the LBS function, the expression parameters obtained in step (3) are converted into the corresponding emotional face mesh; (5) The emotional face mesh, emotion category, audio, and emotion intensity are input into the spatial-audio-emotional attention module to predict the offset of each Gaussian attribute; (6) Add the offset obtained in step (5) to the neutral emotion Gaussian model of the face in step (2) to obtain the Gaussian attributes of the specified emotion, audio, and emotion intensity, and then use the 3D Gaussian sputtering technology to render the virtual emotional portrait.
2. The method for real-time generation of a video emotional virtual image based on audio drive and mixed shapes according to claim 1, characterized in that: In step (1), the basic face model and the Expression blend shape set The acquisition method is obtained through principal component analysis based on the FLAME model. Each model is represented as a set of 3D Gaussian distributions with multiple basic properties, including position , Opacity , Rotate , Zoom ; and The deviation between them is defined as the difference in their Gaussian properties, expressed as ; Any expression of neutral emotion on the face The model can be expressed as: ; in, In any expression The face model below; get a neutral emotional expression Afterwards, through the 3D Gaussian rendering function , get the corresponding neutral emotion grid , which is expressed as follows: 。 3. The method for real-time generation of a video emotional virtual image based on audio drive and mixed shapes according to claim 2, characterized in that: Step (2) Use 3D Gaussian sputtering technology to fit the expression parameters corresponding to each frame of the video for training, and obtain the neutral emotion Gaussian model of the face after training During the training rendering process, the 3D Gaussian distribution is projected onto the 2D plane via the sputtering method; this projection involves a new covariance matrix , which is defined in the camera coordinate system as: ,in is the specified view transformation matrix, is the Jacobian matrix corresponding to the affine approximation of the projective transformation, Represents the 3D Gaussian raw covariance matrix.
4. The method for real-time generation of a video emotional virtual image based on audio drive and mixed shapes according to claim 3, characterized in that: Step (3) Use the A2ET model to input the source face image , audio , emotion tags To obtain accurate expression parameters , which is expressed as follows: ; The A2ET model is an audio-to-emoji converter trained on the massive VoxCeleb2 dataset.
5. The method for real-time generation of a video emotional virtual image based on audio drive and mixed shapes according to claim 4, characterized in that: Step (4) uses the LBS function to convert the acquired expression parameters into the corresponding emotional face mesh, which is expressed as follows: ; in, and represents the standard skinning function and joint regressor in FLAME, and and represents a linear combination of blend shapes that generate poses and expressions, Represents skin weight, using animation coefficients and , and the basis vectors of the pose blend shape and the basis vectors of the emotion blend shapes Defines the weights that each vertex has from different joints, used to smoothly interpolate vertex positions.
6. The method for real-time generation of a video emotional virtual image based on audio drive and mixed shapes according to claim 5, characterized in that: Step (5) Emotional face mesh , emotion categories , audio and emotion intensity T are input into the spatial-audio-emotional attention module to predict the offset of each Gaussian attribute : , , , ; The spatial audio emotion attention module consists of multiple cross-attention layers and feedforward layers Each layer is connected by a skip connection, where is the offset between the neutral emotion grid and the emotional grid, and then uses a multi-resolution hash encoding function To process, represents the initial spatial features of the nth frame, Indicates the The output features of the layer, Represent emotion category coding and intensity coding, Indicates the Intermediate features of the layer cross attention output; For text encoder, is the number of network layers; in order to predict the offset of each Gaussian attribute , using a set of multilayer perceptron regressors , as shown below: ; in, Indicates the The output features of the layer, Represents the Gaussian rotation, scale, color and opacity offsets respectively.
7. The method for real-time generation of a video emotional virtual image based on audio drive and mixed shapes according to claim 6, characterized in that: The step (6) converts the offset obtained in step (5) Same as step (2) face neutral emotion Gaussian model Add up the Gaussian attributes of the specified emotion, audio, and emotion intensity, and then use 3D Gaussian sputtering technology to render the virtual emotional portrait , the Gaussian attributes in the final sentiment space are: ; in, , , , Represent the rotation, scale, color and opacity attributes under neutral mood, Represent the final rotation, scale, color and opacity properties under the specified emotion category, Represents the offset of rotation, scale, color and opacity attributes under the specified emotion category; when all Gaussian parameters in the emotion deformation space are obtained, these parameters are used for rendering; N points are overlapped on a pixel in depth order to form the mixed color of the pixel , as shown below: ; in, Indicates the influence of each Gaussian point on the pixel, represents the transmission term, Represents the Gaussian point set of the pixel, Indicates the color of each Gaussian point under the specified emotion category.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Audio visual emotion recognition method based on multi-layer boosted HMM
CN102930298A
Monocular face avatar generation method based on Gaussian point rendering
CN117974867A
Synthetic audiovisual storyteller
US20150042662A1
Semantic deep face models
US20210279956A1
Aluminum alloy sheet with excellent post-fabrication surface qualities and method of manufacturing same
WO2009123011A1
Cited By
Wearable device video live broadcast method and system based on intelligent AI large model driving
CN120916014A
Character image generation method and system based on audio emotion condition modulation
CN121564162A