Multi-person interaction three-dimensional facial animation generation method
Through a dual-branch optimization model and a chorus dataset, the problems of emotional expression complexity and insufficient interaction effects in the generation of singing-driven multi-person interactive 3D facial animation are solved, and the natural interaction effect of singers' facial animation in multi-person chorus scenes is achieved.
Patent Information
- Application Number
- CN202510750253.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies in the generation of singing-driven multi-person interactive 3D facial animation have problems with the complexity of emotional expression and insufficient interactive effects, especially in the multi-person chorus scenario, there is a lack of effective datasets and methods.
A method for generating multi-person interactive 3D facial animation is designed. Chorus video data is obtained for preprocessing, and a dual-branch optimization model is used to learn the mapping relationship between the rhythmic and content features of audio data and facial action sequences. The parametric face model FLAME is combined to generate singer facial animation with a sense of rhythm and accurate content expression. The interaction relationship between characters is modeled through the CVAE encoder.
It achieves natural interactive effects of singers' facial animations in multi-person chorus scenes, broadens the research on singing-driven multi-person interaction, provides the first chorus dataset, and enhances the facial animation interaction effects between singers.
Smart Images

Figure CN120612410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio-driven three-dimensional facial animation generation, and in particular to a multi-person interactive three-dimensional facial animation generation method. Background Art
[0002] Audio-driven 3D facial animation generation has become a key research direction in multimodal visual synthesis and holds great promise for modern multimedia applications such as avatars, gaming, and filmmaking. Currently, research in this field falls into two main categories: 2D-based and 3D-based approaches. 2D approaches typically rely on techniques such as optical flow, keypoint detection, or disentangled representations to generate facial animation synchronized with audio. However, these approaches have significant limitations when scaling to 3D applications.
[0003] In contrast, 3D methods are mainly divided into two categories based on the difference in feature representation. One category directly predicts vertex displacement bias and outputs a mesh sequence; the other category predicts the coefficients of a parameterized FLAME face model. This method can decouple the various facial regions and achieve separate supervision. In addition, audio-driven facial animation methods can be divided into speech-driven and singing-driven methods based on the driving source. Speech-driven methods have always received widespread attention and have made significant progress. However, singing-driven methods are still in their infancy and face many challenges in the future due to the complexity of their emotional expression.
[0004] Although existing research has achieved certain results, most methods still focus on voice-driven sources and often ignore the complexity of song expression, which is crucial for creating high-dimensional emotional expressions and interactive animations. In addition, voice-driven expressions are relatively simple, while singing, especially in multi-person choral performances, usually requires more exaggerated emotions and wider head movements to convey expression intentions. Therefore, the present invention collects and organizes a multimodal dedicated dataset focusing on chorus, aiming to fill the gaps in existing research. At present, most existing methods focus on single-person facial animation generation, while research on multi-person animation interaction is still in the pioneering stage. To this end, the present invention first explores the generation of song-driven multi-person animation interaction to promote the development of this field. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and propose a method for generating three-dimensional facial animation for multi-person interaction, which can fill the gap in this field in singing-driven multi-person interaction scenarios, obtain singer facial animation with a sense of rhythm and accurate content expression, and the interaction effect between singers is natural and reasonable.
[0006] To achieve the above-mentioned purpose, the present invention provides a technical solution: a method for generating a multi-person interactive 3D facial animation, comprising the following steps:
[0007] S1: Acquire and preprocess the chorus video data to obtain time-synchronized audio data of multiple singers and a set of continuous video frames for each singer;
[0008] S2: Separate the audio tracks of the audio data to obtain vocals and background music, and perform feature extraction on each to extract content features representing the vocals and rhythm features representing the background music. A 3D face reconstruction algorithm is used to extract 3D facial coefficients from the video frame set. These coefficients are represented by a parameterized face model, FLAME, and used to construct the singer's facial action sequence.
[0009] S3: Use the trained dual-branch optimization model to learn the mapping relationship between the rhythm features and content features of audio data and the facial action sequence, and obtain the singer's facial animation with a sense of rhythm and accurate content expression; wherein, the dual-branch optimization model includes branch 1 and branch 2, branch 1 is used to model the mapping relationship between each singer's audio and facial action, specifically, using the content features of the human voice, aligning it with the frame length of the action sequence through a linear interpolation layer, and then inputting it into the Transformer decoder for decoding to obtain each singer's facial action sequence; branch 2 is an audio conditional guidance module based on peer perception, which is used to model the interaction relationship between characters. The audio conditional guidance module uses a CVAE encoder, takes the audio content features and background music rhythm features of the peer as joint conditional features, and maps the joint conditional features to a latent space, samples a latent variable from the latent space, and decodes it through a CVAE decoder to obtain a head posture sequence; finally, the head posture sequence obtained by branch 2 is used to optimize the facial action sequence obtained by branch 1, and finally generates a singer's facial animation with a sense of rhythm and accurate content expression.
[0010] Further, in step S2, the chorus video data is represented as Where M and G represent the audio data and video frame sets respectively, N′ represents the total number of chorus pairs, and i represents the i-th chorus video data; the audio data M is separated into tracks, and the separation adopts the Spleeter algorithm to separate the chorus vocal sequence A=(A1, A2, ..., A i ) and background music B=(B1,B2,...,B i ), where A i 、B i Respectively represent the chorus vocal sequence and background music sequence of the audio data in the i-th pair of chorus video data, audio data M = {A, B}; for the chorus vocal sequence A i Use MossFormer2 algorithm to separate human voice and get and in Represent the audio of singer S1 and singer S2 in the i-th chorus video data respectively; for audio content feature extraction, the audio encoder of the pre-trained Wav2Vec model is used, and for background music rhythm feature extraction, librosa is used to extract feature vectors; the video frame set G is annotated for 3D face reconstruction, and the process is as follows: first, the MTCNN face detection algorithm is used to detect the face images of singers S1 and singer S2 from the video frame set G, and the face image sets I1 and I2 are obtained. Then, according to the index numbers of the face image sets I1 and I2, the common intersection is filtered out to obtain the continuous frame image sets E1 and E2, and the EMOCA v2 face reconstruction algorithm is used to extract the 3D face coefficients of the continuous frame image sets E1 and E2 as the facial action sequences of singers S1 and S2. and
[0011] Further, in step S3, set F i j represents the facial action sequence of the jth singer in the i-th chorus data, where F i j =(f1,f2,...,f t ), j = {S1, S2}, f t Represents the one-dimensional facial action sequence of frame t; each frame action sequence is decomposed into a facial expression sequence α t and head pose sequence β t , therefore, each frame action sequence f t Expressed as: f t ={α t ,β t}, and the facial expression sequence α t and head pose sequence β t They are composed of 53 and 3 coefficients respectively, so each frame of action sequence has a total of 56 coefficients, α t ∈R 53 , β t ∈R 3 , f t ∈R 56 , where R 53 represents the 53-dimensional real number field, R 3 represents the 3D real number field, R 56 Represents a 56-dimensional real number field;
[0012] In the design of branch 1, the audio input of each singer is represented as a one-dimensional signal tensor Where R is the real number domain, b is the batch size, T is the frame length, and the audio encoder of the pre-trained Wav2Vec model is used for feature extraction and converted into a hidden state H with semantic and emotional information. a :
[0013]
[0014] Where k represents the number of time steps;
[0015] The hidden state H a After linear transformation, it is adjusted to the target dimension d so that it can match the input dimension of the subsequent module and obtain the encoding feature C in :
[0016] C in =W a H a +b a ,C in ∈R b×k×d
[0017] Where d is the feature dimension set by the parameter, W a Indicates H a The weight of b a Indicates H a The bias term;
[0018] C in As the context input in the form of self-attention, it is fed into the Transformer decoder to extract the temporal coefficient feature C out , capturing the long-range dependencies of audio features in time series;
[0019] C out =TrandformerDecoder(C in ,C in ),C out ∈R b×k×d
[0020] Where TransformerDecoder is the Transformer decoder;
[0021] The output C of the Transformer decoder out The predicted facial action sequence F is mapped to the facial action space O through the linear layer i j ', respectively, the predicted facial action sequence F i j ' and the corresponding real facial action sequence F i j Input into the parametric face model FLAME to get the vertex coordinates V i j ' and V i j ;
[0022] F i j '=Wv C out +b v
[0023] Where W v Indicates C out Weight, b v Indicates C out The bias term;
[0024] Use the pre-loaded lip mask lip V i j and Extract lip area vertices
[0025]
[0026] Calculate separately MSE is used as the loss function of the lip vertex coordinates L lip , calculate V i j 、V i j 'MSE as the loss function of global vertex coordinates L v :
[0027]
[0028] The final loss function Loss1 of branch 1 is:
[0029] Loss1=θL lip +(1-θ)L v
[0030] Where θ is a learnable parameter.
[0031] Furthermore, in step S3, for interactive generation of head postures, in the design of branch 2, if the head posture sequence of singer S1 is modeled Then the peer perception condition U is and B i ; On the contrary, if we model the head posture sequence of singer S2 Then the peer perception condition U is and B i ; The model expression of branch 2 is:
[0032]
[0033] Where Func(·) represents the function for solving the head pose sequence, and · represents the input peer perception condition;
[0034] Use the ResUnet structure encoder to encode the target singer's head posture sequence, extract high-level temporal features across time, and obtain the posture embedding vector P emb , and concatenate it with the companion perception condition U and the reference posture feature ref to obtain the joint condition feature X in :
[0035] X in =Concat(ref,P emb ,U)
[0036] During training, the joint conditional features X in Input into the CVAE encoder, which consists of a multi-layer perceptron MLP, which predicts the output as the mean μ i and log variance For sampling the latent variable z from a normal distribution:
[0037] μ=W μ MLP(X in )+b μ
[0038]
[0039] Where W μ represents the weight matrix used to calculate the mean, b μ represents μ i The bias term, represents the weight matrix used to calculate the log variance, express The bias term;
[0040] The latent variable z is sampled from the standard normal distribution and concatenated with the reference pose feature ref and the peer perception condition U to obtain the variable z' as the input of the CVAE decoder, and the intermediate decoding feature X is obtained through the MLP network. out , and then reconstruct the posture embedding P through ResUnet emb ', and finally pass through the linear layer Linear pose Output the final head pose sequence P i j ':
[0041] z'=Concat(ref,z,U)
[0042] X out =MLP(z')
[0043] P emb '=ResUnet(X out )
[0044] P ij '=Linear pose (P emb ')
[0045] In the supervision part, the predicted head pose sequence P is calculated i j ' and the true head pose sequence P i j The MSE loss between them is used as the reconstruction loss L rec , since the latent variable z is a priori distributed according to the standard normal distribution, the output of the CVAE encoder has a mean of μ i , the variance is Gaussian distribution, calculate the KL divergence L between the two distributions kl :
[0046]
[0047] z~N(0,1)
[0048]
[0049] Where N represents the normal distribution, dz is the differential element, which represents the integral of the variable z;
[0050] The total loss function of branch 2, Loss2, is composed of the reconstruction loss L rec and KL divergence L kl It consists of two parts:
[0051] Loss2=λL rec +(1-λ)·L kl
[0052] Where λ is a learnable parameter;
[0053] The final head pose sequence P obtained by branch 2 i j '∈R k×3 Result F for optimizing branch 1 i j '∈R k×56 , and finally obtain the facial animation result F of each singer with perceptual interaction refine :
[0054] F refine =(F i j '[:,:50],P i j ',F i j '[:,53:56]).
[0055] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0056] 1. This paper explores the generation of multi-singer facial animation driven by choral singing for the first time, broadens the research field of audio-driven 3D facial animation generation, and lays a foundation for future research.
[0057] 2. This paper preprocesses chorus video data and obtains the first chorus dataset with chorus vocal sequences, background music and 3D face reconstruction annotations, which makes up for the lack of datasets in multi-person interactive facial animation tasks and expands the research potential of downstream tasks in this field.
[0058] 3. The present invention designs a dual-branch optimization model to learn the mapping relationship between the rhythm features and content features of audio data and action sequences, thereby enhancing the interactive effect of multi-person facial animation.
[0059] In summary, this paper explores the generation of multi-singer facial animation driven by choral singing for the first time, contributing the first choral dataset, filling a gap in this field for multi-person collaborative interaction scenarios and inspiring future research directions. Leveraging a trained two-branch optimization model, the mapping relationship between the rhythmic and content features of audio data and action sequences is learned, resulting in rhythmic and accurately expressed facial animations of singers, and enhancing the interactive effects of facial animations between singers. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a framework diagram of the method of the present invention.
[0061] Figure 2 Schematic diagram of the process built for the dataset. DETAILED DESCRIPTION
[0062] The present invention will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0063] like Figure 1 and Figure 2 As shown, this embodiment discloses a method for generating 3D facial animation for multi-person interaction, the details of which are as follows:
[0064] 1) Chorus dataset construction:
[0065] The chorus video data is represented as Among them, M and G represent the audio data and video frame set respectively, N ′ Represents the total number of chorus pairs, i represents the i-th pair of chorus video data; the audio data M is separated into tracks, and the separation adopts the Spleeter algorithm to separate the chorus vocal sequence A=(A1, A2, ..., A i) and background music B=(B1,B2,...,B i ), where A i 、B i Respectively represent the chorus vocal sequence and background music sequence of the audio data in the i-th pair of chorus video data, audio data M = {A, B}; for the chorus vocal sequence A i Use MossFormer2 algorithm to separate human voice and get and in Represent the audio of singer S1 and singer S2 in the i-th chorus video data respectively; for audio content feature extraction, the audio encoder of the pre-trained Wav2Vec model is used, and for background music rhythm feature extraction, librosa is used to extract feature vectors; the video frame set G is annotated for 3D face reconstruction, and the process is as follows: first, the MTCNN face detection algorithm is used to detect the face images of singers S1 and singer S2 from the video frame set G, and the face image sets I1 and I2 are obtained. Then, according to the index numbers of the face image sets I1 and I2, the common intersection is filtered out to obtain the continuous frame image sets E1 and E2, and the EMOCA v2 face reconstruction algorithm is used to extract the 3D face coefficients of the continuous frame image sets E1 and E2 as the facial action sequences of singers S1 and S2. and
[0066] To demonstrate the uniqueness and contribution of the dataset, the following is a comparison of the currently relevant datasets, as shown in Table 1.
[0067] Table 1 Dataset details
[0068] Dataset name Drive type Duration / h BGM 2D 3D Multiplayer MEAD voice 40.0 - √ - - HDTF voice 15.8 - √ - - VOCASET voice 0.5 - - √ - 3D-ETF voice 6.5 - - √ - RAVDESS voice 2.6 - √ - - Song2Face singing 2.0 √ √ - - Musicface singing 40.0 √ √ - - SingingHead singing 27.0 √ √ √ - Chorus Dataset singing 8.0 √ - √ √
[0069] 2) Use the trained dual-branch optimization model to learn the mapping relationship between the rhythm features and content features of audio data and the facial action sequence, so as to obtain the singer's facial animation with a sense of rhythm and accurate content expression; wherein, the dual-branch optimization model includes branch 1 and branch 2, branch 1 is used to model the mapping relationship between the audio and facial action of each singer, specifically by using the content features of the human voice, aligning it with the frame length of the action sequence through a linear interpolation layer, and then inputting it into the Transformer decoder for decoding to obtain the facial action sequence of each singer; branch 2 is an audio conditional guidance module based on peer perception, which is used to model the interaction relationship between characters. The audio conditional guidance module uses a CVAE encoder, takes the audio content features and background music rhythm features of the peer as joint conditional features, and maps the joint conditional features into a latent space, samples a latent variable from the latent space, and decodes it through a CVAE decoder to obtain a head posture sequence; finally, the head posture sequence obtained by branch 2 is used to optimize the facial action sequence obtained by branch 1, and finally generates a singer's facial animation with a sense of rhythm and accurate content expression.
[0070] 2.1) Facial action sequence modeling of branch 1:
[0071] Set F i j represents the facial action sequence of the jth singer in the i-th chorus data, where F i j =(f1,f2,...,f t ), j = {S1, S2}, f t Represents the one-dimensional facial action sequence of frame t; each frame action sequence is decomposed into a facial expression sequence α t and head pose sequence β t , therefore, each frame action sequence f t Expressed as: f t ={α t ,β t}, and the facial expression sequence α t and head pose sequence β t They are composed of 53 and 3 coefficients respectively, so each frame of action sequence has a total of 56 coefficients, α t ∈R 53 , β t ∈R 3 , f t ∈R 56 , where R 53 represents the 53-dimensional real number field, R 3 represents the 3D real number field, R 56 represents a 56-dimensional real number field;
[0072] 2.1.1) In the design of branch 1, the audio input of each singer is represented as a one-dimensional signal tensor Where R is the real number domain, b is the batch size, T is the frame length, and the audio encoder of the pre-trained Wav2Vec model is used for feature extraction and converted into a hidden state H with semantic and emotional information. a :
[0073]
[0074] Where k represents the number of time steps;
[0075] 2.1.2) The hidden state H a After linear transformation, it is adjusted to the target dimension d so that it can match the input dimension of the subsequent module and obtain the encoding feature C in :
[0076] C in =W a H a +b a ,C in ∈R b×k×d
[0077] Where d is the feature dimension set by the parameter, W a Indicates H a The weight of b a Indicates H a The bias term;
[0078] 2.1.3) C in As the context input in the form of self-attention, it is fed into the Transformer decoder to extract the temporal coefficient feature C out , capturing the long-range dependencies of audio features in time series;
[0079] C out =TrandformerDecoder(C in ,C in ),C out ∈R b×k×d
[0080] Where TransformerDecoder is the Transformer decoder;
[0081] 2.1.4) Transformer decoder output C out The predicted facial action sequence F is mapped to the facial action space O through the linear layer i j ', respectively, the predicted facial action sequence F i j' and the corresponding real facial action sequence F i j Input into the parametric face model FLAME to get the vertex coordinates V i j ' and V i j ;
[0082] F i j '=W v C out +b v
[0083] Where W v Indicates C out Weight, b v Indicates C out The bias term;
[0084] Use the pre-loaded lip mask lip V i j and Extract lip area vertices
[0085]
[0086] Calculate separately MSE is used as the loss function of the lip vertex coordinates L lip , calculate V i j 、V i j 'MSE as the loss function of global vertex coordinates L v :
[0087]
[0088] The final loss function Loss1 of branch 1 is:
[0089] Loss1=θL lip +(1-θ)L v
[0090] Where θ is a learnable parameter.
[0091] 2.2) Head pose sequence modeling of branch 2:
[0092] For interactive generation of head poses, in the design of branch 2, if we model the head pose sequence of singer S1 Then the peer perception condition U is and B i ; On the contrary, if we model the head posture sequence of singer S2 Then the peer perception condition U is and B i ; The model expression of branch 2 is:
[0093]
[0094] Where Func(·) represents the function for solving the head pose sequence, and · represents the input peer perception condition;
[0095] 2.2.1) Use the encoder of ResUnet structure to encode the head posture sequence of the target singer, extract the high-level temporal features across time, and obtain the posture embedding vector P emb , and concatenate it with the companion perception condition U and the reference posture feature ref to obtain the joint condition feature X in :
[0096] X in =Concat(ref,P emb ,U)
[0097] 2.2.2) During training, the joint conditional feature X in Input into the CVAE encoder, which consists of a multi-layer perceptron MLP, which predicts the output as the mean μ i and log variance For sampling the latent variable z from a normal distribution:
[0098] μ=W μ MLP(X in )+b μ
[0099]
[0100] Where W μ represents the weight matrix used to calculate the mean, b μ represents μ i The bias term, represents the weight matrix used to calculate the log variance, express The bias term;
[0101] 2.2.3) Sample the latent variable z from the standard normal distribution and concatenate it with the reference pose feature ref and the peer perception condition U to obtain the variable z' as the input of the CVAE decoder, and pass it through the MLP network to obtain the intermediate decoding feature X out , and then reconstruct the posture embedding P through ResUnet emb ', and finally pass through the linear layer Linear pose Output the final head pose sequence P ij ':
[0102] z'=Concat(ref,z,U)
[0103] X out =MLP(z')
[0104] P emb '=ResUnet(X out )
[0105] P i j '=Linear pose (P emb ')
[0106] 2.2.4) In the supervision part, calculate the predicted head pose sequence P i j ' and the true head pose sequence P i j The MSE loss between them is used as the reconstruction loss L rec , since the latent variable z is a priori distributed according to the standard normal distribution, the output of the CVAE encoder has a mean of μ i , the variance is Gaussian distribution, calculate the KL divergence L between the two distributions kl :
[0107]
[0108] z~N(0,1)
[0109]
[0110] Where N represents the normal distribution, dz is the differential element, which represents the integral of the variable z;
[0111] The total loss function of branch 2, Loss2, is composed of the reconstruction loss L rec and KL divergence L kl It consists of two parts:
[0112] Loss2=λL rec +(1-λ)·L kl
[0113] Where λ is a learnable parameter;
[0114] 2.2.5) The final head pose sequence P obtained from branch 2 i j '∈R k×3 Result F for optimizing branch 1 i j '∈Rk×56 , and finally obtain the facial animation result F of each singer with perceptual interaction refine :
[0115] F refine =(F i j '[:,:50],P i j ',F i j '[:,53:56]).
[0116] To verify the effectiveness of the proposed method, LVE, FVE, and FDD were used as evaluation criteria. LVE represents the lip vertex error, which is used to measure the synchronization between sound and image and the accuracy of mapping. FVE represents the vertex error of the facial expression region, which measures the difference between the vertex coordinates of the expression region and the vertex coordinates of the true sequence. FDD represents the facial dynamic variance. A comparative analysis was conducted on the chorus dataset constructed by the proposed method with FaceFormer, CodeTalker, Imitator, SelfTalk, and Diffspeaker. The experimental results are shown in Table 2.
[0117] Table 2 Experimental results analysis table
[0118]
[0119]
[0120] Currently, there are no methods specifically designed for song-driven animation performed by multiple people. To ensure experimental fairness, we extend our single-person animation methods to a multi-person dataset. We fine-tune all methods on a chorus dataset to enable a fair comparison of the results. We then compare our proposed method with several 3D animation generation methods, evaluating their performance across multiple metrics. Notably, most existing methods directly predict vertex offsets, which limits their accuracy in modeling head pose sequences. To ensure a fair comparison, we evaluate all lip-related metrics. Quantitative results, shown in Table 2, show that our proposed method outperforms single-person audio-driven methods. Across multiple metrics, pre-trained models on speech datasets perform the worst, indicating a domain gap between speech and singing. Methods tailored for speech cannot be directly applied to the singing context. Furthermore, experimental results demonstrate that our proposed method achieves state-of-the-art performance across multiple metrics and is worthy of promotion.
[0121] The above embodiments are preferred implementations of the present invention, but the implementations of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A method for generating three-dimensional facial animation for multi-person interaction, characterized in that: The following steps are involved: S1: Acquire and preprocess the chorus video data to obtain time-synchronized audio data of multiple singers and a set of continuous video frames for each singer; S2: Separate the audio tracks of the audio data to obtain vocals and background music, and perform feature extraction on each to extract content features representing vocals and rhythm features representing background music; 3D face reconstruction algorithms are used to extract 3D face coefficients from a set of video frames. These coefficients are represented by a parameterized face model, FLAME, and used to construct the singer's facial action sequence. S3: Use the trained dual-branch optimization model to learn the mapping relationship between the rhythm features and content features of audio data and the facial action sequence, and obtain the singer's facial animation with a sense of rhythm and accurate content expression; wherein, the dual-branch optimization model includes branch 1 and branch 2, branch 1 is used to model the mapping relationship between each singer's audio and facial action, specifically, using the content features of the human voice, aligning it with the frame length of the action sequence through a linear interpolation layer, and then inputting it into the Transformer decoder for decoding to obtain each singer's facial action sequence; branch 2 is an audio conditional guidance module based on peer perception, which is used to model the interaction relationship between characters. The audio conditional guidance module uses a CVAE encoder, takes the audio content features and background music rhythm features of the peer as joint conditional features, and maps the joint conditional features to a latent space, samples a latent variable from the latent space, and decodes it through a CVAE decoder to obtain a head posture sequence; finally, the head posture sequence obtained by branch 2 is used to optimize the facial action sequence obtained by branch 1, and finally generates a singer's facial animation with a sense of rhythm and accurate content expression.
2. The method for generating a multi-person interactive 3D facial animation according to claim 1, wherein: In step S2, the chorus video data is represented as Where M and G represent the audio data and video frame sets respectively, N′ represents the total number of chorus pairs, and i represents the i-th chorus video data; the audio data M is separated into tracks, and the separation adopts the Spleeter algorithm to separate the chorus vocal sequence A=(A1, A2, ..., A i ) and background music B=(B1,B2,...,B i ), where A i 、B i Respectively represent the chorus vocal sequence and background music sequence of the audio data in the i-th pair of chorus video data, audio data M = {A, B}; for the chorus vocal sequence A i Use MossFormer2 algorithm to separate human voice and get and in Represent the audio of singer S1 and singer S2 in the i-th chorus video data respectively; for audio content feature extraction, the audio encoder of the pre-trained Wav2Vec model is used, and for background music rhythm feature extraction, librosa is used to extract feature vectors; the video frame set G is annotated for 3D face reconstruction, and the process is as follows: first, the MTCNN face detection algorithm is used to detect the face images of singers S1 and singer S2 from the video frame set G, and the face image sets I1 and I2 are obtained. Then, according to the index numbers of the face image sets I1 and I2, the common intersection is filtered out to obtain the continuous frame image sets E1 and E2, and the EMOCA v2 face reconstruction algorithm is used to extract the 3D face coefficients of the continuous frame image sets E1 and E2 as the facial action sequences of singers S1 and S2. and 3. The method for generating a multi-person interactive 3D facial animation according to claim 2, wherein: In step S3, set represents the facial action sequence of the jth singer in the i-th chorus data, where j={S1,S2},f t Represents the one-dimensional facial action sequence of frame t; each frame action sequence is decomposed into a facial expression sequence α t and head pose sequence β t , therefore, each frame action sequence f t Expressed as: f t ={α t ,β t }, and the facial expression sequence α t and head pose sequence β t They are composed of 53 and 3 coefficients respectively, so each frame of action sequence has a total of 56 coefficients, α t ∈R 53 , β t ∈R 3 , f t ∈R 56 , where R 53 represents the 53-dimensional real number field, R 3 represents the 3D real number field, R 56 Represents a 56-dimensional real number field; In the design of branch 1, the audio input of each singer is represented as a one-dimensional signal tensor Where R is the real number domain, b is the batch size, T is the frame length, and the audio encoder of the pre-trained Wav2Vec model is used for feature extraction and converted into a hidden state H with semantic and emotional information. a : Where k represents the number of time steps; The hidden state H a After linear transformation, it is adjusted to the target dimension d so that it can match the input dimension of the subsequent module and obtain the encoding feature C in : C in =W a H a +b a ,C in ∈R b×k×d Where d is the feature dimension set by the parameter, W a Indicates H a The weight of b a Indicates H a The bias term; C in As the context input in the form of self-attention, it is fed into the Transformer decoder to extract the temporal coefficient feature C out , capturing the long-range dependencies of audio features in time series; C out =TrandformerDecoder(C in ,C in ),C out ∈R b×k×d Where TransformerDecoder is the Transformer decoder; The output C of the Transformer decoder out The predicted facial action sequence is mapped to the facial action space O through the linear layer The predicted facial action sequences are and the corresponding real facial action sequences Input into the parametric face model FLAME to get the vertex coordinates and Where W v Indicates C out Weight, b v Indicates C out The bias term; Use the pre-loaded lip mask lip right and Extract lip area vertices Calculate separately MSE is used as the loss function of the lip vertex coordinates L lip ,calculate MSE is used as the loss function of the global vertex coordinates L v : The final loss function Loss1 of branch 1 is: Loss1=θL lip +(1-θ)L v Where θ is a learnable parameter.
4. The method for generating a multi-person interactive 3D facial animation according to claim 3, wherein: In step S3, for interactive generation of head poses, in the design of branch 2, if the head pose sequence of singer S1 is modeled Then the peer perception condition U is and B i ; On the contrary, if we model the head posture sequence of singer S2 Then the peer perception condition U is and B i ; The model expression of branch 2 is: Where Func(·) represents the function for solving the head pose sequence, and · represents the input peer perception condition; Use the ResUnet structure encoder to encode the target singer's head posture sequence, extract high-level temporal features across time, and obtain the posture embedding vector P emb , and concatenate it with the companion perception condition U and the reference posture feature ref to obtain the joint condition feature X in : X in =Concat(ref,P emb ,U) During training, the joint conditional features X in Input into the CVAE encoder, which consists of a multi-layer perceptron MLP, which predicts the output as the mean μ i and log variance For sampling the latent variable z from a normal distribution: μ=W μ ·MLP(X in )+b μ Where W μ represents the weight matrix used to calculate the mean, b μ represents μ i The bias term, represents the weight matrix used to calculate the log variance, express The bias term; The latent variable z is sampled from the standard normal distribution and concatenated with the reference pose feature ref and the peer perception condition U to obtain the variable z' as the input of the CVAE decoder, and the intermediate decoding feature X is obtained through the MLP network. out , and then reconstruct the posture embedding P through ResUnet emb ', and finally pass through the linear layer Linear pose Output the final head pose sequence z'=Concat(ref,z,U) X out =MLP(z') P emb '=ResUnet(X out ) In the supervision part, the predicted head pose sequence is calculated and the true head pose sequence The MSE loss between them is used as the reconstruction loss L rec , since the latent variable z is a priori distributed according to the standard normal distribution, the output of the CVAE encoder has a mean of μ i , the variance is Gaussian distribution, calculate the KL divergence L between the two distributions kl : z~N(0,1) Where N represents the normal distribution, dz is the differential element, which represents the integral of the variable z; The total loss function of branch 2, Loss2, is composed of the reconstruction loss L rec and KL divergence L kl Two parts composition: Loss2=λL rec +(1-λ)·L kl Where λ is a learnable parameter; The final head pose sequence obtained by branch 2 Results for optimizing branch 1 The final facial animation result F of each singer with perceptual interaction is obtained refine :