Multi-modal facial animation data set generation method

By introducing large-scale language models, noise removal algorithms and animation curve adjustment methods into the multimodal facial animation dataset generation method, the problems of cumbersome steps, many manual interventions, and imbalance of phoneme distribution in the existing technology are solved, and efficient and automated dataset generation is achieved, improving data quality and applicability.

CN120198557APending Publication Date: 2025-06-24UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510334299.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing multimodal facial animation dataset generation methods have problems such as cumbersome steps, relying on a large number of manual interventions, failure to fully consider phoneme distribution balance, ambient sound interference and actors' personal style differences, resulting in inefficient data processing and inconsistency in the generation results.

Method used

By introducing a large-scale language model (LLM) to pre-calculate phoneme distribution and performing content generation, combining sample-oriented noise removal algorithms and data adjustment methods based on predefined animation curve pairs, the multimodal facial animation data set generation process is realized.

Benefits of technology

It greatly reduces manual intervention, improves the overall efficiency of data generation, and the generated multimodal facial animation data is of high quality, suitable for large-scale application scenarios such as virtual reality, augmented reality and film and television special effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198557A_ABST
    Figure CN120198557A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating a multi-modal facial animation data set. The method comprises the following steps of: obtaining performance contents: generating the performance contents which need to contain all Chinese characters to be generated; audio data processing: noise reduction is performed on the performance audio based on the examples; and facial animation data processing: realizing animation migration from the motion parameters of the performer to the target model through a non-sequential redirection algorithm. According to the method, the LLM, the noise removal algorithm and the data adjustment method are organically combined, so that automation and intelligence of the multi-modal facial animation data set generation process are realized, manual intervention is greatly reduced, and the overall efficiency of data generation is improved; the method is high in data quality and suitable for large-scale application, can generate high-quality multi-modal facial animation data, is suitable for large-scale application scenes such as virtual reality, augmented reality and movie and television special effects, and meets the requirements of the industry for high-quality data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of facial animation datasets, and in particular to a method for generating a multimodal facial animation dataset. Background Art

[0002] As an emerging artificial intelligence technology, facial animation generation technology, supported by multimodal datasets, is bringing revolutionary changes to fields such as game engines, e-commerce platforms, and social networks. By integrating audio, text, and facial information, this technology can capture and express the dynamic relationship between facial expressions and speech more deeply, thereby enhancing the naturalness and expressiveness of animation production.

[0003] In practical applications, the quality of multimodal facial animation datasets is a key factor directly affecting the quality of generated animations. However, existing multimodal data generation methods have significant defects in process design: First, the steps are cumbersome and rely on a large amount of manual intervention, lacking automated and standardized processes, resulting in low data processing efficiency; Second, when processing Chinese speech data, existing methods fail to fully consider the balanced distribution problem of vowel and consonant phonemes, which is likely to lead to inconsistencies between speech and facial expressions when generating high-quality animations; Finally, the recording and subsequent processing of the dataset need to be adjusted repeatedly according to the acting habits of individual actors, which is a complex and highly professional process and difficult to meet the needs of non-professionals. These limitations severely restrict the wide application and development potential of the technology, and innovative solutions are urgently needed to overcome them.

[0004] The existing technology has the following problems:

[0005] (1) The performance content is manually selected and the phoneme distribution balance problem is not considered:

[0006] In the current production process of multimodal facial animation datasets, the performance content of the actors in the dataset mostly comes from specific scripts or manual collection by staff. These contents are only simply screened and do not fully cover the distribution of vowels and consonants in phonemes, resulting in difficulty in achieving an ideal phoneme-emotion correspondence relationship between speech and facial expressions during the generation stage. This deficiency makes it difficult to maintain the consistency between speech and face when generating facial animations with different emotions, affecting the quality and credibility of the generation results.

[0007] (2) The influence of ambient sound on the dataset:

[0008] Current multimodal data generation methods generally adopt the method of directly collecting sound on-site, without fully considering the impact of ambient sound on the dataset. Especially when using professional equipment such as LightStage for multimodal data collection, ambient sound will significantly interfere with the subsequent speech-face consistency modeling process, resulting in the trained generation model being highly sensitive to ambient sound and unable to accurately capture the speech-face correspondence in the target scenario.

[0009] (3) Differences in the personal styles of actors and post-adjustment issues:

[0010] Different actors have their own personal style characteristics during performances, such as raising eyebrows, asymmetrical cheeks, etc. These differences lack unified norms and processing in existing datasets, resulting in a large amount of manual work being required for post-adjustment of the generated facial animations, reducing the standardization and applicability of the dataset. Summary of the Invention

[0011] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for generating a multimodal facial animation dataset.

[0012] The purpose of the present invention is achieved through the following technical solutions:

[0013] In the first aspect of the present invention, there is provided a method for generating a multimodal facial animation dataset, including the following steps:

[0014] Obtaining performance content: Pre-calculate the phoneme distribution according to the reference performance content library provided by the user, and extract Chinese characters to form the Chinese characters to be generated; according to the emotion category and language context selected by the user, and the group of Chinese characters to be generated, prompt a large-scale language model to generate performance content that needs to include all the Chinese characters to be generated;

[0015] Processing audio data: Denoise the performance audio based on an example;

[0016] Processing facial animation data: Realize the animation transfer of the performer's motion parameters to the target model through a non-sequential redirection algorithm, specifically including: calculating and obtaining the initial facial animation parameter sequence; optimizing the motion characteristics of the initial facial animation parameter sequence based on a predefined pair of animation curves to obtain an optimized facial animation parameter sequence; according to the predefined pair of animation curves input by the user, perform a correlation analysis, and perform a symmetry operation on the pair of animation curves whose correlation is lower than the lowest threshold input by the user.

[0017] Further, the pre-calculating the phoneme distribution according to the reference performance content library provided by the user and extracting Chinese characters to form the Chinese characters to be generated includes:

[0018] According to the reference performance content library provided by the user, calculate the phoneme distribution and determine the expected phoneme distribution;

[0019] Calculate the phoneme distribution of the existing performance content library to obtain the current phoneme distribution;

[0020] Compare the occurrence probability of each phoneme in the current phoneme distribution and the expected phoneme distribution. If the occurrence probability in the current phoneme distribution is less than that in the expected phoneme distribution, add it to the phoneme sequence to be generated;

[0021] Randomly select K phonemes from the phoneme sequence to be generated, and based on the dictionary library of the GB3212 standard, randomly select Chinese characters containing the corresponding phonemes for each phoneme; all these Chinese characters form the Chinese characters to be generated.

[0022] Further, the prompting the large-scale language model to generate performance content that needs to include all the Chinese characters to be generated according to the emotion category and language context selected by the user, and the group of Chinese characters to be generated, includes:

[0023] When generating lines, the user selects one from each of the seven emotion categories and ten language contexts, and at the same time randomly selects sentences from the reference performance content library in a non-replacement manner as few samples to prompt the large-scale language model;

[0024] After collecting all the information, send a prompt word to the large-scale language model, asking it to return performance content that matches the specified emotion, specified language context, uses the randomly selected sentence as a prompt example, and needs to include all the Chinese characters to be generated;

[0025] After receiving the generated performance content, check the generated performance content. If it does not contain the specified Chinese characters, regenerate it until the requirements are met.

[0026] Further, the noise reduction of the performance audio based on the example includes:

[0027] Select the first two seconds of silent segment of the recorded audio to analyze the environmental noise, perform short-time Fourier transform on the silent segment to obtain the spectral representation;

[0028] Calculate the average spectrum of the audio, that is, determine the noise spectrum; adjust the parameters of the NLMS filter according to the noise spectrum to optimize the noise suppression effect; after the filter parameter adjustment is completed, apply the adaptive filter for noise reduction processing;

[0029] After the user's performance ends, use the adjusted filter parameters of the NLMS filter to perform noise reduction on the performance audio.

[0030] Further, the calculation to obtain the initial facial animation parameter sequence is: based on the input performer binding model topology and performance action mesh sequence, obtain the animation control parameters through inverse blend shape calculation, specifically including:

[0031] First, according to the preset grouping of facial animation control curves, establish a mapping relationship matrix between each control curve and the corresponding blend shape. Subsequently, adopt a non-temporal optimization strategy and perform the following operations on each frame of the performance mesh sequence:

[0032] a) Construct a linear equation system based on vertex displacement, where the blend shape coefficients are the parameters to be solved;

[0033] b) Use the least squares optimization algorithm to solve the equation, and obtain the optimal weight coefficient combination of each animation control curve by minimizing the geometric deviation between the performance mesh vertices and the target model;

[0034] c) Iteratively adjust the solution result until convergence, and finally output the initial facial animation parameter sequence.

[0035] Furthermore, optimize the motion characteristics of the initial facial animation parameter sequence based on the predefined animation curve pairs to obtain the optimized facial animation parameter sequence, including:

[0036] Optimize the motion characteristics of the initial facial animation parameter sequence, including:

[0037] a) Establish a Pearson correlation coefficient matrix of the predefined animation curve pairs, and calculate the motion correlation index of each curve pair;

[0038] b) Set the correlation threshold τ, and perform symmetry correction on the animation curve pairs that meet the requirements, including:

[0039] i. Extract the parameter sequence of the corresponding control area;

[0040] ii. Calculate the mean vector of the parameter sequence;

[0041] iii. Update the parameter sequence;

[0042] c) Maintain the action continuity through parameter space interpolation, and generate the optimized facial animation parameter sequence after symmetry optimization.

[0043] Furthermore, the method further includes:

[0044] Multimodal synchronization: Synchronize the audio data with the facial animation according to the timestamps of the audio and the facial animation during recording, and cut off the redundant parts in the performance audio content.

[0045] Furthermore, the method further includes:

[0046] Quality assessment: Generate the lip vertex error and expression vertex error of the facial animation through the multimodal facial animation generation model, and evaluate the usability of the generated multimodal facial animation dataset.

[0047] The beneficial effects of the present invention are:

[0048] In an exemplary embodiment of the present invention, by organically combining an LLM, a noise removal algorithm, and a data adjustment method, the automation and intelligence of the multi-modal facial animation dataset generation process are realized, significantly reducing manual intervention and improving the overall efficiency of data generation.

[0049] High data quality, suitable for large-scale applications: It can generate high-quality multi-modal facial animation data, which is suitable for large-scale application scenarios such as virtual reality, augmented reality, and film and television special effects, meeting the industry's demand for high-quality data. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flowchart of a method for generating a multi-modal facial animation dataset provided in an exemplary embodiment of the present invention.

[0051] Figure 2 It is a flowchart for pre-computing phoneme distribution provided in an exemplary embodiment of the present invention.

[0052] Figure 3 It is a flowchart for calculating the phoneme sequence to be generated provided in an exemplary embodiment of the present invention.

[0053] Figure 4 It is a flowchart for extracting Chinese characters provided in an exemplary embodiment of the present invention.

[0054] Figure 5 It is a flowchart for generating prompts and generating performance content provided in an exemplary embodiment of the present invention.

[0055] Figure 6 It is a flowchart for generating audio of performance content provided in an exemplary embodiment of the present invention.

[0056] Figure 7 It is a flowchart for generating an animation parameter sequence of performance content provided in an exemplary embodiment of the present invention.

[0057] Figure 8 It is a flowchart for post-processing data provided in an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] In the description of the present invention, it should be noted that the directions or positional relationships indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0060] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "mounted", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0061] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0062] See Figure 1 , Figure 1 which shows a flowchart of a method for generating a multi-modal facial animation dataset provided in an exemplary embodiment of the present invention, including the following steps:

[0063] Obtaining performance content: Pre-calculate the phoneme distribution according to the reference performance content library provided by the user, and extract Chinese characters to form the Chinese characters to be generated; according to the emotion category and language context selected by the user, and the group of Chinese characters to be generated, prompt the large-scale language model to generate the performance content that needs to include all the Chinese characters to be generated;

[0064] Processing audio data: Denoise the performance audio based on the examples;

[0065] Processing facial animation data: Implement the animation transfer of the performer's motion parameters to the target model through the non-sequential redirection algorithm, specifically including: calculating and obtaining the initial facial animation parameter sequence; optimizing the motion features of the initial facial animation parameter sequence based on the predefined animation curve pair to obtain the optimized facial animation parameter sequence; according to the predefined animation curve pair input by the user, perform correlation analysis, and perform a symmetry operation on the animation curve pair with a correlation lower than the lowest threshold input by the user.

[0066] Specifically, in this exemplary embodiment, by introducing a performance content generation system based on a large language model (LLM), performance content with a balanced phoneme distribution is automatically generated; by introducing an example-guided noise removal algorithm, in a fixed environment such as Lightstage, the influence of environmental noise such as shutter operation on performance data is effectively suppressed; in addition, this exemplary embodiment also introduces a data adjustment method based on predefined animation curves, which can reduce manual intervention in the process of facial animation data processing, significantly reduce the impact of actor's personal style differences on the workload of post-adjustment, and improve the efficiency and quality of data processing.

[0067] Overall, the automation and efficiency of data processing are improved: by organically combining the LLM, noise removal algorithm, and data adjustment method, this exemplary embodiment realizes the automation and intelligence of the multi-modal facial animation dataset generation process, greatly reducing manual intervention and improving the overall efficiency of data generation. High data quality and suitable for large-scale applications: the method of this exemplary embodiment can generate high-quality multi-modal facial animation data, which is suitable for large-scale application scenarios such as virtual reality, augmented reality, and film and television special effects, meeting the industry's demand for high-quality data.

[0068] The following content will illustrate the preferred implementation methods for each step:

[0069] More preferably, in an exemplary embodiment, the pre-computation of phoneme distribution according to the reference performance content library provided by the user and the extraction of Chinese characters to form the Chinese characters to be generated include:

[0070] As Figure 2 shown (phoneme distribution pre-computation process), according to the reference performance content library provided by the user, calculate the phoneme distribution and determine the expected phoneme distribution;

[0071] Calculate the phoneme distribution of the existing performance content library to obtain the current phoneme distribution;

[0072] As Figure 3 shown (phoneme sequence to be generated calculation process), compare the occurrence probability of each phoneme in the current phoneme distribution and the expected phoneme distribution. If the occurrence probability in the current phoneme distribution is less than that in the expected phoneme distribution, then add it to the phoneme sequence to be generated;

[0073] As Figure 4 shown (Chinese character extraction process), randomly select K phonemes from the phoneme sequence to be generated, and based on the dictionary library of GB3212 standard, randomly select Chinese characters containing the corresponding phoneme for each phoneme; all these Chinese characters form the Chinese characters to be generated.

[0074] In a specific exemplary embodiment, assume the reference performance content is: {The weather is nice today. Let's go to the park together. Mom made a delicious dinner.} Phoneme splitting and statistics (decomposed according to Chinese pinyin): {jin(2)tian(3)qi(1)hen(1)hao(1), wo(2)men(2)yi(1)qu(1)gong(1)yuan(2), ma(2)zuo(1)le(1)mei(3)wei(2)de(1)wan(1)can(1)}. Standard phoneme probability distribution (partial): {tian(0.05)|jin(0.05)|mei(0.05), qu(0.05)|gong(0.05)|ma(0.05)} Current performance content: {The flowers in the park are blooming. My brother bought a book.} Current phoneme distribution: {gong(0.1)yuan(0.1)li(0.1)hua(0.1)kai(0.1), ge(0.2)mai(0.1)le(0.1)ben(0.1)shu(0.1)}. By comparison, it is found that the following phonemes need to be supplemented: {jin (standard 0.05 / current 0), mei (standard 0.05 / current 0), qu (standard 0.05 / current 0), tian (standard 0.05 / current 0)}.

[0075] Let K = 3, and randomly select: {jin, mei, qu}. Check the GB3212 dictionary: {jin→today, gold, catty, Tianjin, mei→beautiful, sister, plum, rose, qu→go, take, song, drive} Random combination result: {gold (jin) + plum (mei) + song (qu)}.

[0076] More preferably, in an exemplary embodiment, the large-scale language model is prompted to generate a performance content that needs to include all the to-be-generated Chinese characters according to the emotion category and language context selected by the user, and the to-be-generated Chinese character group, such as Figure 5 shown (prompt word generation and performance content generation flow), including:

[0077] When generating lines, the user selects one from each of the seven emotion categories and ten language contexts, and at the same time randomly selects sentences from the reference performance content library in a non-replacement manner as few samples to prompt the large-scale language model;

[0078] After collecting all the information, send a prompt word to the large-scale language model, asking it to return a performance content that matches the specified emotion, specified language context, and uses the randomly selected sentence as a prompt example and needs to include all the to-be-generated Chinese characters;

[0079] After receiving the generated performance content, check the generated performance content. If it does not include the specified Chinese characters, regenerate it until the requirements are met.

[0080] Preferably, in an exemplary embodiment, the noise reduction of the performance audio based on the example is as follows Figure 6 (as shown in the performance content audio generation process), and includes:

[0081] Select the first two seconds of silent segments of the recorded audio to analyze the environmental noise, perform short-time Fourier transform on the silent segments, and obtain the spectral representation;

[0082] Calculate the average spectrum of the audio, that is, determine the noise spectrum; adjust the NLMS filter parameters according to the noise spectrum to optimize the noise suppression effect; after the filter parameter adjustment is completed, apply an adaptive filter for noise reduction processing;

[0083] After the user's performance ends, use the adjusted filter parameters of the NLMS filter to reduce the noise of the performance audio.

[0084] Preferably, in an exemplary embodiment, the calculation of the initial facial animation parameter sequence is as follows Figure 7 (as shown in the performance content animation parameter sequence generation process), and is: based on the input performer binding model topology and performance action mesh sequence, obtain the animation control parameters through inverse blend shape calculation, specifically including:

[0085] First, according to the preset grouping of the facial animation control curves, establish the mapping relationship matrix between each control curve and the corresponding blend shape; then adopt a non-temporal optimization strategy to perform the following operations frame by frame on the performance mesh sequence:

[0086] a) Construct a linear equation system based on vertex displacement, where the blend shape coefficients are the parameters to be solved;

[0087] b) Solve the equation using the least squares optimization algorithm, and obtain the optimal weight coefficient combination of each animation control curve by minimizing the geometric deviation between the performance mesh vertices and the target model;

[0088] c) Iteratively adjust the solution result until convergence, and finally output the initial facial animation parameter sequence.

[0089] Preferably, in an exemplary embodiment, the motion feature optimization of the initial facial animation parameter sequence based on the predefined animation curve pair to obtain the optimized facial animation parameter sequence is as follows Figure 7 (as shown in the performance content animation parameter sequence generation process), and includes:

[0090] Perform motion feature optimization on the initial facial animation parameter sequence, including:

[0091] a) Establish a Pearson correlation coefficient matrix for predefined animation curve pairs. The Pearson correlation coefficient is used to measure the linear correlation degree of the corresponding time series data in any two animation curves, and calculate the motion correlation index of each curve pair;

[0092] Let the set of predefined animation curves be C = {C1, C2,..., Cn}, and each curve Ci corresponds to time series data {x i,t} (such as control point displacement parameters). The mean value of each curve is

[0093]

[0094] For any two curves Ci and Cj, their Pearson correlation coefficient is:

[0095]

[0096] b) Set a correlation threshold τ (τ ∈ [0, 1]). For animation curve pairs (C i , C j ) that satisfy corr(C i , C j ) < τ, perform symmetry correction:

[0097] i. Extract the corresponding time series data (X i , X j ) of the animation curve pair (C i,t , C j,t )

[0098] ii. Calculate the mean vector of the parameter sequence

[0099]

[0100] where μ is the equalization benchmark of the two curves, used to eliminate the asymmetry deviation.

[0101] iii. Update the corresponding time series data (X i , X j ) of the animation curve pair (C i,t , C j,t ):

[0102]

[0103] where the offset is defined as ΔX i = X i - μ i , ΔX j = X j - μ j , α is the smoothing factor, and ΔP is the original parameter offset;

[0104] c) Generate an optimized facial animation parameter sequence after symmetric optimization by maintaining action continuity through parameter space interpolation.

[0105] For the corrected parameter sequence {X' i}, insert transition frames in the time domain to eliminate the mutations caused by the correction.

[0106] The specific interpolation method is as follows:

[0107] For the parameters X' k and X' k+1 of adjacent frames t i (t k ) and t i (t k+1 ), generate an intermediate frame p' i (t):

[0108]

[0109] More preferably, in an exemplary embodiment, as Figure 8 shown, the method further includes:

[0110] Multimodal synchronization: Synchronize the audio data with the facial animation according to the timestamps of the audio and the facial animation during recording, and cut off the excess in the performance audio content.

[0111] More preferably, in an exemplary embodiment, the method further includes:

[0112] Quality assessment: Generate the lip vertex error and expression vertex error of the facial animation through the multimodal facial animation generation model, and evaluate the practicability of using the generated multimodal facial animation dataset.

[0113] Obviously, the above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, based on the above description, other different forms of changes or modifications can be made. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for generating a multimodal facial animation dataset, characterized in that: The following steps are involved: Performance content acquisition: pre-calculate the phoneme distribution based on the reference performance content library provided by the user, and extract Chinese characters to form the Chinese characters to be generated; According to the emotion category and language context selected by the user and the Chinese character group to be generated, prompting the large-scale language model to generate performance content including all the Chinese characters to be generated; Audio data processing: noise reduction of performance audio based on examples; Facial animation data processing: The animation migration of the performer's motion parameters to the target model is realized through a non-sequential retargeting algorithm, which includes: calculating and obtaining the initial facial animation parameter sequence; Based on the predefined animation curve pairs, the motion features of the initial facial animation parameter sequence are optimized to obtain the optimized facial animation parameter sequence; according to the predefined animation curve pairs input by the user, correlation analysis is performed, and symmetric operations are taken for the animation curve pairs whose correlation is lower than the minimum threshold input by the user.

2. The method for generating a multimodal facial animation dataset according to claim 1, characterized in that: The method of precalculating the phoneme distribution according to the reference performance content library provided by the user and extracting Chinese characters to form the Chinese characters to be generated includes: Calculate the phoneme distribution and determine the expected phoneme distribution based on the reference performance content library provided by the user; Calculate the phoneme distribution of the existing performance content library to obtain the current phoneme distribution; Compare the probability of each phoneme in the current phoneme distribution with the expected phoneme distribution. If the probability of the current phoneme distribution is less than the probability of the expected phoneme distribution, add the phoneme sequence to be generated. In the phoneme sequence to be generated, K phonemes are randomly selected, and based on the dictionary library of the GB3212 standard, Chinese characters containing the corresponding phonemes are randomly selected for each phoneme; all these Chinese characters constitute the Chinese character to be generated.

3. The method for generating a multimodal facial animation dataset according to claim 2, characterized in that: The prompting of the large-scale language model to generate the performance content including all the Chinese characters to be generated according to the emotion category and language context selected by the user and the Chinese character group to be generated includes: When generating lines, users select one of seven emotion categories and ten language contexts. At the same time, sentences are randomly selected from the reference performance content library without replacement as a few samples to prompt the large-scale language model. After all the information is collected, a prompt word is sent to the large-scale language model, requiring it to return a performance content with a specified emotion and a specified language context, using a randomly selected sentence as a prompt example, which needs to contain all the Chinese characters to be generated; After receiving the generated performance content, check the generated performance content, and if it does not contain the specified Chinese characters, regenerate it until the requirements are met.

4. The method for generating a multimodal facial animation dataset according to claim 1, wherein: The example-based noise reduction of the performance audio includes: The first two seconds of silent segments of the recorded audio are selected to analyze the environmental noise, and the silent segments are subjected to short-time Fourier transform to obtain the spectrum representation; Calculate the average frequency spectrum of the audio, that is, determine the noise spectrum; adjust the NLMS filter parameters according to the noise spectrum to optimize the noise suppression effect; after the filter parameters are adjusted, apply the adaptive filter to perform noise reduction processing; After the user's performance is finished, the performance audio is subjected to noise reduction using an NLMS filter using the adjusted filter parameters.

5. The method for generating a multimodal facial animation dataset according to claim 1, characterized in that: The calculation to obtain the initial facial animation parameter sequence is: based on the input performer binding model topology structure and performance action mesh sequence, the animation control parameters are obtained by inverse blend shape calculation, specifically including: First, according to the preset grouping of facial animation control curves, a mapping relationship matrix between each control curve and the corresponding blend shape is established; then, a non-sequential optimization strategy is used to perform the following operations on the performance mesh sequence frame by frame: a) constructing a linear equation system based on vertex displacement, where the blend shape coefficients are used as parameters to be determined; b) using the least squares optimization algorithm to solve the equations and obtain the optimal weight coefficient combination of each animation control curve by minimizing the geometric deviation between the performance mesh vertices and the target model; c) Iteratively adjust the solution results until convergence, and finally output the initial facial animation parameter sequence.

6. The method for generating a multimodal facial animation dataset according to claim 5, characterized in that: The method of optimizing the motion characteristics of the initial facial animation parameter sequence based on the predefined animation curve pair to obtain the optimized facial animation parameter sequence includes: Optimize the motion characteristics of the initial facial animation parameter sequence, including: a) Establishing the Pearson correlation coefficient matrix of predefined animation curve pairs and calculating the motion correlation index of each curve pair; b) Setting the correlation threshold τ, and performing symmetry correction on the animation curves that meet the requirements, including: i. Extract the parameter sequence corresponding to the control area; ii. Calculate the mean vector of the parameter sequence; iii. Update parameter sequence; c) Maintaining motion continuity through parameter space interpolation, generating an optimized facial animation parameter sequence after symmetry optimization.

7. The method for generating a multimodal facial animation dataset according to claim 1, characterized in that: The method further comprises: Multimodal Sync: Synchronize audio data with facial animations based on the timestamps of the audio during recording and the timestamps of the facial animations, and remove the redundant audio content in the performance.

8. The method for generating a multimodal facial animation dataset according to claim 1, characterized in that: The method further comprises: Quality evaluation: The lip vertex error and expression vertex error of facial animation are generated by the multimodal facial animation generation model to evaluate the practicality of using the generated multimodal facial animation dataset.