High-simulation audio-driven character expression generation method based on multi-modal large model feedback mechanism
Through the multimodal large model feedback mechanism, combined with multi-scale speech features and three-dimensional face geometric model, the problem of insufficient mimicry in virtual character expression drivers is solved, and high-precision and high-realistic expression generation is achieved.
Patent Information
- Application Number
- CN202510388882.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, in the virtual character expression drive, it is difficult to achieve insufficient delicacy and realism, especially in terms of subtle expressions and expression transitions, there is a problem of insufficient mimicry.
The multimodal large model feedback mechanism is adopted, and the multi-scale speech feature extraction and three-dimensional face geometric model driving is used, combined with expression quality evaluation and weight update, to achieve high-realistic character expression generation.
It improves the accuracy and naturalness of expression generation, especially in subtle expressions and expression transitions, and enhances the mimicry and emotional communication effect of expressions.
Smart Images

Figure QLYQS_2 
Figure QLYQS_8 
Figure QLYQS_10
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for generating human expressions driven by high-fidelity audio based on a multimodal large model feedback mechanism. Background Art
[0002] In the field of computer vision and computer graphics, character expression driving has always been a hot topic of research, especially in applications such as virtual characters, animation production, and augmented reality. Although technologies such as deep learning and generative adversarial networks (GANs) have made significant progress in recent years, and can capture the expressions of real faces and map them to virtual characters, the current facial expression driving still has the problem of insufficient realism. Specifically, although the system can capture and reproduce some basic expression changes well, it is still difficult to achieve a completely consistent effect with real human expressions in terms of subtle expressions, subtle muscle movements, and natural transitions of expressions. In addition, factors such as lighting, facial texture, and individual differences also increase the complexity of the expression driving model, causing the generated expressions to sometimes appear stiff or unnatural. Therefore, how to further improve the delicacy and realism of expressions remains a challenge that needs to be solved in this field. Summary of the invention
[0003] The purpose of the present invention is to solve the shortcomings of the prior art mentioned in the background technology, and to propose a rare earth oxide coal-fired catalyst and a preparation method thereof.
[0004] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: first, the input audio and facial image are obtained, multi-scale speech features of the input audio are extracted, and a preliminary three-dimensional facial geometric model is constructed based on the input image. Then, the three-dimensional facial geometric model is driven by the audio data to generate an initial expression; then, the realism of the expression is analyzed in real time through the feedback mechanism of the large model, and the expression generation is dynamically optimized to generate an expression driving vector for optimizing the facial expression, thereby realizing highly realistic facial expression driving.
[0005] A method for generating highly realistic audio-driven human expressions based on a multimodal large model feedback mechanism comprises the following steps: S1. Obtain input audio, and generate a multi-dimensional feature representation including an audio detail feature vector and a global semantic vector through multi-scale speech feature extraction; S2, obtaining a face image, inputting the face image into a human head reconstruction model, and realizing generation of a three-dimensional head model of the face based on the input image; S3, inputting the multi-dimensional feature representation of the audio detail feature vector and the global semantic vector into the pre-trained expression driving model, constructing the motion vector of the three-dimensional face geometric model, and driving the three-dimensional head model of the face to generate an initial expression sequence; S4 renders the initial expression sequence as a video, inputs it into the multi-modal large model for expression quality evaluation, and obtains the expression evaluation score; S5 updates the weights based on the expression evaluation score. After the weight update at specific intervals, the expression evaluation based on the feedback mechanism of the multi-modal large model is performed again until the standard expression score is reached, and the weight update is stopped; S6 outputs the expression driving vector in real time based on the expression driving model after weight update; during the generation process, the realism of the expression is analyzed in real time, and the parameters in the feature extraction and expression generation processes are dynamically adjusted, so as to continuously correct the output of the model, and make the generated expression driving vector reflect the semantics and emotional expressions of the input audio and images.
[0006] Preferably, the implementation method of multi-scale speech feature extraction for the input audio in S1 is as follows: First, the EMD is used to decompose the original speech signal, then the Pearson correlation coefficient is calculated for the decomposed IMF components, and the Hilbert-Huang transform energy entropy is also introduced as a non-linear correlation metric. According to the correlation coefficient and non-linear metric results, a weighted fusion strategy is used to select the most representative 5 IMF components as effective IMF components, and the selected IMF components are input into the multi-scale feature fusion module, while driving the model to learn and extract deep acoustic feature information. Such a processing method can effectively capture the time-varying information in the speech signal and the connection between adjacent frames, help the driving model to learn and extract deep acoustic feature information more deeply, and improve the accuracy of speech-driven facial expressions; The calculation method of the audio detail feature vector in S1 is as follows: First, the first-order difference, second-order difference, and FBank features are extracted from the effective IMF components, and then, based on these features, an MFbank feature map is constructed to capture the voiceprint features in the audio.
[0007] Preferably, the calculation method of the audio global semantic vector in S1 is as follows: perform improved PLP (Perceptual Linear Prediction) feature extraction on the original speech signal, specifically including: in the preprocessing stage, adopt the adaptive dynamic range control technology to enhance the high-frequency components of the speech signal and ensure the effective retention of high-frequency information; in the process of PLP feature extraction, introduce a multi-resolution filter bank, combine the Mel filter bank and the high-frequency compensation filter bank to capture the low-frequency and high-frequency features of the speech signal simultaneously; perform spectral enhancement on the extracted PLP features, and use the short-time Fourier transform and inverse filtering technology to further restore and enhance the high-frequency information; then, input the enhanced PLP features into the global semantic encoder, which adopts a multi-scale attention mechanism to perform weighted fusion on the low-frequency and high-frequency features to generate the audio global semantic vector; finally, optimize the global semantic vector through a non-linear mapping network to ensure that it can represent both the global semantic information of the speech signal and retain high-frequency details; The calculation method of the final audio multi-dimensional features in S1 is: splice the PLP feature map and the MFbank feature map to finally obtain multi-scale speech features.
[0008] Preferably, the 3D face geometric model in S2 is an improved face model that fuses the BFM model and the FLAME model. The specific implementation method is as follows: First, adopt a parameter-sharing framework to fuse the global shape optimization ability of BFM and the detailed expression ability of FLAME; Second, optimize the vertices of the key areas of the face through a vertex constraint strategy to enhance the accuracy of expression generation; Finally, introduce a regularization loss function and a dynamic resolution technology to ensure that the model can adapt to the computational requirements of different scenarios while maintaining a high degree of fidelity.
[0009] Preferably, the input audio features of the preliminary expression-driven model in S3 are the previously extracted multi-scale speech features, and the middle layer of the model consists of a convolutional neural network layer based on the Encoder-Decoder architecture; The Encoder of the Encoder-Decoder architecture in S3 is used to extract audio features and can be expressed as: ; where represents feature splicing, is a 5-layer convolutional network (convolution kernel size 3×3, stride 1, ReLU activation), is the aforementioned MFbank feature map, is the aforementioned PLP feature map. The decoder is used to generate the driving vector and can be expressed as: ; Among them is a 5-layer convolutional network (convolution kernel size 3×3, stride 1, ReLU activation).
[0010] Preferably, in S5, the multi-modal large model base for optimizing expression generation at specific interval steps is Qwen2-VL. In the specific usage process, in order to better evaluate expressions, here we construct a face expression image-description data pair to perform SFT fine-tuning on it.
[0011] Preferably, the feedback mechanism of the multi-modal large model in S5 is optimized through the following improved objective function: ; Among them, the generation accuracy term is used to measure the difference between the finally generated expression driving vector and the vector true value, and is calculated using the mean squared error function; the realism term is used to measure the realism of the expression vector, is the feedback score calculated by the multi-modal large model feedback; the cross-modal alignment regularization term introduces KL divergence to measure the audio feature distribution and the visual expression distribution for consistency, forcing the model to learn semantic alignment between modalities. Its calculation method is to perform Gaussian kernel density estimation (KDE) modeling on MFbank and PLP features and compare with the distribution of the expression driving vector; At the same time, here the hyperparameters α, β, γ are dynamically optimized through the meta-learning framework: ; Among them, σ is the Sigmoid function, and are learnable weight matrices to achieve real-time balance between "generation accuracy" and "realism"; and are respectively the in the generation accuracy term at time t and the feedback score calculated by the multi-modal large model feedback , here we control the network weights based on multiple indicators through such hyperparameters for balanced optimization to avoid falling into local optima.
[0012] Preferably, in S6, the real-time analysis of the realism of expressions is completed through prompt engineering.
[0013] Preferably, in the calculation of the realism of expressions in S6, three indicators are specifically constructed for lip expression verification. At the same time, here we set its role as a lip expression research expert in the systemprompt of the multi-modal large model to reduce the influence of human face differences: a. Lip Closure Degree Definition: This index measures the distance between the upper and lower lips, reflecting the degree of lip closure; Principle: The degree of lip closure is crucial for speech pronunciation or expression; Calculation method: By detecting the vertical distance between the two endpoints of the midlines of the upper and lower lips, the degree of lip closure is measured. The specific calculation formula is as follows: ; where, and are the coordinate distances of the midlines of the upper and lower lips respectively, is the reference distance for the initial lip opening and closing state; Lip Corner Symmetry Definition: This index measures the symmetry of the left and right lip corners, reflecting whether the lips are evenly stretched or lifted; Principle: Natural expressions are usually symmetrical. If the heights or positions of the lip corners are inconsistent when smiling or speaking, it may affect the naturalness of the expression or the clarity of pronunciation; Calculation method: By detecting the relative height and position differences of the left and right lip corners, the symmetry is calculated. The specific calculation formula is as follows: ; where, and are the heights of the left and right lip corners, is the width of the mouth; Lip Curvature Definition: This index measures the degree of curvature of the lip shape, especially the shape change of the lips when smiling or speaking; Principle: Natural smiles or expressions are usually accompanied by lip curvature changes. Especially when the lips are lifted or lowered, if the curvature is too rigid or straight, it may lead to unnatural expressions or pronunciations. For example, when pronouncing certain phonemes (such as / p / , / b / ), the lips should be completely closed, while when pronouncing vowels, the lips will be partially open; Calculation method: By fitting the curve shapes of the upper and lower lips, the curvature of the lip midline is calculated. Usually, a parabola fitting is used to quantify the curvature of the lips. The specific calculation formula is as follows: ; where, is the actual position of the upper / lower lip of the lips, is the fitted parabola curve. We use such an index as the systermprompt to input into the large model to achieve expression evaluation; At the same time, here we adaptively adjust the index weights according to the audio spectrum energy distribution. The weights of each score are expressed as: ; where, is the MFbank energy entropy of the current frame (measuring the complexity of speech information); is a learnable weight vector optimized through offline pre-training; is the Sigmoid function, which constrains the weight range to (-1, 1). This can prevent the network from experiencing gradient explosion during the learning process. By using such an adaptive adjustment of the metric weights, when the input audio is a special audio, such as the plosive sound / p / , significantly increases → the (lip closure degree weight) is automatically increased. At this time, the system preferentially optimizes the lip closure action.
[0014] Preferably, in S6, an improved PD control strategy is used to dynamically adjust the BlendShape parameters in the feature extraction and expression generation processes, specifically including: First, on the basis of the traditional calculation of the error change rate, a dynamic proportional-derivative gain calculation link is introduced, and the error change rate is smoothed by the weighted average method to avoid oscillations caused by noise or mutations; Second, an adaptive gain adjustment mechanism is adopted to dynamically adjust the proportional gain (Kp) and derivative gain (Kd) according to the steady-state value and instantaneous change trend of the error, ensuring that the adjustment process responds quickly and converges stably; Finally, by restricting the gain range and introducing dead zone control, over-adjustment of the parameters is prevented, thereby ensuring the stability of the BlendShape parameters and the robustness of the expression generation.
[0015] Compared with the prior art, the beneficial effects of the present invention are: The technical solution of the present invention constructs multi-scale speech features, and based on the speech features, preliminary face expression driving is performed, and then real-time dynamic feedback of face expressions is performed based on the multi-modal large model to ensure the high-precision and high-fidelity generation of face expressions. At the same time, the present invention specifically designs indexes for lip closure degree, lip angle symmetry, and lip curvature to achieve high-precision mouth shape control. Specific Embodiments
[0016] The content of the present invention can be more easily understood by referring to the following detailed description of the preferred embodiments of the present invention and the included examples. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs. When there is a contradiction, the definition in this specification shall prevail.
[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0018] A method for generating highly realistic audio-driven human expressions based on a multi-modal large model feedback mechanism, comprising the following steps: S1. Obtain the input audio, and through multi-scale speech feature extraction, generate a multi-dimensional feature representation including an audio detail feature vector and a global semantic vector; S2. Obtain a face image, and input the face image into a head reconstruction model to generate a 3D head model of the face based on the input image; S3. Input the multi-dimensional feature representation of the audio detail feature vector and the global semantic vector into a pre-trained expression-driven model, construct a 3D face geometric model motion vector, and drive the 3D head model of the face to generate an initial expression sequence; S4. Render the initial expression sequence into a video, input it into a multi-modal large model for expression quality evaluation, and obtain an expression evaluation score; S5. Update the weights based on the expression evaluation score. After the weight update at specific intervals, perform expression evaluation again based on the feedback mechanism of the multi-modal large model until the standard expression score is reached, and then stop the weight update; S6. Based on the expression-driven model after weight update, output an expression-driven vector in real time; during the generation process, analyze the realism of the expression in real time, dynamically adjust the parameters in the feature extraction and expression generation processes, so as to continuously correct the output of the model, and make the generated expression-driven vector reflect the semantics and emotional expressions of the input audio and image.
[0019] The implementation method of the multi-scale speech feature extraction constructed for the input audio in S1 is as follows: First, use EMD to decompose the original speech signal, then calculate the Pearson correlation coefficient for the decomposed IMF components, and also introduce the Hilbert-Huang transform energy entropy as a non-linear correlation metric. According to the correlation coefficient and non-linear metric results, adopt a weighted fusion strategy to select the 5 most representative IMF components as effective IMF components, and input the selected IMF components into the multi-scale feature fusion module to simultaneously drive the model to learn and extract deep acoustic feature information. Such a processing method can effectively capture the time-varying information in the speech signal and the connection between adjacent frames, help the driving model to learn and extract deep acoustic feature information more deeply, and improve the accuracy of speech-driven facial expressions; The calculation method of the audio detail feature vector in S1 is as follows: First, perform first-order difference, second-order difference, and FBank feature extraction on the effective IMF components, and then construct an MFbank feature map based on these features to capture the voiceprint features in the audio.
[0020] The calculation method of the audio global semantic vector in S1 is as follows: perform improved PLP feature extraction on the original speech signal, specifically including: in the preprocessing stage, adopt the adaptive dynamic range control technology to enhance the high-frequency components of the speech signal and ensure that the high-frequency information is effectively retained; in the process of PLP feature extraction, introduce a multi-resolution filter bank, combine the Mel filter bank and the high-frequency compensation filter bank to capture the low-frequency and high-frequency features of the speech signal simultaneously; perform spectral enhancement on the extracted PLP features, and use the short-time Fourier transform and inverse filtering technology to further restore and enhance the high-frequency information; then, input the enhanced PLP features into the global semantic encoder, which adopts a multi-scale attention mechanism to perform weighted fusion on the low-frequency and high-frequency features to generate the audio global semantic vector; finally, optimize the global semantic vector through a non-linear mapping network to ensure that it can represent the global semantic information of the speech signal and retain the high-frequency details; The calculation method of the final audio multi-dimensional features in S1 is: splice the PLP feature map and the MFbank feature map to finally obtain multi-scale speech features.
[0021] The 3D face geometric model in S2 is an improved face model that fuses the BFM model and the FLAME model. The specific implementation method is: first, adopt a parameter sharing framework to fuse the global shape optimization ability of BFM and the detailed expression ability of FLAME; second, optimize the vertices in the key areas of the face through a vertex constraint strategy to enhance the accuracy of expression generation; finally, introduce a regularization loss function and a dynamic resolution technology to ensure that the model can adapt to the computational requirements of different scenarios while maintaining a high degree of fidelity.
[0022] The input audio features of the preliminary expression-driven model in S3 are the originally extracted multi-scale speech features, and the middle layer of the model consists of a convolutional neural network layer based on the Encoder-Decoder architecture; The Encoder of the Encoder-Decoder architecture in the model in S3 is used to extract audio features and can be expressed as: ; where represents feature splicing, is a 5-layer convolutional network (convolution kernel size 3×3, stride 1, ReLU activation), is the aforementioned MFbank feature map, is the aforementioned PLP feature map. The decoder is used to generate the driving vector and can be expressed as: ; where It is a 5-layer convolutional network (convolution kernel size 3×3, stride 1, ReLU activation).
[0023] In S5, the multi-modal large model base for optimizing expression generation at specific interval steps is Qwen2-VL. In the specific usage process, in order to better evaluate expressions, here we constructed a face expression image-description data pair to perform SFT fine-tuning on it.
[0024] The feedback mechanism of the multi-modal large model described in S5 is optimized through the following improved objective function: ; Among them, the generation accuracy term is used to measure the difference between the finally generated expression driving vector and the vector ground truth, and is calculated using the mean squared error function; the realism term is used to measure the realism of the expression vector, is the feedback score calculated by the multi-modal large model feedback; the cross-modal alignment regularization term introduces KL divergence to measure the audio feature distribution and the visual expression distribution for consistency, forcing the model to learn semantic alignment between modalities. Its calculation method is to perform Gaussian kernel density estimation (KDE) modeling on MFbank and PLP features and compare with the distribution of the expression driving vector; At the same time, the hyperparameters α, β, γ are dynamically optimized through the meta-learning framework here: ; Among them, σ is the Sigmoid function, and are learnable weight matrices to achieve real-time balance between "generation accuracy" and "realism"; and are respectively the in the generation accuracy term at time t and the feedback score calculated by the multi-modal large model feedback . Here, we control the network weights based on multiple indicators through such hyperparameters for balanced optimization to avoid falling into local optima.
[0025] In S6, the real-time analysis of the realism of expressions is completed through prompt engineering.
[0026] During the calculation of the realism of expressions in S6, three indicators are specifically constructed to check lip expressions. At the same time, here we set the role of the multi-modal large model in the systemprompt as an expert in lip expression research to reduce the influence of human face differences: a. Lip Closure Degree Definition: This metric measures the distance between the upper and lower lips, reflecting the degree of lip closure; Principle: The degree of lip closure is crucial for speech pronunciation or expression; Calculation method: By detecting the vertical distance between the two endpoints of the midlines of the upper and lower lips, the degree of lip closure is measured. The specific calculation formula is as follows: ; where, and are the coordinate distances of the midlines of the upper and lower lips respectively, is the reference distance for the initial lip opening and closing state; Lip Corner Symmetry Definition: This metric measures the symmetry of the left and right lip corners, reflecting whether the lips are evenly stretched or lifted; Principle: Natural expressions are usually symmetric. If the heights or positions of the lip corners are inconsistent when smiling or speaking, it may affect the naturalness of the expression or the clarity of pronunciation; Calculation method: By detecting the relative height and position differences of the left and right lip corners, the symmetry is calculated. The specific calculation formula is as follows: ; where, and are the heights of the left and right lip corners, is the width of the mouth; Lip Curvature Definition: This metric measures the degree of curvature of the lip shape, especially the shape change of the lips when smiling or speaking; Principle: Natural smiles or expressions are usually accompanied by changes in lip curvature. Especially when the lips are lifted or lowered, if the curvature is too rigid or straight, it may lead to unnatural expressions or pronunciations; For example, when pronouncing certain phonemes (such as / p / , / b / ), the lips should be completely closed, while when pronouncing vowels, the lips will be partially open; Calculation method: By fitting the curve shapes of the upper and lower lips, the curvature of the midline of the lips is calculated. Usually, a parabola fitting is used to quantify the curvature of the lips. The specific calculation formula is as follows: ; where, is the actual position of the upper / lower lip of the lips, is the fitted parabola curve. We use such a metric as the systermprompt to input into the large model to realize expression evaluation; At the same time, here we adaptively adjust the metric weights according to the audio spectrum energy distribution. The weights of each score are expressed as: ; where, is the MFbank energy entropy of the current frame (measuring the complexity of speech information); is a learnable weight vector, optimized through offline pre-training; is the Sigmoid function, which constrains the weight range to (-1, 1). This can prevent the network from experiencing gradient explosion during the learning process. By using such an adaptive adjustment of the metric weights, when the input audio is a special audio, such as the plosive sound / p / , significantly increases → the (lip closure degree weight) is automatically increased. At this time, the system preferentially optimizes the lip closure action.
[0027] In S6, an improved PD control strategy is used to dynamically adjust the BlendShape parameters in the feature extraction and expression generation processes. Specifically, first, on the basis of the traditional calculation of the error change rate, a dynamic proportional-derivative gain calculation link is introduced, and the error change rate is smoothed by the weighted average method to avoid oscillations caused by noise or mutations; second, an adaptive gain adjustment mechanism is adopted to dynamically adjust the proportional gain (Kp) and derivative gain (Kd) according to the steady-state value and instantaneous change trend of the error, ensuring that the adjustment process responds quickly and converges stably; finally, by restricting the gain range and introducing dead zone control, over-adjustment of the parameters is prevented, thereby ensuring the stability of the BlendShape parameters and the robustness of the expression generation.
[0028] Embodiment
[0029] This embodiment provides a high-precision head animation generation method based on a single image and audio, including the following steps: S1 Obtain the input audio, and through, generate a multi-dimensional feature representation including an audio detail feature vector and a global semantic vector; S2 Obtain a face image and construct a preliminary three-dimensional face geometric model based on the input image; S3 Use the audio data to drive the three-dimensional face geometric model to generate an initial expression; S4 Through the feedback mechanism of the large model, optimize the expression generation at specific interval steps to generate an expression driving vector for optimizing the human face expression; S5 Perform real-time analysis on the realism of the expression, dynamically adjust the parameters in the feature extraction and expression generation processes, thereby continuously correcting the output of the model, so that the generated expression driving vector more accurately reflects the semantics and emotional expressions of the input audio and image; The implementation method of multi-scale speech feature extraction for the input audio in S1 is as follows: First, the EMD is used to decompose the original speech signal. Then, the Pearson correlation coefficient is calculated for the decomposed IMF components, and the 5 IMF components with the largest correlation are selected as the effective IMF components. Such a processing method can effectively capture the time-varying information in the speech signal and the connection between adjacent frames, helping to drive the model to learn more deeply and extract deep acoustic feature information, and improving the accuracy of speech-driven facial expressions. The calculation method of the audio detail feature vector in S1 is as follows: First, the first-order difference, second-order difference, and FBank features are extracted from the effective IMF components. Then, the MFbank feature map is constructed based on these features.
[0030] The calculation method of the audio global semantic vector in S1 is: Extract the PLP (Perceptual Linear Prediction) features from the original speech signal.
[0031] The calculation method of the final audio multi-dimensional feature in S1 is: Concatenate the PLP feature map and the MFbank feature map to finally obtain the multi-scale speech features.
[0032] The 3D face geometric model in S2 is an improved face model that fuses the BFM model and the FLAME model.
[0033] The input audio feature of the preliminary expression-driven model in S3 is the previously extracted multi-scale speech features, and the middle layer of the model is composed of convolutional neural network layers.
[0034] The multi-modal large model base for optimizing expression generation at specific interval steps in S4 is Qwen2-VL.
[0035] The real-time analysis of the authenticity of the expression in S5 is completed through prompt engineering.
[0036] During the calculation of the authenticity of the expression in S5, it is best to construct three indicators for better lip expression inspection: Lip Closure Degree Definition: This indicator measures the distance between the upper and lower lips and reflects the degree of lip closure.
[0037] Principle: The degree of lip closure is crucial for speech pronunciation or expression. For example, when pronouncing certain phonemes (such as / p / , / b / ), the lips should be completely closed, while the lips will be partially open when pronouncing vowels.
[0038] Calculation method: Measure the degree of lip closure by detecting the vertical distance between the two endpoints of the midlines of the upper and lower lips. The specific calculation formula is as follows: ; Among them, and are the coordinate distances of the midlines of the upper and lower lips respectively, is the reference distance for the initial lip opening and closing state.
[0039] LipCornerSymmetry Definition: This index measures the symmetry of the left and right lip corners and reflects whether the lips are evenly stretched or lifted. Principle: Natural expressions are usually symmetric. If the heights or positions of the lip corners are inconsistent when smiling or speaking, it may affect the naturalness of the expression or the clarity of pronunciation.
[0040] Calculation method: Calculate the symmetry by detecting the relative height and position differences of the left and right lip corners. The specific calculation formula is as follows: ; Among them, and are the heights of the left and right lip corners, is the width of the mouth.
[0041] LipCurvature Definition: This index measures the degree of curvature of the lip shape, especially the shape change of the lips when smiling or speaking. Principle: Natural smiles or expressions are usually accompanied by changes in the curvature of the lips, especially when the lips are lifted or lowered. If the curvature is too rigid or straight, it may lead to unnatural expressions or pronunciations.
[0042] Calculation method: Calculate the degree of curvature of the midline of the lips by fitting the curve shapes of the upper and lower lips. Usually, a parabola fitting is used to quantify the curvature of the lips. The specific calculation formula is as follows: ; Among them, is the actual position of the upper / lower lip of the lips, is the fitted parabola curve; Dynamically adjust the parameters in the feature extraction and expression generation processes using the PD control strategy.
[0043] The examples involved in this article are only illustrative and are used to explain some features of the method described in the present invention. The appended claims are intended to claim the broadest possible scope that can be conceived, and the embodiments presented herein are only illustrative of the selected implementation manners according to the combinations of all possible embodiments. Therefore, the intention of the applicant is that the appended claims are not limited by the selection of the examples that illustrate the features of the present invention. Some numerical ranges used in the claims also include sub-ranges within them, and the variations within these ranges should also be interpreted as being covered by the appended claims whenever possible.
Claims
1. A high-fidelity audio-driven character expression generation method based on a multimodal large model feedback mechanism, characterized in that It includes the following steps: S1. Obtain the input audio, and through multi-scale speech feature extraction, generate a multi-dimensional feature representation including audio detail feature vectors and global semantic vectors; S2. Obtain a face image, and input the face image into a head reconstruction model to generate a three-dimensional head model of the face based on the input image; In S2, the three-dimensional geometric model of the face is an improved face model that fuses the BFM model and the FLAME model. The specific implementation method is as follows: First, adopt a parameter sharing framework to fuse the global shape optimization ability of BFM and the detailed expression ability of FLAME; Second, optimize the vertices in the key areas of the face through a vertex constraint strategy to enhance the accuracy of expression generation; Finally, introduce a regularization loss function and a dynamic resolution technology to ensure that the model can adapt to the computational requirements of different scenarios while maintaining a high degree of realism; S3. Input the multi-dimensional feature representation of the audio detail feature vectors and global semantic vectors into a pre-trained expression-driven model to construct a three-dimensional face geometric model motion vector, and drive the three-dimensional head model of the face to generate an initial expression sequence; In S3, the input audio features of the preliminary expression-driven model are the multi-scale speech features extracted originally, and the middle layer of the model is composed of a convolutional neural network layer based on the Encoder-Decoder architecture; The Encoder of the model's Encoder-Decoder architecture in S3 is used to extract audio features , which can be expressed as: ; Among them represents feature concatenation is a 5-layer convolutional network (convolution kernel size 3×3, stride 1, ReLU activation) is the aforementioned MFbank feature map is the aforementioned PLP feature map. The decoder is used to generate the driving vector , which can be expressed as: ; Among them is a 5-layer convolutional network (convolution kernel size 3×3, stride 1, ReLU activation); S4. Render the initial expression sequence into a video, input it into a multi-modal large model for expression quality evaluation to obtain an expression evaluation score, update the weight based on the expression evaluation score, and after the weight is updated at specific interval steps, perform expression evaluation again based on the feedback mechanism of the multi-modal large model until the standard expression score is reached and the weight update is stopped; In S4, the base of the multi-modal large model for optimizing the expression generation at specific interval steps is Qwen2-VL. In the specific use process, in order to better perform expression evaluation, here we construct a face expression image-description data pair to perform SFT fine-tuning on it; In S4, the feedback mechanism of the multi-modal large model is optimized through the following improved objective function: ; Among them, the generation accuracy term is used to measure the difference between the finally generated expression-driven vector and the vector ground truth, and is calculated using the mean square error function; the realism term is used to measure the realism of the expression vector, which is the feedback score calculated for the multimodal large model; the cross-modal alignment regularization term introduces the KL divergence to measure the audio feature distribution and the visual expression distribution for consistency, forcing the model to learn the semantic alignment between modalities, which is calculated by performing Gaussian kernel density estimation (KDE) modeling on the MFbank and PLP features and comparing with the distribution of the expression-driven vector; At the same time, the hyperparameters α, β, γ are dynamically optimized through a meta-learning framework: ; where σ is the Sigmoid function, and are learnable weight matrices to achieve real-time balance between "generation accuracy" and "fidelity"; and are respectively the in the generation accuracy term at time t and the feedback score calculated from the feedback of the multi-modal large model , and here we control the network weights based on multiple indicators through such hyperparameters for balanced optimization to avoid falling into local optima; S5. Based on the expression-driven model after weight update, output an expression-driven vector in real time; during the generation process, perform real-time analysis on the realism of the expression, dynamically adjust the parameters in the feature extraction and expression generation processes, so as to continuously correct the output of the model, and make the generated expression-driven vector reflect the semantics and emotional expressions of the input audio and image; In S5, the real-time analysis of the realism of the expression is completed through prompt engineering. In the calculation process of the realism of the expression, three indicators are specifically constructed to check the lip expression. At the same time, here we set the role of the multi-modal large model in the systemprompt as an expert in lip expression research to reduce its influence on the differences in human face; a. Lip Closure Degree Definition: This metric measures the distance between the upper and lower lips, reflecting the degree of lip closure; Principle: The degree of lip closure is crucial for speech pronunciation or expression; Calculation method: By detecting the vertical distance between the two endpoints of the midlines of the upper and lower lips, the degree of lip closure is measured. The specific calculation formula is as follows: ; Among them, and are the coordinate distances of the midlines of the upper and lower lips respectively, is the reference distance for the initial lip opening and closing state; Lip Corner Symmetry Definition: This metric measures the symmetry of the left and right lip corners, reflecting whether the lips are evenly stretched or lifted; Principle: Natural expressions are usually symmetric. If the heights or positions of the lip corners are inconsistent when smiling or speaking, it may affect the naturalness of the expression or the clarity of pronunciation; Calculation method: By detecting the relative height and position differences of the left and right lip corners, the symmetry is calculated. The specific calculation formula is as follows: ; Among them, and are the heights of the left and right lip corners, is the mouth width; Lip Curvature Definition: This metric measures the degree of curvature of the lip shape, especially the shape change of the lips when smiling or speaking; Principle: Natural smiles or expressions are usually accompanied by lip curvature changes, especially when the lips are lifted or lowered. If the curvature is too rigid or straight, it may lead to unnatural expressions or pronunciations; For example, when pronouncing certain phonemes (such as / p / , / b / ), the lips should be fully closed, while when pronouncing vowels, the lips will be partially open; Calculation method: By fitting the curve shapes of the upper and lower lips, the curvature of the lip midline is calculated. Usually, a parabola fitting is used to quantify the curvature of the lips. The specific calculation formula is as follows: ; Among them, is the actual position of the upper / lower lip, is the fitted parabolic curve. We use such an index as the systermprompt to input into the large model to realize expression evaluation. At the same time, here we adaptively adjust the index weights according to the audio spectrum energy distribution. The weights of each score are expressed as: ; Among them, is the MFbank energy entropy of the current frame (measuring the complexity of speech information); is a learnable weight vector, optimized through offline pre-training; is the Sigmoid function, which constrains the weight range to (-1, 1). This can prevent the network from experiencing gradient explosion during the learning process. By using such an adaptive adjustment of the metric weights, when the input audio is a special audio, such as the plosive sound / p / , significantly increases → the (lip closure weight) is automatically increased. At this time, the system preferentially optimizes the lip closure action; In S5, an improved PD control strategy is used to dynamically adjust the BlendShape parameters in the feature extraction and expression generation processes. Specifically, first, based on the traditional calculation of the error change rate, a dynamic proportional-derivative gain calculation link is introduced, and the error change rate is smoothed by the weighted average method to avoid oscillations caused by noise or mutations; Second, an adaptive gain adjustment mechanism is adopted to dynamically adjust the proportional gain (Kp) and derivative gain (Kd) according to the steady-state value and instantaneous change trend of the error, ensuring that the adjustment process responds quickly and converges stably; Finally, by restricting the gain range and introducing dead zone control, over-adjustment of the parameters is prevented, thus ensuring the stability of the BlendShape parameters and the robustness of expression generation.
2. A highly realistic audio-driven character expression generation method based on a multimodal large model feedback mechanism according to claim 1, characterized in that, The implementation method of multi-scale speech feature extraction for the input audio in S1 is as follows: First, EMD is used to decompose the original speech signal, then the Pearson correlation coefficient is calculated for the decomposed IMF components, and the Hilbert-Huang transform energy entropy is also introduced as a non-linear correlation metric. According to the correlation coefficient and non-linear metric results, a weighted fusion strategy is adopted to select the most representative 5 IMF components as effective IMF components, and the selected IMF components are input into the multi-scale feature fusion module to simultaneously drive the model to learn and extract deep acoustic feature information; The calculation method of the audio detail feature vector in S1 is as follows: First, the first-order difference, second-order difference, and FBank features are extracted from the effective IMF components, and then, based on these features, an MFbank feature map is constructed to capture the voiceprint features in the audio.
3. A high-fidelity audio-driven character expression generation method based on a multi-modal large model feedback mechanism according to claim 2, characterized in that The calculation method of the audio global semantic vector in S1 is as follows: perform improved PLP feature extraction on the original speech signal, specifically including: in the preprocessing stage, adopt the adaptive dynamic range control technology to enhance the high-frequency components of the speech signal and ensure that the high-frequency information is effectively retained; in the process of PLP feature extraction, introduce a multi-resolution filter bank, combine the Mel filter bank and the high-frequency compensation filter bank to simultaneously capture the low-frequency and high-frequency features of the speech signal; perform spectral enhancement on the extracted PLP features, and use the short-time Fourier transform and inverse filtering technology to further restore and enhance the high-frequency information; then, input the enhanced PLP features into the global semantic encoder, which adopts a multi-scale attention mechanism to perform weighted fusion on the low-frequency and high-frequency features to generate the audio global semantic vector; finally, optimize the global semantic vector through a non-linear mapping network to ensure that it can represent both the global semantic information of the speech signal and retain the high-frequency details; The calculation method of the final audio multi-dimensional features in S1 is as follows: splice the PLP feature map and the MFbank feature map to finally obtain multi-scale speech features.
Citation Information
Cited By
Multi-modal expression generation system and dynamic optimization method in virtual-real fusion scene
CN121999100A