Portrait Action Driving Method, Device, Storage Medium, and Computer Equipment
By obtaining the audio expression characteristics of the audio data and generating action driving parameters, combined with overlapping frame prediction technology, the problem of poor adaptability of portrait action driving models in the existing technology is solved, and high coordination and natural portrait action animation generation is achieved, meeting the needs of high-end video production.
Patent Information
- Application Number
- CN202411784368.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The existing portrait action driving technology relies on a single-drive model, and multiple models are needed to control the action postures in different areas of the 3D portrait. The model is poorly adaptable and it is difficult to meet the needs of high-end video production.
A portrait action driving method is adopted to obtain the audio expression characteristics of the target audio data, determine the target expression parameter generation model based on overlapping frame prediction technology, generate action driving parameters, and drive the target portrait head action based on these parameters to generate portrait action animation with high coordination and naturalness.
It improves the coordination and nature of the driving effect, avoids the limitations of portraits and styles, and meets the needs of the high-end video production field.
Smart Images

Figure CN119540416B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a portrait action driving method, device, storage medium, and computer device. Background Art
[0002] With the wide application of AIGC (Artificial Intelligence Generated Content) technology in the video field, the voice-driven 3D portrait action generation technology is rapidly rising, showing great potential in multiple fields such as education, entertainment, and virtual reality. Especially in high-cost video shooting and production scenarios, such as film and television, games, and animations, this technology enables 3D characters or animals to speak, sing, and have conversations through voice driving, significantly reducing the cost and time of video production.
[0003] However, most current portrait action driving technologies rely on single driving models, and multiple models need to be constructed separately to precisely control the action postures of different regions of 3D portraits. In addition, these models not only require customized training for each portrait, but also usually use classification methods to define action styles during training, resulting in poor adaptability of the models. Generally speaking, the existing methods not only limit the application scope of the technology, but also lack coordination between the actions of different parts and the voice in the generated driving effects, making it difficult to meet the requirements of the high-end video production field. Summary of the Invention
[0004] The purpose of this application aims to at least solve one of the above technical defects, especially the technical defect that the existing methods not only limit the application scope of the technology, but also lack coordination between the actions of different parts and the voice in the generated driving effects, making it difficult to meet the requirements of the high-end video production field.
[0005] This application provides a portrait action driving method, and the method includes:
[0006] Obtain target audio data, and use a preset target feature extraction model to extract the audio expression features of the target audio data; the audio expression features include lip movement features, expression style features, and head posture features;
[0007] Determine a target expression parameter generation model; the target expression parameter generation model is trained using a target loss function based on the overlapping frame prediction technology.
[0008] Input the audio expression features into the target expression parameter generation model to obtain the action driving parameters output by the target expression parameter model.
[0009] Obtain a target portrait, and drive the head movement of the target portrait based on the action driving parameters to generate a portrait action animation.
[0010] Optionally, the target feature extraction model includes a temporal convolutional feature extractor and a multi-layer transformer encoder;
[0011] The extracting the audio expression features of the target audio data by using a preset target feature extraction model includes:
[0012] Input the target audio data into the temporal convolutional feature extractor to perform low-level feature extraction on the target audio data through the temporal convolutional feature extractor to generate audio phoneme features; the audio factor features include rhythm features, emotional features, and intonation features;
[0013] Use the multi-layer transformer encoder to encode the high-level features of the audio phoneme features and output them as audio expression features.
[0014] Optionally, the performing low-level feature extraction on the target audio data by the temporal convolutional feature extractor to generate audio phoneme features includes:
[0015] Extract the temporal audio features of the target audio data through the temporal convolutional feature extractor, and convert the temporal audio features into frequency-domain audio features to perform multi-dimensional feature extraction on the frequency-domain audio features to form audio phoneme features.
[0016] Optionally, the encoding the high-level features of the audio phoneme features by using the multi-layer transformer encoder and outputting them as audio expression features includes:
[0017] Use the multi-head self-attention mechanism in the multi-layer transformer encoder to read the temporal dependence relationship and cross-feature dependence relationship of the audio phoneme features, and based on the temporal dependence relationship and the cross-feature dependence relationship, use a layer-by-layer feature interaction weighting mechanism to perform high-level encoding on the audio phoneme features to generate audio expression features.
[0018] Optionally, the determining the target expression parameter generation model includes:
[0019] Construct a pre-trained model by using a Transformers architecture and a CNN architecture, and introduce an overlapping frame prediction technique into the pre-trained model to obtain an initial expression parameter generation model;
[0020] Input the pre-acquired sample audio data into the initial expression parameter generation model to obtain the predicted action driving parameters output by the initial expression parameter generation model;
[0021] Taking the predicted action driving parameter to approach the true action driving parameter of the sample audio data as the goal, and training the initial expression parameter generation model by using the target loss function;
[0022] When the initial expression parameter generation model meets the preset training conditions, the trained initial expression parameter generation model is used as the target expression parameter generation model.
[0023] Optionally, the calculation formula of the target loss function includes:
[0024]
[0025] In the formula, represents the target loss function; represents the true action driving parameter; represents the predicted action driving parameter; where represents the time window from the p-th moment to the w-th moment.
[0026] Optionally, the driving of the head movement of the target portrait based on the action driving parameter to generate a portrait action animation includes:
[0027] Determine a portrait action driving model; the portrait action driving model is composed of a lip movement driving module, an expression style driving module, a head pose driving module, and an animation generation module;
[0028] Taking the action driving parameter and the target portrait as input data, and inputting the input data into the lip movement driving module, the expression style driving module, and the head pose driving module respectively, to obtain a first driving result output by the lip movement driving module, a second driving result output by the expression style driving module, and a third driving result output by the head pose driving module;
[0029] Aligning and fusing the first driving result, the second driving result, and the third driving result in the animation generation module based on the target audio data to generate a portrait action animation.
[0030] This application also provides a portrait action driving device, including:
[0031] A feature extraction module, configured to obtain target audio data, and extract audio expression features of the target audio data by using a preset target feature extraction model; the audio expression features include lip movement features, expression style features, and head pose features;
[0032] A model determination module, configured to determine a target expression parameter generation model; the target expression parameter generation model is trained by using a target loss function based on an overlapping frame prediction technique;
[0033] A parameter prediction module, configured to input the audio expression features into the target expression parameter generation model to obtain action driving parameters output by the target expression parameter model;
[0034] An action driving module, configured to obtain a target portrait and drive the head movement of the target portrait based on the action driving parameters to generate a portrait action animation.
[0035] The present application further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the portrait action driving method according to any one of the above embodiments.
[0036] The present application further provides a computer device, including: one or more processors, and a memory;
[0037] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the portrait action driving method according to any one of the above embodiments are executed.
[0038] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0039] For the portrait action driving method, device, storage medium and computer device provided by the present application, when it is necessary to drive the movement of a target portrait, the target audio data for action driving can be obtained first, and then the audio expression features of the target audio data are extracted by using a pre-trained target feature extraction model. The target feature extraction model can capture features at multiple levels in the audio data at the same time, realize the many-to-many mapping between the lip movement features, expression style features and head pose features in the audio expression features, and further improve the coordination of the driving effect; then the target expression parameter generation model can be determined. Since the model is trained based on the overlapping frame prediction technique, the action driving parameters predicted by the target expression parameter generation model for the audio expression features can achieve seamless conversion in the time series, and further improve the naturalness of the driving effect. Therefore, driving the head movement of the target portrait based on the action driving parameters can generate a portrait action animation with high coordination and high naturalness. Through this method, the present application can directly copy the feature distributions in different forms in the target audio data, realize the many-to-many feature mapping, so as to avoid the limitations in aspects such as portraits and styles, and further improve the coordination and naturalness of the driving effect. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic flowchart of a portrait action driving method provided by an embodiment of the present application;
[0042] Figure 2 It is a schematic flowchart of a process for extracting audio expression features provided by an embodiment of the present application;
[0043] Figure 3 It is a schematic flowchart of a process for determining a target expression parameter generation model provided by an embodiment of the present application;
[0044] Figure 4 It is a schematic structural diagram of a portrait action driving device provided by an embodiment of the present application;
[0045] Figure 5 It is a schematic internal structure diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0047] Most current portrait action driving technologies rely on single driving models and need to separately construct multiple models to precisely control the action postures of different regions of 3D portraits; in addition, these models not only need to be customized and trained for each portrait, but also usually use classification methods to define action styles during the training process, resulting in poor adaptability of the models. Generally speaking, the existing methods not only limit the application scope of the technology, but also lack coordination between the actions of different parts and the speech in the generated driving effects, making it difficult to meet the requirements of the high-end video production field.
[0048] Based on this, the present application proposes the following technical solutions. For details, please refer to the following:
[0049] In one embodiment, as Figure 1 shown, Figure 1Schematic flowchart of a portrait motion driving method provided by an embodiment of the present application; the present application provides a portrait motion driving method, which specifically includes the following:
[0050] S110: Obtain target audio data, and use a preset target feature extraction model to extract the audio expression features of the target audio data.
[0051] In this step, when the user needs to perform motion driving on a portrait, the target audio data for motion driving can be selected on a computer device, so that after the computer device obtains the target audio data, the pre-trained target feature extraction model can be used to extract the audio expression features of the target audio data as the target motion driving basis for the corresponding portrait.
[0052] Among them, the audio expression features refer to the dynamic information related to the human expression hidden in the target audio data, including lip movement features, expression style features, and head pose features. Specifically, the lip movement features refer to the mouth shape and movement features related to the speech pronunciation process, such as the degree of lip opening and closing or the shape of lip opening and closing; the expression style features refer to the expression style features conveyed by the emotional or tone information hidden in the audio data, such as happy, angry, or confused; the head pose features refer to the head movement features expressed in the audio data, such as nodding, shaking the head, or tilting the head.
[0053] It can be understood that the target feature extraction model of the present application refers to a model that extracts features from the input target audio data and obtains multimodal fused audio expression features. It is trained with various types of audio data as training samples and the audio expression features corresponding to the audio data as training labels. The target feature extraction model trained in this way can simultaneously capture the complex correlations between features at multiple levels in the audio data, realize the many-to-many mapping between lip movement features, expression style features, and head pose features. This mapping can not only show the strong correlation between the audio data and the driving actions, but also effectively capture the potential cross-feature interaction relationships.
[0054] For example, in a target audio data, through the feature extraction results of the target feature extraction model, it can be found that some high pitches or rapid speech rates simultaneously drive more exaggerated lip movements, richer expression changes, and more frequent head movements. Therefore, this model can ensure the tight synchronization and emotional consistency of actions, expressions, and postures during the portrait motion driving process, thereby improving coordination.
[0055] S120: Determine the target expression parameter generation model; the target expression parameter generation model is trained using a target loss function based on the overlapping frame prediction technique.
[0056] In this step, after extracting the audio expression features of the target audio data through step S110, the computer device can also determine a target expression parameter generation model for converting the audio expression features into corresponding action driving parameters for portrait action driving.
[0057] It should be noted that in order for the prediction results of the target expression parameter generation model to achieve seamless conversion in the time series, when training the target expression parameter generation model, the computer device can use the target loss function for training based on the overlapping frame prediction technology.
[0058] Among them, the overlapping frame prediction technology refers to a time series prediction method aimed at achieving the temporal smoothness and consistency of prediction results by introducing partial overlapping regions between consecutive time windows; it mainly uses the data of the current window and the last part of the data of the previous window as conditional inputs, which can not only capture the short-term dependencies of the time series but also ensure seamless conversion between windows, avoiding breaks or discontinuities caused by independent window predictions.
[0059] Specifically, when applying the target expression parameter generation model, the overlapping frame prediction technology can divide the time series into windows with a fixed length, and there are a certain number of overlapping frames between the end of each window and the beginning of the next window. For example, for a window with a fixed length, in addition to inputting the audio data of the current window, the last few frames of the action driving parameters generated by the previous window are also provided as conditional inputs to the model, so that the model can utilize the context information of the time series, fully refer to historical features in the current prediction process, capture long-term and short-term dependencies, and thus achieve the temporal continuity and natural transition of the prediction results.
[0060] S130: Input the audio expression features into the target expression parameter generation model to obtain the action driving parameters output by the target expression parameter model.
[0061] In this step, after determining the target expression parameter generation model through step S120, the computer device can input the audio expression features into the target expression parameter generation model and perform parameter prediction on the audio expression features through the target expression parameter generation model. Since the target expression parameter generation model is trained using the target loss function based on the overlapping frame prediction technology, the computer device can directly obtain high-quality action driving parameters output by the target expression parameter model based on the audio expression features.
[0062] It can be understood that the action driving parameter is a fixed-length numerical vector extracted and mapped from audio expression features, which is used to drive the expression and action changes of the portrait. The action driving parameter can include multiple elements, and each element represents a control signal of the portrait in a specific action dimension, such as controlling eye movement, controlling lip movement, controlling head posture, etc. Through the action driving parameter, the present application can encode the multi-dimensional expression features in a numerical form into a compact representation, so that each dimension directly corresponds to a specific action of portrait driving.
[0063] Specifically, the target expression parameter generation model of the present application is trained by the method of adding noise and denoising. In the noise addition stage, the target expression parameter model can sample random noise data from the standard Gaussian distribution and gradually superimpose it on the real action driving parameter to construct the noise addition process. In each step of prediction in the denoising stage, the model can receive the noise data of the current step, audio expression features, including partial features of the current window and the previous window, and the motion parameters generated by denoising in the previous step as inputs, and estimate the noise residual of the current step through network inference, and then update and remove the noise in the reverse direction based on the estimation result, and generate an intermediate result closer to the real parameter; after multiple denoising iterations, the model can gradually optimize the highly noisy initial data to the final noiseless result, and output it as the action driving parameter, which is not only smooth and continuous in time, but also highly matched with the audio data semantically, and can maintain natural coordination between various action dimensions.
[0064] S140: Obtain a target portrait, and drive the head action of the target portrait based on the action driving parameter to generate a portrait action animation.
[0065] In this step, after obtaining the action driving parameter through step S130, the computer device can also obtain the target portrait, and then can drive the head action of the target portrait based on the action driving parameter to generate a portrait action animation with high coordination and high naturalness.
[0066] It can be understood that the action driving parameter here can capture and restore the rich dynamic features in the target audio data through accurate modeling and multi-dimensional feature mapping. It realizes the many-to-many feature mapping including lip movement, expression style, head posture, etc. by directly copying and mapping the feature distributions in different forms in the target audio data, so that the generated animation is closer to natural performance. Therefore, the present application can effectively avoid the limitations of portrait structure, style, or expression mode through this parameter generation mechanism independent of specific portrait characteristics, and improve the application range.
[0067] In the above embodiments, when it is necessary to perform action driving on a target portrait, target audio data for action driving can be obtained first, and then the audio expression features of the target audio data can be extracted by using a pre-trained target feature extraction model. The target feature extraction model here can capture features at multiple levels in the audio data at the same time, realizing a many-to-many mapping between the lip movement features, expression style features, and head pose features in the audio expression features, thereby improving the coordination of the driving effect. Then, a target expression parameter generation model can be determined. Since this model is trained based on the overlapping frame prediction technology, the action driving parameters obtained by predicting the audio expression features through the target expression parameter generation model can achieve seamless conversion in the time series, thereby improving the naturalness of the driving effect. Therefore, driving the head movement of the target portrait based on the action driving parameters can generate a portrait action animation with high coordination and high naturalness. Through this method, the present application can directly copy the feature distributions in different forms in the target audio data, realize many-to-many feature mapping, thereby avoiding limitations in aspects such as portraits and styles, and further improving the coordination and naturalness of the driving effect.
[0068] In one embodiment, the target feature extraction model in step S110 may include a time-domain convolutional feature extractor and a multi-layer transformer encoder; as Figure 2 shown, Figure 2 is a schematic flowchart of a process for extracting audio expression features provided by an embodiment of the present application; Figure 2 in, the process of extracting the audio expression features of the target audio data by using the preset target feature extraction model may include:
[0069] S111: Input the target audio data into the time-domain convolutional feature extractor to perform low-level feature extraction on the target audio data through the time-domain convolutional feature extractor, and generate audio phoneme features.
[0070] S112: Use the multi-layer transformer encoder to encode the high-level features of the audio phoneme features and output them as audio expression features.
[0071] In this embodiment, the target feature extraction model may use the Wav2Vec2 model as the basic architecture, which is mainly composed of a time-domain convolutional feature extractor and a multi-layer transformer encoder. Among them, the time-domain convolutional feature extractor can perform low-level feature extraction on the input target audio data and generate output audio phoneme features; the multi-layer transformer encoder can encode the high-level features of the audio phoneme features and output them as audio expression features.
[0072] Among them, the audio phoneme features refer to the features related to speech phonemes in audio data, which can reflect the basic constituent units of audio data and their changes over time. The audio phoneme features of this application can include rhythm features, emotional features, and intonation features; specifically, the rhythm features are related to the rhythm and duration changes in the audio data, the emotional features reflect the emotional expression in the audio data, and the intonation features describe the pitch change patterns in the audio data.
[0073] It can be understood that the time-domain convolutional feature extractor can efficiently process time-series data, retain the details of the audio data, while compressing and abstracting its key information, providing a high-quality basic feature representation for subsequent processing. Based on this, this application can capture the short-term dynamic change features in the target audio data through the time-domain convolutional feature extractor and generate representative audio phoneme features, including rhythm features, emotional features, and intonation features, which respectively correspond to the time rhythm pattern, emotional expression, and speech intonation change and other information in the target audio data.
[0074] In addition, based on the audio phoneme features generated by the time-domain convolutional feature extractor, the multi-layer transformer encoder can further encode the high-level features of the audio phoneme features. Here, the multi-layer transformer encoder can model the complex long-range dependencies in the feature sequence through the self-attention mechanism and pay attention to the interaction and correlation between features at different levels; after multi-layer encoding processing, the output features are integrated into a high-level semantic representation, that is, the audio demeanor features. The audio demeanor features of this application include lip movement features, expression style features, and head pose features, which can accurately describe the mapping relationship between audio and visual performances, such as how the language rhythm drives the lip movement changes, how the emotional intonation affects the expression style, and how the language intonation is accompanied by the head pose adjustment, etc.
[0075] All in all, by combining the time-domain convolutional feature extractor with the multi-layer transformer encoder, the target feature extraction model can efficiently extract multi-level feature information from the target audio data and establish a natural association from audio to action. This structural design not only enhances the multi-dimensional expression ability of the audio data, but also provides a richer and more accurate feature basis for subsequent action-driven parameter generation and portrait action driving, ensuring a high degree of coordination and naturalness of the driving effect.
[0076] In one embodiment, the process of generating audio phoneme features by performing low-level feature extraction on the target audio data through the time-domain convolutional feature extractor in step S111 may include:
[0077] S1111: Extract the time-domain audio features of the target audio data through a time-domain convolutional feature extractor, and convert the time-domain audio features into frequency-domain audio features to perform multi-dimensional feature extraction on the frequency-domain audio features to form audio phoneme features.
[0078] In this embodiment, after the computer device inputs the target audio data into the time-domain convolutional feature extractor, the time-domain convolutional feature extractor can first extract the time-domain audio features of the target audio data, and then convert the time-domain audio features into frequency-domain audio features. Furthermore, multi-dimensional feature extraction can be performed on the frequency-domain audio features to form audio phoneme features.
[0079] Specifically, the time-domain convolutional feature extractor can first perform time-domain analysis on the target audio signal and extract the time-domain audio features in the target audio data according to the analysis results, so that the time-domain audio features can be used to reflect the local changes of the target audio data in the time series. Then, the time-domain convolutional feature extractor can use the Fourier transform to convert the time-domain audio features from the form of a time-domain signal to the form of a frequency-domain signal, thereby obtaining the frequency-domain audio features, enabling the frequency-domain audio features to reveal the distribution of frequency components in the target audio data, and further capturing spectral features, such as energy changes in the low-frequency and high-frequency parts. After obtaining the frequency-domain audio features, the time-domain convolutional feature extractor can further perform multi-dimensional feature extraction on it, and extract key features related to phonemes, such as rhythm features, emotion features, and intonation features, from the frequency-domain audio features through multi-layer convolution operations and feature aggregation to form audio phoneme features, providing an encoding basis for the subsequent generation process of audio expression features.
[0080] In one embodiment, the process of encoding the audio phoneme features into high-level features by using a multi-layer transformer encoder in step S112 and outputting them as audio expression features may include:
[0081] S1121: Use the multi-head self-attention mechanism in the multi-layer transformer encoder to read the temporal dependence and cross-feature dependence of the audio phoneme features, and based on the temporal dependence and cross-feature dependence, use a layer-by-layer feature interaction weighting mechanism to perform high-level encoding on the audio phoneme features to generate audio expression features.
[0082] In this embodiment, after the computer device inputs the audio phoneme features output by the time-domain convolutional feature extractor into the multi-layer transformer encoder, the multi-layer transformer encoder can read the temporal dependence and cross-feature dependence of the audio phoneme features through the multi-head self-attention mechanism, and then based on the temporal dependence and cross-feature dependence, use a layer-by-layer feature interaction weighting mechanism to perform high-level encoding on the audio phoneme features to generate audio expression features.
[0083] It can be understood that the multi-head self-attention mechanism here can simultaneously analyze different dimensions of the input features through multiple attention heads, thereby effectively capturing the temporal dependence relationship and cross-feature dependence relationship between audio phoneme features. Among them, the temporal dependence relationship refers to the variation law of audio phoneme features in time. For example, the different dependence patterns formed by the changes in rhythm, intonation, emotional fluctuations, etc. on the time axis; the cross-feature dependence relationship refers to the mutual influence and correlation of different-level features in audio phoneme features. For example, the correlation between rhythm features and emotional features, intonation features, and how these features work together to convey the meaning and emotional color in speech.
[0084] Specifically, after obtaining the temporal dependence relationship and cross-feature dependence relationship of audio phoneme features through the multi-head self-attention mechanism, the multi-layer transformer encoder can adopt a layer-by-layer feature interaction weighting mechanism. By weighting features at different levels, it ensures that in the processing of each layer, the audio phoneme features can be effectively fused and strengthened between different layers. Therefore, in the process of high-level encoding of audio phoneme features, the multi-layer transformer encoder not only considers the processing of features in the current layer, but also can weight based on the features of the previous layer, gradually refining and enhancing the semantic expression of audio phoneme features, and then generating high-quality audio expression features, providing rich basic data for generating high-coordination and high-naturalness portrait action animations.
[0085] In one embodiment, as Figure 3 shown, Figure 3 is a schematic flowchart of the process for determining a target expression parameter generation model provided by an embodiment of the present application; Figure 3 In, the process of determining the target expression parameter generation model in step S120 may include:
[0086] S121: Construct a pre-training model using the Transformers architecture and the CNN architecture, and introduce an overlapping frame prediction technique into the pre-training model to obtain an initial expression parameter generation model.
[0087] S122: Input the pre-acquired sample audio data into the initial expression parameter generation model to obtain the predicted action driving parameters output by the initial expression parameter generation model.
[0088] S123: Aim at the predicted action driving parameters approaching the true action driving parameters of the sample audio data, and use the target loss function to train the initial expression parameter generation model.
[0089] S124: When the initial expression parameter generation model meets the preset training conditions, use the trained initial expression parameter generation model as the target expression parameter generation model.
[0090] In this embodiment, when the computer device trains the target expression parameter generation model, it can use the Transformers architecture and the CNN (Convolutional Neural Networks) architecture to construct a pre-trained model. Then, an overlapping frame prediction technique is introduced into the pre-trained model to obtain an initial expression parameter generation model. Furthermore, the initial expression parameter generation model can be trained using the pre-acquired sample audio data, and the trained initial expression parameter generation model is used as the target expression parameter generation model.
[0091] It can be understood that the Transformers architecture is good at processing sequence data and can capture long-term dependencies in the input data, while the CNN architecture excels in local feature extraction. Therefore, by integrating the advantages of these two architectures, the constructed pre-trained model has better prediction ability. Then, the computer device can further introduce the overlapping frame prediction technique into the pre-trained model, so as to ensure that the formed initial expression parameter generation model can perceive historical context information during cross-window prediction, thereby improving the fluency and accuracy of the prediction results.
[0092] Specifically, when training the initial expression parameter generation model, the computer device can first obtain the sample audio data, which is labeled with corresponding sample labels, that is, the true action driving parameters of the sample audio data. After inputting this sample audio data into the initial expression parameter generation model, the predicted action driving parameters output by the initial expression parameter generation model can be obtained. Then, the computer device can aim at making the predicted action driving parameters approach the true action driving parameters of the sample audio data, and use the target loss function to train the initial expression parameter generation model. When the initial expression parameter generation model meets the preset training conditions, such as the number of iterations reaches the set value, it is regarded as the training is completed, and the trained initial expression parameter generation model is used as the target expression parameter generation model.
[0093] Furthermore, the present application can also store the trained target expression parameter generation model in the computer device. When the computer device needs to predict the action driving parameters subsequently, it can directly call the target expression parameter generation model to predict the parameters of the input audio expression features, thereby improving the efficiency of portrait action driving.
[0094] In one embodiment, the calculation formula of the target loss function in step S120 may include:
[0095]
[0096] In the formula, represents the target loss function; Represents the true action driving parameter; Represents the predicted action driving parameter; where Represents the time window from the p-th moment to the w-th moment.
[0097] In this embodiment, the target loss function Can measure the performance of the model by calculating the squared error between the true action driving parameter and the predicted true action driving parameter. During the model iteration process, by continuously minimizing , the model can adjust its internal parameters in multiple rounds of iteration to make the predicted action driving parameter it outputs closer to the true action driving parameter, and thus enable the model to finally generate high-quality, highly coordinated and natural and smooth action driving parameters, providing a data basis for portrait action driving.
[0098] It can be understood that in the above formula, the time window Specifies the time range considered by the model during prediction, including the historical action driving parameters of the last frames of the previous window and the action driving parameters of the current window's frames. Therefore, this target loss function can enable the model to learn the continuity and correlation of action driving parameters in time, ensuring the smoothness and naturalness of the prediction results in terms of actions and expressions.
[0099] In one embodiment, the process of driving the head action of the target portrait based on the action driving parameter in step S140 to generate a portrait action animation may include:
[0100] S141: Determine the portrait action driving model; the portrait action driving model is composed of a lip movement driving module, an expression style driving module, a head pose driving module, and an animation generation module.
[0101] S142: Use the action driving parameter and the target portrait as input data, and input the input data into the lip movement driving module, the expression style driving module, and the head pose driving module respectively to obtain the first driving result output by the lip movement driving module, the second driving result output by the expression style driving module, and the third driving result output by the head pose driving module.
[0102] S143: Align and fuse the first driving result, the second driving result, and the third driving result in the animation generation module based on the target audio data to generate a portrait action animation.
[0103] In this embodiment, when the computer device performs portrait action driving, it can first determine the portrait action driving model. Since the portrait action driving model is a model that drives the portrait according to the input action driving parameters, the computer device inputs the action driving parameters and the target portrait as input data into the portrait action driving model, and can obtain the portrait action animation directly output by the portrait action driving model.
[0104] Specifically, the portrait action driving model can be composed of a lip movement driving module, an expression style driving module, a head pose driving module, and an animation generation module. Among them, the lip movement driving module is mainly responsible for driving the lip movement of the target portrait, the expression style driving module is mainly responsible for driving the expression style of the target portrait, the head pose driving module is mainly responsible for driving the head pose of the target portrait, and the animation generation module is mainly responsible for aligning and fusing the driving results of each module to generate a portrait action animation. Therefore, when the computer device performs portrait action driving, it can use the action driving parameters and the target portrait as input data, and input the input data into the lip movement driving module, the expression style driving module, and the head pose driving module respectively, to obtain the first driving result output by the lip movement driving module, the second driving result output by the expression style driving module, and the third driving result output by the head pose driving module. Finally, in the animation generation module, the first driving result, the second driving result, and the third driving result are aligned and fused based on the target audio data to generate a portrait action animation with high coordination and high naturalness.
[0105] It can be understood that through this modular structure design, the portrait action driving model can efficiently convert the action driving parameters corresponding to the target audio data into the performance of the target portrait. Specifically, the portrait action driving model can not only accurately capture the multi-level features in the audio data, but also generate smooth and coordinated dynamic effects through time frame alignment and fusion, so as to realize a portrait action animation that is synchronized with the target audio data and highly natural through independent driving and precise control of lip movement, expression style, and head pose.
[0106] Next, the portrait action driving device provided in the embodiments of the present application will be described. The portrait action driving device described below can be correspondingly referred to the portrait action driving method described above.
[0107] In one embodiment, as Figure 4 shown, Figure 4 is a schematic structural diagram of a portrait action driving device provided in an embodiment of the present application; the present application also provides a portrait action driving device, including a feature extraction module 210, a model determination module 220, a parameter prediction module 230, and an action driving module 240, specifically including the following:
[0108] A feature extraction module 210, configured to obtain target audio data and extract audio expression features of the target audio data by using a preset target feature extraction model; the audio expression features include lip movement features, expression style features, and head pose features.
[0109] A model determination module 220, configured to determine a target expression parameter generation model; the target expression parameter generation model is trained by using a target loss function based on an overlapping frame prediction technique.
[0110] A parameter prediction module 230, configured to input the audio expression features into the target expression parameter generation model to obtain action driving parameters output by the target expression parameter model.
[0111] An action driving module 240, configured to obtain a target portrait and drive the head action of the target portrait based on the action driving parameters to generate a portrait action animation.
[0112] In the above embodiment, when it is necessary to drive the action of the target portrait, the target audio data for action driving can be obtained first, and then the audio expression features of the target audio data are extracted by using a pre-trained target feature extraction model. The target feature extraction model here can capture features at multiple levels in the audio data at the same time, realize the many-to-many mapping between the lip movement features, expression style features, and head pose features in the audio expression features, and further improve the coordination of the driving effect. Then, the target expression parameter generation model can be determined. Since the model is trained based on the overlapping frame prediction technique, the action driving parameters predicted from the audio expression features by the target expression parameter generation model can achieve seamless conversion in the time series, and further improve the naturalness of the driving effect. Therefore, driving the head action of the target portrait based on the action driving parameters can generate a portrait action animation with high coordination and high naturalness. Through this method, the present application can directly copy the feature distributions in different forms in the target audio data, realize the many-to-many feature mapping, so as to avoid the limitations in aspects such as portraits and styles, and further improve the coordination and naturalness of the driving effect.
[0113] In one embodiment, the target feature extraction model in the feature extraction module 210 may include a time-domain convolutional feature extractor and a multi-layer transformer encoder; the feature extraction module 210 may further include:
[0114] A phoneme feature extraction sub-module, configured to input the target audio data into the time-domain convolutional feature extractor to perform low-level feature extraction on the target audio data through the time-domain convolutional feature extractor and generate audio phoneme features; the audio phoneme features include rhythm features, emotion features, and intonation features.
[0115] The facial expression feature encoding sub-module is used to encode high-level features of audio phoneme features using a multi-layer transformer encoder and output them as audio facial expression features.
[0116] In one embodiment, the phoneme feature extraction sub-module may include:
[0117] The feature extraction unit is used to extract the time-domain audio features of the target audio data through a time-domain convolutional feature extractor, convert the time-domain audio features into frequency-domain audio features, and perform multi-dimensional feature extraction on the frequency-domain audio features to form audio phoneme features.
[0118] In one embodiment, the phoneme feature extraction sub-module may include:
[0119] The feature encoding unit is used to read the temporal dependence and cross-feature dependence of the audio phoneme features using the multi-head self-attention mechanism in the multi-layer transformer encoder, and based on the temporal dependence and cross-feature dependence, perform high-level encoding on the audio phoneme features using a layer-by-layer feature interaction weighting mechanism to generate audio facial expression features.
[0120] In one embodiment, the model determination module 220 may include:
[0121] The model construction sub-module is used to construct a pre-trained model using the Transformers architecture and the CNN architecture, and introduce an overlapping frame prediction technique into the pre-trained model to obtain an initial expression parameter generation model.
[0122] The model prediction sub-module is used to input the pre-acquired sample audio data into the initial expression parameter generation model to obtain the predicted action driving parameters output by the initial expression parameter generation model.
[0123] The model training sub-module is used to aim at the predicted action driving parameters approaching the true action driving parameters of the sample audio data, and use the target loss function to train the initial expression parameter generation model.
[0124] The model generation sub-module is used to, when the initial expression parameter generation model meets the preset training conditions, use the trained initial expression parameter generation model as the target expression parameter generation model.
[0125] In one embodiment, the model determination module 220 may further include:
[0126]
[0127] Where, represents the target loss function; represents the true action driving parameter; represents the predicted action driving parameter; where, Represents a time window from the p-th moment to the w-th moment.
[0128] In one embodiment, the action driving module 240 may include:
[0129] A model determination sub-module for determining a portrait action driving model; the portrait action driving model is composed of a lip movement driving module, an expression style driving module, a head pose driving module, and an animation generation module.
[0130] A model driving sub-module for using the action driving parameters and the target portrait as input data, and inputting the input data into the lip movement driving module, the expression style driving module, and the head pose driving module respectively, to obtain a first driving result output by the lip movement driving module, a second driving result output by the expression style driving module, and a third driving result output by the head pose driving module.
[0131] An animation generation sub-module for aligning and fusing the first driving result, the second driving result, and the third driving result in the animation generation module based on the target audio data to generate a portrait action animation.
[0132] In one embodiment, the present application further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the portrait action driving method as described in any one of the above embodiments.
[0133] In one embodiment, the present application further provides a computer device, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the portrait action driving method as described in any one of the above embodiments.
[0134] Schematically, as Figure 5 shown, Figure 5 is an internal structural schematic diagram of a computer device provided by an embodiment of the present application. The computer device 300 may be provided as a server. Referring to Figure 5 , the computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by a memory 301 for storing instructions executable by the processing component 302, such as application programs. The application programs stored in the memory 301 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the portrait action driving method of any of the above embodiments.
[0135] The computer device 300 may further include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM, or the like.
[0136] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0137] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0138] The various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0139] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A portrait action driving method, characterized in that: The method comprises: Acquire target audio data, and use a preset target feature extraction model to extract audio facial features of the target audio data; the audio facial features include lip movement features, expression style features, and head posture features; Determine a target expression parameter generation model; the target expression parameter generation model is based on overlapping frame prediction technology and is trained using a target loss function; Inputting the audio expression feature into the target expression parameter generation model to obtain the action driving parameters output by the target expression parameter model; Acquire a target portrait, and drive the target portrait to perform a head movement based on the movement driving parameters to generate a portrait movement animation; Wherein, the target feature extraction model includes a time-domain convolutional feature extractor and a multi-layer transformer encoder; The step of extracting the audio state features of the target audio data by using a preset target feature extraction model includes: Inputting the target audio data into the time-domain convolutional feature extractor, so as to perform low-level feature extraction on the target audio data through the time-domain convolutional feature extractor to generate audio phoneme features; the audio phoneme features include rhythm features, emotion features and intonation features; The multi-head self-attention mechanism in the multi-layer transformer encoder is used to read the temporal dependency and cross-feature dependency of the audio phoneme features, and based on the temporal dependency and the cross-feature dependency, a layer-by-layer feature interaction weighting mechanism is used to perform advanced encoding on the audio phoneme features to generate audio mood features.
2. The portrait motion driving method according to claim 1, characterized in that: The step of performing low-level feature extraction on the target audio data by the time domain convolution feature extractor to generate audio phoneme features comprises: The time domain audio features of the target audio data are extracted by the time domain convolution feature extractor, and the time domain audio features are converted into frequency domain audio features, so as to perform multi-dimensional feature extraction on the frequency domain audio features to form audio phoneme features.
3. The portrait motion driving method according to claim 1, characterized in that: The step of determining a target expression parameter generation model comprises: A pre-trained model is constructed using the Transformers architecture and the CNN architecture, and an overlapping frame prediction technology is introduced into the pre-trained model to obtain an initial expression parameter generation model; Inputting the pre-acquired sample audio data into the initial expression parameter generation model to obtain the predicted action driving parameters output by the initial expression parameter generation model; The predicted action driving parameters are close to the real action driving parameters of the sample audio data, and the initial expression parameter generation model is trained using a target loss function; When the initial expression parameter generation model meets the preset training conditions, the trained initial expression parameter generation model is used as the target expression parameter generation model.
4. The portrait motion driving method according to claim 1 or 3, characterized in that: The calculation formula of the objective loss function includes: ; In the formula, represents the target loss function; represents the real action driving parameters; represents the predicted action driving parameter; where, represents the time window from the pth moment to the wth moment.
5. The portrait motion driving method according to claim 1, characterized in that: The driving of the target portrait to perform a head movement based on the movement driving parameters to generate a portrait movement animation includes: Determine a portrait action driving model; the portrait action driving model is composed of a lip action driving module, an expression style driving module, a head posture driving module and an animation generation module; The action driving parameters and the target portrait are used as input data, and the input data are respectively input into the lip action driving module, the expression style driving module and the head posture driving module to obtain a first driving result output by the lip action driving module, a second driving result output by the expression style driving module and a third driving result output by the head posture driving module; In the animation generation module, the first driving result, the second driving result and the third driving result are aligned and fused in time frames based on the target audio data to generate a portrait action animation.
6. A portrait motion driving device, characterized in that: include: A feature extraction module is used to obtain target audio data and extract audio expression features of the target audio data using a preset target feature extraction model; the audio expression features include lip movement features, expression style features and head posture features; A model determination module is used to determine a target expression parameter generation model; the target expression parameter generation model is obtained by training with a target loss function based on an overlapping frame prediction technology; A parameter prediction module, used for inputting the audio expression feature into the target expression parameter generation model to obtain the action driving parameters output by the target expression parameter model; An action driving module, used for acquiring a target portrait, and driving the head action of the target portrait based on the action driving parameters to generate a portrait action animation; Wherein, the target feature extraction model in the feature extraction module includes a time-domain convolution feature extractor and a multi-layer transformer encoder; The feature extraction module also includes: Inputting the target audio data into the time-domain convolutional feature extractor, so as to perform low-level feature extraction on the target audio data through the time-domain convolutional feature extractor to generate audio phoneme features; the audio phoneme features include rhythm features, emotion features and intonation features; The multi-head self-attention mechanism in the multi-layer transformer encoder is used to read the temporal dependency and cross-feature dependency of the audio phoneme features, and based on the temporal dependency and the cross-feature dependency, a layer-by-layer feature interaction weighting mechanism is used to perform advanced encoding on the audio phoneme features to generate audio mood features.
7. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the portrait motion driving method according to any one of claims 1 to 5.
8. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the portrait motion driving method according to any one of claims 1 to 5 are performed.
Citation Information
Patent Citations
Image processing method and device based on audio driving and storage medium
CN117974850A
Audio-driven video generation method and device, computer equipment and storage medium
CN118413722A