Digital human video generation method and device and computer equipment

CN120807731AActive Publication Date: 2025-10-17SHENZHEN PEMI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510932640.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-17
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

When existing technologies process complex contexts and emotional expressions, there is a problem in the generation of digital human videos in which lip movements and facial expressions cannot be fully synchronized.

Method used

By acquiring text and audio information, performing semantic parsing and speech feature extraction, generating facial expression parameter sequences and lip shape parameter sequences, and generating a stylized 3D face model through style control parameters, driving the expression and lip shape, and finally performing time alignment and synchronous rendering to generate a digital human video.

Benefits of technology

It achieves a high degree of synchronization between facial expressions and lip movements in digital human videos, improves naturalness and accuracy, can adjust the appearance and performance style of digital humans according to different needs, and improves synchronization issues in complex contexts and emotional expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807731A_ABST
    Figure CN120807731A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital humans, and provides a digital human video generation method and device and computer equipment, and the method comprises the steps: obtaining text information, audio information and style control parameters, carrying out the semantic analysis and feature extraction of the text information and the audio information, and obtaining a digital human video; after obtaining text-driven expression features and audio-driven expression features, carrying out cooperative modulation to obtain a facial expression parameter sequence and a mouth shape parameter sequence; and generating a stylized three-dimensional face model of the digital human based on the style control parameters, driving the stylized three-dimensional face model by using the facial expression parameter sequence and the mouth shape parameter sequence, and generating a corresponding digital human video. Through cooperative modulation of the text information, the audio information and the style control parameters, synchronization of facial expressions and mouth shapes is enhanced, naturalness and accuracy of digital human video generation are improved, and the problem that the mouth shapes and the facial expressions cannot be completely synchronized when complex contexts and emotion expressions are processed is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital humans, and in particular to a digital human video generation method and device and a computer device. BACKGROUND

[0002] With the rapid development of artificial intelligence and computer graphics technology, digital human technology has been widely applied in many fields. As a representative of virtual characters, digital humans can provide efficient and interactive solutions in virtual reality, film production, game development, intelligent customer service and other industries. The generation of digital humans involves multiple disciplines such as graphics, computer vision, speech synthesis, natural language processing, etc. In particular, in the film production, advertising, education and entertainment industries, the application of digital humans has gradually become one of the core technologies.

[0003] In related technical means, digital human video generation technology generally relies on recording real person performances or generating text-based speech, and then combining these data with three-dimensional face models through computer graphics technology to match expressions, mouth shapes and actions. Common implementation methods include text-based speech and facial expression synthesis, generating corresponding facial animations through speech-driven technology, and then synchronizing these facial animations with target audio to enable virtual characters to interact with preset dialogue content.

[0004] For the above technical solutions, although the digital human video generation technology based on speech and expression driving can effectively realize the synchronization of virtual characters and audio content, in dealing with complex contexts and emotional expressions, the existing technology often appears the problem of incoordination between speech and facial expressions, especially in complex or variable emotional expression scenarios, there is a problem that the mouth shape and facial expression cannot be completely synchronized. SUMMARY

[0005] In order to improve the problem that the mouth shape and facial expression cannot be completely synchronized when dealing with complex contexts and emotional expressions, the present application provides a digital human video generation method and device and a computer device.

[0006] The present invention provides a method for generating a digital human video, comprising: acquiring text information, audio information and style control parameters, performing semantic analysis on the text information to obtain text-driven expression features; performing speech feature extraction on the audio information to obtain a rhythm feature vector and a pronunciation dynamic vector, fusing the rhythm feature vector with the pronunciation dynamic vector to obtain an audio-driven expression feature; co-modulating the text-driven expression feature and the audio-driven expression feature to obtain a facial expression parameter sequence and a lip shape parameter sequence; generating a stylized three-dimensional face model of the digital human based on the style control parameters, performing expression-driven operation on the stylized three-dimensional face model using the facial expression parameter sequence to obtain a facial animation sequence, performing lip shape-driven operation on the stylized three-dimensional face model using the lip shape parameter sequence to obtain a lip shape animation sequence; performing time alignment processing on the facial animation sequence and the lip shape animation sequence to obtain complete animation action data, synchronously rendering the animation action data and the audio information to generate a corresponding digital human video.

[0007] As a preferred solution, the step of performing semantic analysis on the text information to obtain text-driven expression features includes: performing sentence processing on the text information to obtain a sentence unit set and inter-sentence dependency relations, extracting intra-sentence semantic features based on the sentence unit set, and extracting a global semantic context vector and inter-sentence sentiment flow features according to the inter-sentence dependency relations; fusing the intra-sentence semantic features with the global semantic context vector to obtain a multi-level semantic feature vector, performing cross-attention mapping on the inter-sentence sentiment flow features and the multi-level semantic feature vector to obtain a sentiment-enhanced semantic vector and a sentiment feature representation; using the sentiment feature representation to perform style adjustment on the sentiment-enhanced semantic vector to obtain a sentiment-enhanced semantic vector that is consistent with the text meaning. Figure One The semantic feature vector and the emotional feature vector are fused to obtain text-driven expression features.

[0008] Preferably, the step of performing speech feature extraction on the audio information to obtain a prosody feature vector and a pronunciation dynamic vector, and fusing the prosody feature vector and the pronunciation dynamic vector to obtain audio-driven expression features comprises: performing frame division processing on the audio information to extract time domain envelope features and frequency domain spectrum features of each frame, extracting short-time prosody patterns based on the time domain envelope features, fusing the short-time prosody patterns and rhythm information of adjacent frames to obtain a time-continuous prosody feature vector, extracting frequency domain pronunciation features using the frequency domain spectrum features, performing principal component decoupling processing on the frequency domain pronunciation features to extract a pronunciation state sequence and pronunciation boundary change information, encoding and mapping the pronunciation state sequence and the pronunciation boundary change information to obtain a pronunciation dynamic vector, and fusing the time-continuous prosody feature vector and the pronunciation dynamic vector in a time dimension to obtain audio-driven expression features.

[0009] Preferably, the step of modulating the text-driven expression features and the audio-driven expression features to obtain a facial expression parameter sequence and a lip shape parameter sequence comprises: inputting the text-driven expression features into an expression channel encoder to extract a multi-level facial expression driving vector, inputting the audio-driven expression features into a lip shape channel encoder to extract a multi-frequency pronunciation mapping vector, performing time sequence alignment processing on the multi-level facial expression driving vector and the multi-frequency pronunciation mapping vector to obtain expression feature pairs on the same time axis, constructing a fusion expression state graph based on the expression feature pairs, performing feature decoding on the fusion expression state graph to obtain facial expression parameters and lip shape parameters, and modulating the facial expression parameters and the lip shape parameters to eliminate expression conflicts and rhythm misplacement to obtain a facial expression parameter sequence and a lip shape parameter sequence.

[0010] Preferably, the step of generating a stylized three-dimensional face model of the digital human based on the style control parameter, driving the stylized three-dimensional face model with the facial expression parameter sequence to obtain a facial animation sequence, and driving the stylized three-dimensional face model with the lip shape parameter sequence to obtain a lip shape animation sequence comprises: loading a corresponding basic three-dimensional face model and style mapping weight from a preset digital human parameter template library based on the style control parameter, performing style deformation and material mapping on the basic three-dimensional face model using the style mapping weight to obtain a stylized three-dimensional face model; inputting the facial expression parameter sequence into a preset muscle control module in the stylized three-dimensional face model to drive changes of facial expression control points in the stylized three-dimensional face model to obtain an expression driving result; inputting the lip shape parameter sequence into a preset lip deformation module in the stylized three-dimensional face model to drive geometric deformation of a lip and tooth region in the stylized three-dimensional face model to obtain a lip shape driving result; performing frame-by-frame pose fusion processing on the expression driving result to generate a facial animation sequence, and performing frame-by-frame pose fusion processing on the lip shape driving result to generate a lip shape animation sequence.

[0011] Preferably, the step of performing time alignment processing on the facial animation sequence and the lip shape animation sequence to obtain complete animation action data, and synchronously rendering the animation action data and the audio information to generate a corresponding digital human video comprises: performing time axis resampling on the facial animation sequence and the lip shape animation sequence, and establishing a uniform frame rate and a reference starting frame based on the resampling result, performing time alignment processing on the facial animation sequence and the lip shape animation sequence using the uniform frame rate and the reference starting frame to extract an expression action vector and a lip shape action vector of each frame; performing fusion weight modeling on the expression action vector and the lip shape action vector to construct a fusion driving vector sequence, inputting the fusion driving vector sequence into a preset action combination module in the stylized three-dimensional face model to generate complete animation action data; performing alignment slicing processing on the audio information to extract an audio frame segment corresponding to a frame, and performing frame-level synchronous rendering on the animation action data and the audio frame segment to generate a corresponding digital human video.

[0012] Preferably, the step of aligning and slicing the audio information, extracting audio frame segments corresponding to frames, and frame-level synchronously rendering the animation action data and the audio frame segments to generate a corresponding digital human video comprises: dividing the audio information into audio frame segments with the same time interval according to a preset frame rate, extracting audio features from each audio frame segment in the audio frame segment set to obtain time-domain audio features and frequency-domain audio features; constructing an audio feature vector according to the time-domain audio features and the frequency-domain audio features, time-aligning the audio feature vector with a motion feature vector of a corresponding frame in the animation action data to obtain an aligned audio feature vector; frame-level weight adjusting the aligned audio feature vector and the motion feature vector to obtain synchronous rendering parameters, inputting the synchronous rendering parameters into a preset rendering module for audio-video synchronous rendering to generate a corresponding digital human video.

[0013] The application further provides a digital human video generation device, comprising: an acquisition module configured to acquire text information, audio information, and style control parameters, and perform semantic analysis on the text information to obtain text-driven expression features; a fusion module configured to extract speech features from the audio information to obtain prosody feature vectors and pronunciation dynamic vectors, and fuse the prosody feature vectors and the pronunciation dynamic vectors to obtain audio-driven expression features; a modulation module configured to modulate the text-driven expression features and the audio-driven expression features cooperatively to obtain a facial expression parameter sequence and a lip shape parameter sequence; a driving module configured to generate a stylized three-dimensional face model of a digital human based on the style control parameters, drive the stylized three-dimensional face model with the facial expression parameter sequence to obtain a facial animation sequence, and drive the stylized three-dimensional face model with the lip shape parameter sequence to obtain a lip shape animation sequence; and a generation module configured to perform time alignment processing on the facial animation sequence and the lip shape animation sequence to obtain complete animation action data, and synchronously render the animation action data and the audio information to generate a corresponding digital human video.

[0014] The application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, and the processor implements the digital human video generation method according to any one of the above embodiments when executing the computer program.

[0015] Compared with the prior art, the application has the following beneficial effects: high synchronization. Through the cooperative modulation of text information and audio information, natural and synchronous facial expressions and mouth shape animations are generated according to the text and audio, realizing high coordination between the digital person and the text information and the audio information. At the same time, the introduction of the style control parameter can adjust the appearance and performance style of the digital person according to different needs, making it more personalized and expressive, enhancing the synchronization of facial expressions and mouth shapes, improving the naturalness and accuracy of digital person video generation, and improving the problem that the mouth shape and facial expression cannot be completely synchronized when processing complex context and emotional expression. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0017] The structures, proportions, sizes, etc. shown in the drawings of the present specification are only used to cooperate with the content disclosed in the present specification, to enable those skilled in the art to understand and read, and are not used to limit the conditions that the present application can be implemented, so they do not have technical significance. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effects that the present application can produce and the purposes that the present application can achieve, should still fall within the scope of the technical content disclosed by the present application.

[0018] Figure 1 is a flowchart of the digital person video generation method provided by the embodiment of the present application; Figure 2 is a structural schematic block diagram of the digital person video generation device provided by the embodiment of the present application; Figure 3 is a structural schematic block diagram of the computer device provided by the embodiment of the present application.

[0019] Explanation of reference signs: 10, digital person video generation device; 11, acquisition module; 12, fusion module; 13, modulation module; 14, driving module; 15, generation module; 20, computer device; 21, memory; 22, processor. DETAILED DESCRIPTION

[0020] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.

[0021] The flowcharts shown in the drawings are only illustrative, and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further decomposed, combined or partially merged, so that the actual execution order can be changed according to actual situations.

[0022] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clear from the context, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0023] It should be further understood that the term "and / or" used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0024] The technical solutions of the present application will be further described below with reference to the drawings and through specific embodiments.

[0025] Embodiment 1: As shown in the figure, the present application provides a digital human video generation method, including steps S100 to S500. Figure 1

[0026] Step S100, obtaining text information, audio information and style control parameters, performing semantic analysis on the text information to obtain text-driven expression features.

[0027] In this step, first, the input text information is obtained through a natural language processing module, and the text is preprocessed through methods such as word segmentation and syntax analysis; specifically, the syntax structure in the text is extracted through a syntax dependency analysis module, a set of sentence elements is generated, and a BERT (Bidirectional Encoder Representations from Transformers) model is used to understand the semantics of the text to obtain deep semantic features of each sentence. The BERT model is a deep learning model commonly used in natural language processing tasks, which can effectively capture the relationship between words to obtain more accurate semantic representations.

[0028] ​For example, given the input text "He said happily: 'It's a good weather today!'", the BERT model can extract the sentiment vector of the text, such as "happiness" is associated with "good weather", to generate a semantic vector with emotional color.

[0029] Step S200, speech feature extraction is performed on the audio information to obtain prosody feature vector and pronunciation dynamic vector, and the prosody feature vector and the pronunciation dynamic vector are fused to obtain audio-driven expression features.

[0030] In this step, the audio information is first sent to the acoustic feature extraction module, and the short-time Fourier transform (STFT) is used to extract the frequency domain features of the audio. Specifically, the audio signal is first processed by frame, then the time domain envelope and frequency domain spectrum of each frame are calculated, and the audio features of each frame are obtained. By using an acoustic model (such as a deep neural network DNN or a long short-term memory network LSTM), the prosody features are extracted to obtain the prosody feature vector of the audio. In addition, the pronunciation dynamic vector is obtained by modeling the pronunciation state of the audio, and the MFCC (Mel-Frequency Cepstral Coefficients) is used to extract the pronunciation features of the audio, combined with the speech information before and after, to obtain the dynamic change vector of the pronunciation. The prosody feature vector and the pronunciation dynamic vector are fused by weighted average method to obtain the audio-driven expression features.

[0031] For example, in the audio information "He said happily: 'It's a good weather today!'", by extracting its pitch, intonation, speech rate and other prosodic features, as well as the boundaries and pause information of pronunciation, an audio-driven expression feature vector can be generated to describe the prosodic changes in the audio.

[0032] Step S300, the text-driven expression features and the audio-driven expression features are modulated cooperatively to obtain the facial expression parameter sequence and the mouth shape parameter sequence.

[0033] In this step, the text-driven expression features and the audio-driven expression features are fused through a dual-channel cooperative modulation network. Specifically, a fusion mechanism based on the Transformer model is used to interact the sentiment features of the text and the prosodic features of the audio in multiple channels to generate a global expression feature. Then, this global expression feature is sent to the facial expression generation module and the mouth shape generation module. Through a shared neural network model, the facial expression parameter sequence and the mouth shape parameter sequence are generated respectively.

[0034] For example, in the text "He said happily: 'It's a good weather today!'", the Transformer model will capture the emotion of "happiness", and at the same time generate appropriate facial expression and mouth shape parameters according to the "tone and speed of speaking" in the audio, to ensure that the facial expression and mouth shape are synchronized with the speech content.

[0035] Step S400, generate a stylized three-dimensional face model of the digital person based on the style control parameters, drive the stylized three-dimensional face model with the facial expression parameter sequence to obtain a facial animation sequence, and drive the stylized three-dimensional face model with the lip shape parameter sequence to obtain a lip shape animation sequence.

[0036] In this step, first, a suitable three-dimensional model of the digital person is selected according to the style control parameters, and the shape and material are adjusted according to the user's set style. Specifically, the style control parameters include facial features (such as the shape of the eyes and mouth), facial expression style (such as exaggerated or natural), etc., and the style transfer network is used to stylize the basic three-dimensional model. Then, the stylized three-dimensional face model is driven by the facial expression parameter sequence and the lip shape parameter sequence, respectively. The expression driving adjusts the facial muscles through the muscle deformation module, and the lip shape driving controls the lip movement through the lip deformation module, thereby generating a facial animation sequence and a lip shape animation sequence.

[0037] For example, by adjusting the style control parameters, the shape of the digital person's mouth can be adjusted to a smiling or closed mouth state, and the facial expression produces corresponding eyebrow lifting, eye smiling, and other expression features according to the "happy" emotion.

[0038] Step S500, time align the facial animation sequence and the lip shape animation sequence to obtain complete animation action data, and synchronize the animation action data with the audio information for rendering to generate a corresponding digital person video.

[0039] In this step, the facial animation sequence and the lip shape animation sequence are time aligned to ensure their accurate matching on the time axis. Specifically, the time axis of the facial expression and the lip shape action is aligned through the dynamic time warping (DTW) algorithm to ensure the accurate synchronization of the lip shape and the audio content. Then, the synchronized animation data and the audio information are time domain synthesized, and the digital person video is generated through the rendering engine. In the rendering process, according to the facial animation and the lip shape animation of each frame, the time domain information of the audio is combined to generate the final digital person video frame sequence, which is post-processed and enhanced to improve the picture quality.

[0040] For example, through the dynamic time warping algorithm, the pauses and accents in the audio can be accurately synchronized with the micro-expression changes in the facial animation, ensuring that the lip shape of the digital person in the final video completely matches the speech content and the expression is natural.

[0041] In this embodiment, by acquiring text information, audio information and style control parameters, first, the text information is semantically analyzed to obtain text-driven expression features. Then, the audio information is used to extract speech features to obtain prosodic feature vectors and pronunciation dynamic vectors, and the two are fused to obtain audio-driven expression features. Then, the text-driven expression features and the audio-driven expression features are modulated in coordination to obtain facial expression parameter sequences and mouth shape parameter sequences. Based on the style control parameters, a stylized three-dimensional face model of a digital person is generated, and the facial expression parameter sequence is used to drive the expression of the stylized three-dimensional face model to obtain a facial animation sequence; the mouth shape parameter sequence is used to drive the mouth shape of the stylized three-dimensional face model to obtain a mouth shape animation sequence. Finally, the facial animation sequence and the mouth shape animation sequence are time-aligned to obtain complete animation action data, and the animation action data and the audio information are synchronized to render a corresponding digital person video. The high coordination between the digital person and the text information and the audio information is realized, especially in the synchronization of facial expressions and mouth shapes, which improves the naturalness and accuracy of digital person video generation and improves the problem of incomplete synchronization of mouth shapes and facial expressions when dealing with complex contexts and emotional expressions. In addition, the introduction of the style control parameter can adjust the appearance and performance style of the digital person according to different needs, making it more personalized and expressive. Through the coordinated modulation of the text-driven expression features and the audio-driven expression features, the conflict between the mouth shape and the expression can be effectively eliminated, the performance ability of the digital person in multiple emotions and multiple contexts is improved, and the synchronization problem in the existing technology in complex emotions and speech expression is overcome. A more accurate and flexible solution is provided for the application of virtual characters in film and television production, game interaction and artificial intelligence field.

[0042] Embodiment 2 In step S100, the text information is semantically analyzed to obtain text-driven expression features, which specifically includes: The text information is processed by sentence to obtain a sentence element set and an inter-sentence dependency relationship, intra-sentence semantic features are extracted from the sentence element set based on a multi-layer semantic encoding network, and global semantic context vectors and inter-sentence emotional flow features are extracted from the inter-sentence dependency relationship based on a context modeling network.

[0043] Intratextual semantic features are extracted by a multi-layer semantic encoding network. Specifically, in this step, the input text is bidirectionally encoded by a BERT model (Bidirectional Encoder Representations from Transformers) to generate context information for each sentence and capture intratextual semantic relationships. When processing text, the BERT model can simultaneously consider the left and right context of a word, thereby generating a deep semantic representation of each word. This model can effectively understand the relationships between individual words in a set of sentence elements and generate intratextual semantic feature vectors.

[0044] For example, when the input text is “He said happily: ‘It’s a good weather today!’”, the BERT model can identify the semantic relationship between “happiness” and “good weather” and generate corresponding semantic feature vectors for these words, reflecting the emotional color in the text.

[0045] Global semantic context vectors and inter-sentence sentiment flow features are extracted by a context modeling network. Specifically, in this step, an LSTM (Long Short-Term Memory network) or a Transformer model is used to process inter-sentence dependencies in the text. LSTM can record long-span context information through memory cells, while Transformer can capture the relationships between parts of a sentence through self-attention mechanisms to generate global semantic context vectors. At the same time, the model extracts inter-sentence sentiment flow features to represent the transfer relationship of sentiment between sentences.

[0046] For example, when processing the text “He said happily: ‘It’s a good weather today!’”, the LSTM or Transformer model can capture the transfer of “happiness” sentiment from the previous sentence to the sentence “It’s a good weather today” and generate global semantic context vectors and inter-sentence sentiment flow features containing sentiment information.

[0047] Intratextual semantic features are fused with global semantic context vectors to obtain multi-level semantic feature vectors. Inter-sentence sentiment flow features are cross-attention mapped with multi-level semantic feature vectors to obtain sentiment-enhanced semantic vectors and emotional feature representations.

[0048] Fusion is performed through a cross-attention mechanism. Specifically, the intratextual semantic features are weighted and fused with the global semantic context vectors using a multi-head self-attention mechanism. The multi-head attention mechanism can focus on different parts of the information simultaneously through multiple independent attention calculation paths, thereby obtaining multi-level context information. Then, the inter-sentence sentiment flow features are cross-mapped with the fused semantic features to ensure that sentiment information can effectively influence the expression of the text, generating sentiment-enhanced semantic vectors and emotional feature representations.

[0049] For example, in processing the text "He said happily: 'It's a good day!'", the cross-attention mechanism combines the emotional information of "happiness" with the content of "good weather" in the text, and the finally generated emotion-enhanced semantic vector will highlight the positive emotion, ensuring that the digital person's expression and tone in the generated digital person video accurately reflect this emotion.

[0050] The emotion-enhanced semantic vector is style-adjusted using an emotion feature representation to obtain a semantic feature vector and an emotion feature vector corresponding to the text intent. Figure One

[0051] The style-adjusted network is adjusted, specifically, a style transfer network is used to adjust the emotion-enhanced semantic vector. The style-adjusted network adjusts the activation value, weight, and other parameters, so that the emotion-enhanced semantic vector is more in line with the preset style requirements when expressing, for example, adjusting the tone of the voice or the expression style of the digital person. This ensures that the generated semantic features and emotion features are highly consistent with the emotional intent of the input text.

[0052] For example, in the text "He said happily: 'It's a good day!'", the style adjustment will ensure that the "happiness" emotion in the emotion-enhanced semantic vector is amplified, thereby generating a digital person expression feature that conforms to the "happy" expression and positive emotion.

[0053] The semantic feature vector and the emotion feature vector are fused to obtain a text-driven expression feature.

[0054] Through weighted fusion, specifically, the semantic feature vector and the emotion feature vector are fused in a weighted superposition manner, ensuring that the semantic and emotional content of the text can complement and strengthen each other when generating the text-driven expression feature. The fused expression feature will accurately reflect the text information and drive the digital person's expressions, movements, and other forms of expression.

[0055] For example, for the sentence "It's a good day!", the final text-driven expression feature will include emotional components such as "happy" and "positive", and will drive the digital person to display natural facial expressions such as a cheerful smile and a smile.

[0056] In step S200, the audio information is subjected to speech feature extraction to obtain a prosody feature vector and a pronunciation dynamic vector, and the prosody feature vector and the pronunciation dynamic vector are fused to obtain an audio-driven expression feature. The step specifically includes: The audio information is subjected to frame division processing, and the time domain envelope feature and the frequency domain spectrum feature of each frame are extracted. Based on a preset acoustic feature extraction network, the short-time prosody pattern is extracted from the time domain envelope feature, and the short-time prosody pattern is fused with the rhythm information of adjacent frames to obtain a time-continuous prosody feature vector. ​

[0057] Through frame division and feature extraction, specifically, the audio signal is processed by short-time Fourier transform (STFT) for frame division, usually with a frame length of 20 ms and a frame shift of 10 ms. Through STFT, the audio signal of each frame is converted into time-domain envelope features and frequency-domain spectrogram features. The time-domain envelope features are used to reflect the amplitude variation of the audio signal, while the frequency-domain spectrogram features provide the frequency information of the audio signal. Then, the time-domain envelope features are processed based on a preset acoustic feature extraction network (such as a network based on a convolutional neural network CNN) to extract short-time prosody patterns. Next, the short-time prosody patterns are fused with the rhythm information of adjacent frames to obtain a prosody feature vector containing time continuity.

[0058] For example, for the input audio signal "today is a good weather", multiple audio frames are obtained through STFT frame division, and time-domain envelope features and frequency-domain spectrogram features are extracted for each frame. The acoustic feature extraction network further processes the time-domain envelope features to extract rhythm patterns and fuse them with the rhythm information of adjacent frames to finally generate a prosody feature vector containing time continuity. These prosody features describe the rhythm, pitch variation, and other prosodic characteristics in the audio signal.

[0059] Frequency-domain pronunciation features are extracted using frequency-domain spectrogram features; principal component decoupling processing is performed on the frequency-domain pronunciation features to extract pronunciation state sequences and pronunciation boundary change information, and the pronunciation state sequences and pronunciation boundary change information are encoded and mapped to obtain pronunciation dynamic vectors.

[0060] Through frequency-domain feature processing, specifically, the frequency-domain spectrogram features are processed by the Mel frequency cepstral coefficient (MFCC) algorithm to extract frequency-domain pronunciation features. MFCC is widely used in speech processing and can extract frequency spectrum features representing pronunciation information, reflecting the acoustic characteristics of audio. Then, principal component analysis (PCA) or other dimension reduction algorithms are used to perform principal component decoupling on the extracted frequency-domain pronunciation features to extract pronunciation state sequences and pronunciation boundary change information. The pronunciation state sequence represents the state change of pronunciation, and the pronunciation boundary change information represents the pause, sentence break, and other information of speech. Through encoding and mapping, these features are converted into pronunciation dynamic vectors to drive lip synthesis.

[0061] For example, after MFCC processing of the audio signal "today is a good weather", frequency-domain pronunciation features reflecting the frequency characteristics of different syllables can be obtained. Then, PCA is used to reduce the dimension of the pronunciation features to extract the pronunciation state change between syllables (such as the pronunciation state change from "today" to "is"), which is then converted into a pronunciation dynamic vector through encoding and mapping.

[0062] The time-continuous prosody feature vector and the pronunciation dynamic vector are fused in the time dimension using an attention mechanism to obtain audio-driven expression features.

[0063] Through time dimension fusion, specifically, the prosody feature vector and the pronunciation dynamic vector are fused in the time dimension using a multi-head self-attention mechanism (Multi-head Self-Attention). The attention mechanism can effectively combine the prosody information of each frame with the corresponding pronunciation state and pronunciation boundary information to generate a global audio-driven expression feature. During the fusion process, the attention mechanism automatically allocates weights according to the time characteristics of the audio signal, ensuring that important pronunciation information and prosody information can be effectively combined to obtain more accurate audio-driven expression features.

[0064] For example, for the audio signal "today is a good weather", the prosody feature and the pronunciation dynamic feature are fused through the attention mechanism. In this fusion process, the rhythm, tone, etc. of the two syllables "today" and "good weather" in the audio are combined with the corresponding pronunciation dynamics (such as changes in mouth shape, pauses in speech), thereby obtaining an accurate audio-driven expression feature.

[0065] In step S300, the step of modulating the text-driven expression feature and the audio-driven expression feature to obtain the facial expression parameter sequence and the mouth shape parameter sequence, specifically includes: The text-driven expression feature is input into an expression channel encoder to extract a multi-level facial expression driving vector, and the audio-driven expression feature is input into a mouth shape channel encoder to extract a multi-frequency pronunciation mapping vector.

[0066] The multi-level facial expression driving vector and the multi-frequency pronunciation mapping vector are extracted through the expression channel encoder and the mouth shape channel encoder. Specifically, the expression channel encoder adopts a combined model based on a convolutional neural network (CNN) and a Transformer architecture. After the text-driven expression feature is input into the expression channel encoder, low-level emotional features (such as emotional tone) are first extracted through a convolutional layer, and then a global emotional pattern is captured by a Transformer module, finally generating a multi-level facial expression driving vector. At the same time, the audio-driven expression feature is input into the mouth shape channel encoder, which is based on an LSTM (Long Short-Term Memory Network) model, captures the dynamic changes of pronunciation, and generates a multi-frequency pronunciation mapping vector through segmented spectral analysis. The time-dependent mechanism of LSTM enables it to extract the evolution law of pronunciation from the audio-driven expression feature.

[0067] For example, given the text-driven expression feature "happy emotional tone" and the audio signal "today is a good weather", the expression channel encoder extracts multi-level facial expression features of "happy" emotion such as mouth opening, eye changes, etc. from the text-driven expression feature; the mouth shape channel encoder extracts the pronunciation characteristics within the frequency band such as the time sequence information of the mouth closing, opening and tongue movement from the audio-driven expression feature.

[0068] The multi-level facial expression driving vector and the multi-frequency band pronunciation mapping vector are time-aligned to obtain expression feature pairs on the same time axis, a fusion expression state graph is constructed based on the expression feature pairs, and the facial expression parameters and the mouth shape parameters are obtained by feature decoding of the fusion expression state graph.

[0069] The multi-level facial expression driving vector and the multi-frequency band pronunciation mapping vector are time-aligned by using a dynamic time warping (DTW) algorithm. DTW can effectively align feature sequences with inconsistent time lengths, ensuring the consistency of the two sequences in the time dimension. Then, the aligned expression features are used to construct a fusion expression state graph. The fusion expression state graph combines the emotional and action information of all key nodes on the time axis of expression and mouth shape. Subsequently, the fusion expression state graph is feature-decoded by a graph convolutional network (GCN) to generate facial expression parameters (control parameters for facial muscles) and mouth shape parameters (pronunciation parameters for driving the tongue, lips, etc.).

[0070] For example, in the input text expression "today is a good weather", the multi-level facial expression driving vector representing "happy" emotion and the multi-frequency band pronunciation dynamic vector on the corresponding time axis are aligned by DTW to generate a fusion expression state graph, which allocates mouth opening characteristics and eye squinting expression characteristics to the two words "today" and assigns specific tongue, lip closing characteristics to "good weather", and finally predicts facial expression parameters and mouth shape parameters.

[0071] The facial expression parameters and the mouth shape parameters are modulated in coordination to eliminate expression conflicts and rhythm misalignment, and the facial expression parameter sequence and the mouth shape parameter sequence are obtained.

[0072] By utilizing a collaborative modulation model based on a multi-modal generative adversarial network (MM-GAN), input facial expression parameters and mouth shape parameters are jointly optimized to actively detect potential conflicts between expressions and mouth shapes. For example, certain expression characteristics may produce unnatural facial tension, or mouth shapes may not be synchronized with facial expressions when expressing specific pronunciations. The generative network of the MM-GAN is responsible for generating consistency parameters for expressions and mouth shapes, while the adversarial discriminant network is used to evaluate whether the generated results are harmonious and natural. In the process of adversarial optimization, conflicts between expressions and mouth shapes (such as exaggerated expressions or mismatched mouth shapes) are fed back and adjusted through a loss function, ultimately generating natural and coherent facial expression parameter sequences and mouth shape parameter sequences.

[0073] For example, in "Today is a good weather", if the smile expression generated by the expression parameter sequence conflicts with the closed mouth shape of the "weather" syllable pronunciation in the mouth shape parameter sequence, the MM-GAN will detect and optimize the smile amplitude of the expression to be consistent with normal speaking, while adjusting the smoothness of the mouth shape when connecting vowels and consonants, ultimately generating both natural and coherent expression parameter sequences and mouth shape parameter sequences.

[0074] In step S400, a stylized three-dimensional face model of a digital person is generated based on a style control parameter, the stylized three-dimensional face model is driven by a facial expression parameter sequence to obtain a facial animation sequence, and the stylized three-dimensional face model is driven by a mouth shape parameter sequence to obtain a mouth shape animation sequence. The specific steps include: A corresponding basic three-dimensional face model and style mapping weight are loaded from a preset digital person parameter template library based on the style control parameter, and the basic three-dimensional face model is subjected to style deformation and material mapping using the style mapping weight to obtain the stylized three-dimensional face model.

[0075] By loading the basic three-dimensional face model and the style mapping weight, specifically, first, according to the style control parameter set by the user, a basic three-dimensional face model is selected from the digital person parameter template library. The template library contains various three-dimensional models, such as "male", "female", "child", and basic models of different races or ages. Then, the corresponding style mapping weight is loaded according to the style control parameter, which includes style-related deformation weights (such as facial proportion adjustment) and material weights (such as skin texture, eye color, etc.). Through the style transfer module in the GAN generative adversarial network, the basic three-dimensional face model is subjected to style deformation and material mapping, and finally a stylized three-dimensional face model that meets the style requirements is generated.

[0076] For example, when the user selects the style control parameter "young woman" and "fair skin", the template library will load a basic model "female character", and adjust the shape ratio of eye size, face width, etc. through style mapping weight, while optimizing the skin texture to make the skin appear fair and flawless. The finally generated stylized three-dimensional face model has the characteristics of young women and the material effect that meets the personalized requirements.

[0077] The facial expression parameter sequence is input into the preset muscle control module of the stylized three-dimensional face model to drive the facial expression control points in the stylized three-dimensional face model to change, and an expression driving result is obtained.

[0078] The muscle control module drives the expression control points. Specifically, the facial expression parameter sequence is input into the facial muscle control module in the stylized three-dimensional face model to realize expression change. The muscle control module is built based on Blendshape technology and skeletal animation technology. Each point represents the variation amplitude of a key area of the face (such as eyebrows, corners of the mouth, and corners of the eyes), and by decoding the parameters to the corresponding Blendshape weight, the specific expression control points in the stylized three-dimensional face model are driven to change. According to the time sequence of the expression parameters, the expression driving result is gradually generated.

[0079] For example, the input expression parameter sequence corresponds to the emotion of "joy", and by driving the facial muscle control module, the corners of the mouth on the model will gradually lift up, the eye muscles will present a natural slight squinting state, and the eyebrows will slightly lift up. The final expression driving result produces a facial expression matching the "joy" emotion, showing an obvious smile.

[0080] The mouth shape parameter sequence is input into the preset lip deformation module of the stylized three-dimensional face model to drive the geometric deformation of the lip region in the stylized three-dimensional face model, and a mouth shape driving result is obtained.

[0081] The lip deformation module drives the change of the lip region. Specifically, the mouth shape parameter sequence is input into the lip deformation module of the stylized three-dimensional face model, which is based on the MorphTarget Animation method and the tongue-tooth dynamic binding technology to drive the mouth position, mouth opening angle, tongue movement path, etc. in real time. Through the geometric deformation of the lip region, the pronunciation dynamic sequence is mapped to the real mouth shape change, which can naturally transition from vowels to consonants.

[0082] For example, in the pronunciation dynamic sequence, the phonemes of "today is a good weather" correspond to the mouth shape driving result. "Today" corresponds to the opening of the closed lips; "day" corresponds to slightly opening the mouth and showing the tongue tip; "good weather" corresponds to the smooth lip-tooth change of consecutive vowel phonemes. The finally generated mouth shape driving result can be synchronized with the actual pronunciation content.

[0083] The expression-driven result is subjected to frame-by-frame pose fusion processing to generate a facial animation sequence, and the lip-driven result is subjected to frame-by-frame pose fusion processing to generate a lip animation sequence.

[0084] The animation sequence is generated through frame-by-frame pose fusion processing. Specifically, a pose processing mechanism based on a time series prediction model (such as a GRU, Gated Recurrent Unit) and optimized weight fusion is used to perform frame-by-frame calculation on the expression-driven result and the lip-driven result. This model can predict the pose transition between each frame and optimize the motion coordination between the expression region and the lip region. Through weight distribution, the pose results of the facial animation and the lip animation are accurately fused to generate a coherent and natural animation sequence.

[0085] For example, when generating the facial animation sequence, the process of "gradually deepening the smile" is controlled through frame-by-frame pose prediction, while the dynamic relationship between the eyebrows and the corners of the mouth is corrected to make the expression movement more natural and contextually appropriate. For the lip animation sequence, the pronunciation closure state of "weather" is predicted by optimizing the lip-teeth coordination to seamlessly connect with the facial expression. The final generated animation sequence can not only express a happy emotion but also accurately match the lip shape for each phoneme pronunciation.

[0086] In step S500, the facial animation sequence and the lip animation sequence are subjected to time alignment processing to obtain complete animation action data, and the animation action data and the audio information are synchronously rendered to generate a corresponding digital human video. Specifically, the steps include: The facial animation sequence and the lip animation sequence are subjected to time axis resampling, and a uniform frame rate and a reference starting frame are established based on the resampling results. The facial animation sequence and the lip animation sequence are subjected to time alignment processing using the uniform frame rate and the reference starting frame to extract the expression action vector and the lip action vector of each frame.

[0087] The time axis of the facial animation sequence and the lip animation sequence is resampled based on a dynamic time warping (DTW) algorithm. First, frame rate standardization is performed on the two animation sequences through discrete sampling to generate uniform time sampling points. Then, based on the reference starting frame, the timestamps in the animation sequences are dynamically aligned through the DTW algorithm to ensure that each frame point of the two sequences has a clear time correspondence. Finally, for each frame of the sequence, the expression action vector and the lip action vector of the corresponding frame are extracted to complete the alignment processing.

[0088] For example, given a facial animation sequence and a lip-sync animation sequence that does not perfectly match in time (e.g., the lip-sync sequence can be shorter due to different syllable speeds), the system normalizes both sequences to 30 FPS by resampling, aligns the frame "smiling with the pronunciation of 'tian'" by DTW, extracts the smile expression motion vector and the lip motion vector of the open and closed lips, and ensures the consistency of the corresponding frames in time.

[0089] The expression motion vector and the lip motion vector are fused to model the weight, a fusion driving vector sequence is constructed, and the fusion driving vector sequence is input into a preset motion combination module in the stylized three-dimensional face model to generate complete animation motion data.

[0090] By fusion weight modeling, specifically, the expression motion vector and the lip motion vector are combined into a single fusion driving vector sequence with the help of a multi-modal fusion mechanism. An attention-based fusion model is used to model the weight of the expression and the lip motion, and to optimize their collaborative relationship in the time range. The attention mechanism can assign weights to the expression motion and the lip motion according to the temporal context of the current frame, ensuring the consistency and smoothness of the motion. The fused driving vector sequence is input into the motion combination module in the stylized three-dimensional face model, which defines complete animation motion data using previously generated facial muscle control points and lip geometry deformation.

[0091] For example, in the processing of the frame sequence "today is a good day", the attention mechanism assigns a higher weight to the lip motion in the "today" frame, because the pronunciation of opening the mouth is the dynamic focus of the frame; while for the "is" frame, the expression motion of smiling and the slight corner of the mouth may jointly obtain a higher weight. Finally, the motion combination module generates complete animation motion data that is continuous and smooth, showing a seamless transition from a happy smile to clear pronunciation.

[0092] The audio information is aligned and sliced, the audio frame segment corresponding to the frame is extracted, the animation motion data is frame-level synchronized with the audio frame segment, and the corresponding digital human video is generated.

[0093] By audio slicing and frame-level synchronous rendering, specifically, first, according to a uniform frame rate, the audio information is time-fragmented, and the timestamps of the audio sequence and the animation action data are aligned. Then, based on a time-domain signal processing method (such as short-time Fourier transform, STFT), the audio segment features of each frame are extracted, ensuring that each frame of audio is synchronized with the corresponding animation action. Next, the animation action data and the corresponding audio frame segment are simultaneously input into the rendering engine for frame-level synchronous rendering. In the rendering engine, based on a real-time rendering algorithm accelerated by GPU processing, combined with the adjusted time sequence relationship of consecutive frames of animation and audio, the final video frame sequence is generated frame by frame.

[0094] For example, if each phoneme in the "today is a good weather" audio occupies two frames, and the animation action data contains more complex expression and mouth shape changes (such as smiling, opening mouth, closing mouth actions), the system will render the first frame of audio feature "today" with "open mouth + expression change", and the second frame of audio "weather" matches the expression change of "close mouth + smile transition", completing the time synchronization of audio and video frame by frame, and finally generating a continuous and smooth digital human video.

[0095] Among them, the step of aligning and slicing the audio information, extracting the corresponding audio frame segment, and performing frame-level synchronous rendering of the animation action data and the audio frame segment to generate the corresponding digital human video specifically includes: The audio information is divided into audio frame segment sets by a preset frame rate at the same time interval, and audio feature extraction is performed on each audio frame segment in the audio frame segment set to obtain time-domain audio features and frequency-domain audio features.

[0096] Through audio slicing and feature extraction, specifically, based on a set frame rate (such as 30FPS), the audio signal is time-axis sliced to discretize the audio information at the same time interval, obtaining an audio frame segment set. Each audio frame segment corresponds to a time interval. For each frame segment, the short-time Fourier transform (STFT, Short-Time Fourier Transform) is used to extract time-domain audio features and frequency-domain audio features. The time-domain audio features include the change of audio amplitude over time, and the frequency-domain audio features include the frequency spectrum distribution of the audio signal. The time-domain features are used to represent the dynamic changes of the audio, and the frequency-domain features are used to describe the frequency components of the audio.

[0097] For example, given a speech signal "today is a good weather", the system will slice the audio into small segments according to the preset frame rate of 30FPS, and each frame corresponds to 33ms of audio data. For the first frame slice "today", the STFT calculates its time-domain amplitude change curve and frequency-domain spectrum distribution. The time-domain features reflect the intensity of the tone gradually rising in the syllable, and the frequency-domain features show that the main energy is concentrated in the mid-frequency band.

[0098] The audio feature vector is constructed according to the time-domain audio feature and the frequency-domain audio feature, the audio feature vector is time-aligned with the motion feature vector of the corresponding frame in the animation motion data, the consistency of the audio and the motion data on the time axis is ensured, and an aligned audio feature vector is obtained.

[0099] By constructing the audio feature vector and time alignment, specifically, the time-domain audio feature and the frequency-domain audio feature of each frame slice are combined into an audio feature vector using a multi-dimensional feature splicing method. The audio feature vector contains amplitude change information in the time dimension and energy distribution pattern information in the frequency domain. Then, the audio feature vector is time-axis aligned with the corresponding frame motion feature vector in the animation motion data through a dynamic time warping (DTW) algorithm. The DTW can automatically optimize the matching relationship of the audio and the motion data in the time dimension, ensuring that the changes of the audio features and the motion features (such as mouth shape and expression) are completely synchronized in time.

[0100] For example, for the audio segment of the syllable “tian”, the spectral feature shows that the phoneme pronunciation energy is concentrated in the medium-high frequency band, which is associated with a specific mouth shape (slightly open mouth and tongue tip lifting). The DTW algorithm aligns the audio feature vector and the animation motion feature according to the timestamp, so that the starting sound wave of the “tian” audio is accurately matched with the mouth opening action in the animation, generating an aligned audio feature vector.

[0101] The aligned audio feature vector and the motion feature vector are adjusted in the frame level, the intensity of the audio and the rhythm of the motion are combined, the synchronous rendering parameters are obtained, the synchronous rendering parameters are input into a preset rendering module for audio-video synchronous rendering, and a corresponding digital human video is generated.

[0102] Through weight adjustment and synchronous rendering, specifically, in the frame-level weight adjustment stage, the audio feature vector and the motion feature vector are weighted and distributed through a multi-modal fusion attention mechanism. The attention mechanism analyzes the intensity changes of the audio (such as volume size and pitch lifting) and the rhythm and dynamic amplitude of the motion, and gives each frame an appropriate weight, ensuring that the synchronous rendering result of the audio and the motion conforms to the natural visual and auditory rules. Subsequently, the synchronous rendering parameters are input into a preset rendering module, the rendering module is based on GPU acceleration and real-time light and shadow calculation, the audio features and the animation motion feature sequence are fused frame by frame, and a continuous digital human video frame sequence is generated.

[0103] For example, in the pronunciation segment of "weather", "tian" as a loud syllable will get a higher audio weight, and its mouth opening and expression in the action feature will become more prominent; while "qi" as a lighter syllable will be assigned a lower weight, and its mouth slightly closes and the expression amplitude decreases. These weights are fused in real time frame by frame through the rendering module, and the final generated digital human video naturally presents the strength and dynamics of the "tian" character, while seamlessly transitioning to the gentle action of "qi".

[0104] In the embodiment, by combining text-driven expression features and audio-driven expression features, multi-level semantic encoding and speech feature extraction techniques are used to achieve natural and smooth digital human video generation. First, the BERT model is used to perform semantic analysis on the text, and a multi-layer semantic encoding network is used to extract intra-sentence semantic features and emotional information, and a context modeling network is used to extract inter-sentence dependency relationship and emotional flow feature, thereby obtaining emotion-enhanced semantic vectors and emotional feature representations. Then, through the audio feature extraction technology based on LSTM and MFCC, the prosody features and pronunciation dynamic features are extracted, and the attention mechanism is used to fuse them to ensure the time sequence consistency of the audio features and the animation actions. Subsequently, a deep fusion technology is used to modulate the text and audio features cooperatively to generate facial expressions and mouth shape parameters. A stylized three-dimensional face model is generated through style control parameters, and facial animation and mouth shape animation sequences are driven by facial expressions and mouth shape parameters. Finally, through time alignment and weight adjustment optimization, audio and video synchronization is optimized to generate high-quality digital human videos. In the whole process, the technical means include dynamic time warping (DTW) for audio and animation alignment, multi-modal generative adversarial network (MM-GAN) for expression and mouth shape modulation, and GPU accelerated real-time rendering module to ensure the naturalness and consistency of the generated video. Through these comprehensive means, the embodiment effectively solves the problems of uncoordinated expression and audio in traditional digital human generation technology, and the inability to meet the stylization requirements, and provides a higher quality and more personalized digital human video generation method.

[0105] Embodiment 3: As shown in Figure 2 The present application also provides a digital human video generation device 10, which comprises an acquisition module 11, a fusion module 12, a modulation module 13, a driving module 14 and a generation module 15.

[0106] The acquisition module 11 is mainly used for acquiring text information, audio information and style control parameters, performing semantic analysis on the text information to obtain text-driven expression features.

[0107] The fusion module 12 is mainly used for performing speech feature extraction on the audio information to obtain prosody feature vectors and pronunciation dynamic vectors, fusing the prosody feature vectors and the pronunciation dynamic vectors to obtain audio-driven expression features.

[0108] The modulation module 13 is mainly used for modulating the text-driven expression features and the audio-driven expression features cooperatively to obtain the facial expression parameter sequence and the mouth shape parameter sequence.

[0109] The driving module 14 is mainly used for generating a stylized three-dimensional face model of the digital human based on the style control parameter, driving the stylized three-dimensional face model by the facial expression parameter sequence to obtain a facial animation sequence, and driving the stylized three-dimensional face model by the mouth shape parameter sequence to obtain a mouth shape animation sequence.

[0110] The generation module 15 is mainly used for performing time alignment processing on the facial animation sequence and the mouth shape animation sequence to obtain complete animation action data, and synchronously rendering the animation action data and the audio information to generate a corresponding digital human video.

[0111] In the embodiment, through the cooperative work of the acquisition module 11, the fusion module 12, the modulation module 13, the driving module 14 and the generation module 15, the generation of the high-quality digital human video is realized. The acquisition module 11 performs semantic analysis on the text information, extracts text-driven expression features by using an advanced BERT model, and captures inter-sentence semantic dependency and emotional flow features by combining context modeling technology. The fusion module 12 extracts prosody feature vectors and pronunciation dynamic vectors based on the audio information by combining short-time Fourier transform (STFT) and Mel frequency cepstrum coefficient (MFCC), and performs time dimension fusion on the two by using an attention mechanism to generate accurate audio-driven expression features. The modulation module 13 deeply combines the text-driven expression features and the audio-driven expression features by cooperative modulation to generate facial expression parameter sequences and mouth shape parameter sequences containing emotions and speech coordination. The driving module 14 generates a stylized three-dimensional face model from a digital human parameter template library based on the style control parameter, and drives the muscle control module and the lip deformation module in the model by the expression parameters and the mouth shape parameters to realize the generation of facial animation and mouth shape animation. Finally, the generation module 15 performs time axis alignment processing on the facial animation sequence and the mouth shape animation sequence by dynamic time warping (DTW), and synchronously renders the audio information to generate a high-quality digital human video containing natural speech and expression video. The embodiment realizes high integration of text analysis, audio fusion, expression modulation, stylized driving and rendering generation process through modular design, significantly improves the expressiveness and personalized support of the digital human, and has efficient processing capability and stable generation quality.

[0112] It should be noted that, for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing embodiment 1, which will not be described herein again.

[0113] Embodiment 4: As Figure 3 shown, the present application also provides a computer device 20 comprising a memory 21 and a processor 22, the memory 21 storing a computer program executable on the processor 22, and the processor 22 implementing the digital human video generation method of embodiment 1 when executing the computer program.

[0114] In this embodiment, the complete function of the digital human video generation method is realized by the memory 21 and the processor 22 in the computer device 20. The computer program stored in the memory 21 includes modules such as acquisition, processing, fusion, driving and rendering, and the processor 22 performs logical operations and data processing of these modules. By running the computer program, the processor 22 first performs semantic analysis on the text information, extracts text-driven expression features using deep learning models such as BERT and context modeling network. Then, the processor 22 performs scientific frame processing on the audio information, extracts time domain and frequency domain features using short-time Fourier transform (STFT), and generates prosody feature vector and pronunciation dynamic vector through Mel frequency cepstrum coefficient (MFCC) algorithm, and finally completes audio feature fusion using attention mechanism to generate audio-driven expression features. Subsequently, the processor 22 cooperates with the modulation of text-driven expression features and audio-driven expression features to generate facial expression parameter sequence and mouth shape parameter sequence, and generates stylized three-dimensional face model based on style control parameters. With the help of facial expression parameters and mouth shape parameters, the processor 22 realizes animation driving through muscle modules and lip deformation modules in the driving model. Finally, the processor 22 performs time alignment on the facial animation and mouth shape animation through dynamic time warping (DTW) algorithm, and completes synchronous rendering combined with audio information to generate high-quality digital human video. Through the high-performance processing capability of the computer device 20, this embodiment integrates the collaborative work of multiple modules in the hardware and software system, and can efficiently complete the multi-modal fusion of text, speech and animation to generate natural and realistic digital human video, providing stable and efficient technical support for the application of film and television, virtual interaction and other fields.

[0115] The structures, proportions, sizes, etc. shown in the drawings of the present specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and do not have technical substantial significance, and any modification of structure, change of proportion relationship or adjustment of size, without affecting the effect and purpose that the present application can produce, should still fall within the scope of the technical content disclosed by the present application.

[0116] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a digital human video, characterized in that: include: Acquiring text information, audio information, and style control parameters, performing semantic analysis on the text information, and obtaining text-driven expression features; Performing speech feature extraction on the audio information to obtain a prosodic feature vector and a pronunciation dynamic vector, and fusing the prosodic feature vector with the pronunciation dynamic vector to obtain an audio-driven expression feature; Co-modulating the text-driven expression feature and the audio-driven expression feature to obtain a facial expression parameter sequence and a lip shape parameter sequence; generating a stylized 3D face model of a digital human based on the style control parameters, performing expression-driven operation on the stylized 3D face model using the facial expression parameter sequence to obtain a facial animation sequence, and performing lip-shaped operation on the stylized 3D face model using the lip-shaped parameter sequence to obtain a lip-shaped animation sequence; The facial animation sequence and the lip animation sequence are time-aligned to obtain complete animation action data, and the animation action data and the audio information are synchronously rendered to generate a corresponding digital human video.

2. The method for generating a digital human video according to claim 1, wherein: The step of performing semantic analysis on the text information to obtain text-driven expression features includes: Sentence processing is performed on the text information to obtain a sentence unit set and inter-sentence dependency relations, intra-sentence semantic features are extracted based on the sentence unit set, and a global semantic context vector and inter-sentence sentiment flow features are extracted based on the inter-sentence dependency relations; The intra-sentence semantic features are fused with the global semantic context vector to obtain a multi-level semantic feature vector, and the inter-sentence sentiment flow features are cross-attention mapped with the multi-level semantic feature vector to obtain a sentiment enhancement semantic vector and a sentiment feature representation; The emotion-enhanced semantic vector is style-adjusted using the emotion feature representation to obtain a semantic feature vector and an emotion feature vector consistent with the text intention, and the semantic feature vector and the emotion feature vector are fused to obtain a text-driven expression feature.

3. The method for generating a digital human video according to claim 1, wherein: The step of extracting speech features from the audio information to obtain a prosodic feature vector and a pronunciation dynamic vector, and fusing the prosodic feature vector with the pronunciation dynamic vector to obtain an audio-driven expression feature comprises: Performing frame processing on the audio information, extracting time domain envelope features and frequency domain spectrogram features of each frame, extracting a short-term rhythmic pattern based on the time domain envelope features, and fusing the short-term rhythmic pattern with rhythm information of adjacent frames to obtain a time-continuous rhythmic feature vector; Extracting frequency domain pronunciation features using the frequency domain spectrogram features; performing principal component decoupling processing on the frequency domain pronunciation features to extract a pronunciation state sequence and pronunciation boundary change information; encoding and mapping the pronunciation state sequence and the pronunciation boundary change information to obtain a pronunciation dynamic vector; The time-continuous prosodic feature vector is fused with the pronunciation dynamic vector in a time dimension to obtain an audio-driven expression feature.

4. The method for generating a digital human video according to claim 1, wherein: The step of collaboratively modulating the text-driven expression feature and the audio-driven expression feature to obtain a facial expression parameter sequence and a lip shape parameter sequence comprises: Input the text-driven expression features into an expression channel encoder to extract multi-level facial expression driving vectors, input the audio-driven expression features into a lip channel encoder to extract multi-band pronunciation mapping vectors; Performing time alignment processing on the multi-level facial expression driving vector and the multi-band pronunciation mapping vector to obtain expression feature pairs on the same time axis, constructing a fused expression state map based on the expression feature pairs, and performing feature decoding on the fused expression state map to obtain facial expression parameters and lip shape parameters; The facial expression parameters and the lip shape parameters are collaboratively modulated to eliminate expression conflicts and rhythm dislocations, thereby obtaining a facial expression parameter sequence and a lip shape parameter sequence.

5. The method for generating a digital human video according to claim 1, wherein: The steps of generating a stylized 3D face model of a digital human based on the style control parameters, performing expression-driven operation on the stylized 3D face model using the facial expression parameter sequence to obtain a facial animation sequence, and performing lip-shaped operation on the stylized 3D face model using the lip-shaped parameter sequence to obtain a lip-shaped animation sequence include: Based on the style control parameters, a corresponding basic 3D face model and style mapping weights are loaded from a preset digital human parameter template library, and the style mapping weights are used to perform style deformation and material mapping on the basic 3D face model to obtain a stylized 3D face model; Inputting the facial expression parameter sequence into a muscle control module preset in the stylized three-dimensional face model to drive changes in facial expression control points in the stylized three-dimensional face model to obtain an expression driving result; Inputting the lip shape parameter sequence into a lip deformation module preset in the stylized 3D face model to drive the geometric deformation of the lip and teeth area in the stylized 3D face model to obtain a lip shape driving result; The expression driving result is subjected to frame-by-frame posture fusion processing to generate a facial animation sequence, and the lip-type driving result is subjected to frame-by-frame posture fusion processing to generate a lip-type animation sequence.

6. The method for generating a digital human video according to claim 1, wherein: The steps of performing time alignment processing on the facial animation sequence and the lip animation sequence to obtain complete animation action data, and synchronously rendering the animation action data and the audio information to generate a corresponding digital human video include: resampling the facial animation sequence and the lip-sync animation sequence on a time axis, establishing a unified frame rate and a reference start frame based on the sampling results, performing time alignment processing on the facial animation sequence and the lip-sync animation sequence using the unified frame rate and the reference start frame, and extracting an expression action vector and a lip-sync action vector for each frame; Performing fusion weight modeling on the facial expression motion vector and the lip motion vector to construct a fusion drive vector sequence, and inputting the fusion drive vector sequence into a preset action combination module in the stylized three-dimensional face model to generate complete animation motion data; The audio information is aligned and sliced, audio frame segments of corresponding frames are extracted, the animation action data and the audio frame segments are synchronously rendered at the frame level, and a corresponding digital human video is generated.

7. The method for generating a digital human video according to claim 6, wherein: The steps of aligning and slicing the audio information, extracting audio frame segments of corresponding frames, performing frame-level synchronous rendering on the animation action data and the audio frame segments, and generating corresponding digital human videos include: The audio information is segmented into equal time intervals at a preset frame rate to obtain a set of audio frame segments, and audio features are extracted from each audio frame segment in the set of audio frame segments to obtain time-domain audio features and frequency-domain audio features; constructing an audio feature vector based on the time-domain audio features and the frequency-domain audio features, and performing time-series alignment on the audio feature vector and the motion feature vector of the corresponding frame in the animation motion data to obtain an aligned audio feature vector; Frame-level weight adjustment is performed on the aligned audio feature vector and the action feature vector to obtain synchronous rendering parameters, and the synchronous rendering parameters are input into a preset rendering module to perform audio and video synchronous rendering to generate a corresponding digital human video.

8. A digital human video generation device, characterized in that: include: An acquisition module is used to acquire text information, audio information, and style control parameters, perform semantic analysis on the text information, and obtain text-driven expression features; A fusion module is used to extract speech features from the audio information to obtain a prosodic feature vector and a pronunciation dynamic vector, and fuse the prosodic feature vector with the pronunciation dynamic vector to obtain an audio-driven expression feature; A modulation module, configured to collaboratively modulate the text-driven expression feature and the audio-driven expression feature to obtain a facial expression parameter sequence and a lip shape parameter sequence; a driving module for generating a stylized 3D face model of a digital human based on the style control parameters, performing expression-driven operation on the stylized 3D face model using the facial expression parameter sequence to obtain a facial animation sequence, and performing lip-shaped operation on the stylized 3D face model using the lip-shaped parameter sequence to obtain a lip-shaped animation sequence; The generation module is used to perform time alignment processing on the facial animation sequence and the lip animation sequence to obtain complete animation action data, synchronously render the animation action data and the audio information, and generate corresponding digital human video.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the method for generating a digital human video according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Public opinion emotion heat entropy calculation method based on bidirectional LSTM

    CN110162626A

  • Method and system for realizing three-dimensional facial expression and mouth shape by synchronously driving text

    CN114882154A

  • Animation fusion method and device

    CN114898019A

  • Animation data generation method and device, terminal equipment and storage medium

    CN117830480A

  • Multi-mode integrated digital human generation method, device, equipment and medium

    CN118644594A

Cited By

  • Digital human generation method based on sound driving

    CN121304866A

  • Digital human facial expression generation method and system based on generative adversarial network

    CN121353482A