Digital human video generation method and device, equipment, medium and program product
By performing emotion and rhythm analysis on the text to be output by the digital human, especially the analysis of tone and parallelism, voice and visual data that match the emotion and rhythm of the output text are generated, which solves the problem of unnatural expression of the digital human and improves the user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-10
AI Technical Summary
Existing digital humans lack accuracy in emotional expression, resulting in a reduced interactive experience. In particular, when dealing with flexible and ever-changing real-world communication scenarios, facial expressions and body movements appear rigid and unnatural.
By performing emotion and prosody analysis on the text to be output by the digital human, including tonal and parallelism analysis, voice and visual data matching emotion and prosody are generated. The pre-trained large model is then fine-tuned to construct an emotion and prosody analysis model, generating richer and more culturally rich digital human videos.
It improves the generation quality of digital human videos, ensuring that voice and actions match the emotions and rhythm of the text, and enhancing the user's interactive experience.
Smart Images

Figure CN121842472A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital human technology, and in particular to a method, apparatus, device, medium, and program product for generating digital human videos. Background Technology
[0002] With the rapid development of artificial intelligence technology, digital humans (also known as virtual humans, virtual avatars, etc.) are increasingly widely used in education, customer service, entertainment, and live streaming. Especially in international Chinese language education, digital humans are used to construct highly realistic immersive language interaction environments, breaking away from the traditional learning approach and providing learners with a scientific, engaging, and efficient immersive language interaction environment. This helps learners practice speaking and engage in cross-cultural communication, ultimately improving their fluency, listening comprehension, and confidence in cross-cultural communication.
[0003] To enable learners to engage in immersive, real-time voice or text conversations with different digital human characters, these characters need to possess diverse emotions, such as an enthusiastic convenience store clerk, an impatient taxi driver, a meticulous university professor, or a lively classmate. Therefore, digital humans need to be equipped with different emotional expression capabilities.
[0004] Currently, the emotional responses of digital humans largely rely on pre-set scripts and keywords in the output text. However, emotional interactions triggered by pre-set scripts and keywords are inaccurate. Their patterned and predictable nature makes them ill-suited to the flexible and ever-changing real-world communication scenarios. This can lead to digital humans' facial expressions, body language, and tone of voice often appearing rigid and unnatural, such as a stiff "fake enthusiasm" or a noticeable "mechanical" feel. Consequently, digital humans currently lack accurate emotional expression, resulting in a diminished user experience. Summary of the Invention
[0005] This invention provides a method, apparatus, device, medium, and program product for generating digital human videos, in order to solve the deficiency of digital humans in the prior art in lacking accurate emotional expression, and to realize the generation of digital human videos with fuller emotional expression and richer cultural connotation.
[0006] This invention provides a method for generating digital human videos, comprising: Obtain the output text to be output by the digital human; The output text is subjected to sentiment and prosody analysis to generate sentiment and prosody analysis data; the sentiment and prosody analysis includes the analysis of the tone and / or parallelism of the output text. Based on the emotion and prosody analysis data, voice data and visual data that match the emotion and prosody of the output text are generated; the visual data includes facial motion data and / or body motion data. Based on the voice data and the visual data, a digital human video is generated that matches the emotion and rhythm of the output text.
[0007] According to a method for generating digital human videos provided by the present invention, the step of performing sentiment and prosody analysis on the output text to generate sentiment and prosody analysis data includes... Based on the output text and the preset prompt word template, an input prompt word is generated; The input prompt words are input into the sentiment and prosody analysis model to obtain the sentiment and prosody analysis data output by the sentiment and prosody analysis model; The sentiment and prosody analysis model is obtained by fine-tuning a pre-trained large model. The sentiment and prosody analysis model is used to perform sentiment and prosody analysis on the output text. The sentiment and prosody analysis also includes the analysis of the number of vowels in the output text.
[0008] According to a method for generating digital human videos provided by the present invention, the emotion and prosody analysis further includes the analysis of the emotion category of the output text and the dimensional scoring of the emotion category, wherein the dimensional scoring includes at least one of valence score, arousal score and dominance score.
[0009] According to a method for generating digital human videos provided by the present invention, the visual data is generated based on the following manner: The emotion and prosody analysis data is input into the action generation model to obtain the visual data output by the action generation model; the emotion and prosody analysis data includes the digital human's output text and the emotion tag of each word in the output text; The action generation model is built based on a large model; the action generation model is used to synchronize facial expressions and lip movements based on the text to be output and each of the emotion tags, so as to ensure that the facial expressions and lip movements of the same frame represented by the facial action data are synchronized.
[0010] According to a method for generating digital human videos provided by the present invention, the motion generation model is generated based on the following method: Based on the sample sentiment and prosody analysis data, and the visual data labels corresponding to the sample sentiment and prosody analysis data, the large model is fine-tuned to obtain the action generation model. The visual data tags include action data that represents cultural symbols.
[0011] According to a method for generating digital human videos provided by the present invention, the voice data is generated based on the following method: The emotion and prosody analysis data are input into the speech synthesis model to obtain the speech data output by the speech synthesis model; The speech synthesis model is built based on a large model and is used to control at least one parameter among pitch, volume, speech rate and timbre based on the emotion and prosody analysis data.
[0012] According to a method for generating digital human videos provided by the present invention, the emotion and prosody analysis data includes the output text of the digital human and the emotion tag of each word in the output text; The step of generating speech and visual data that match the emotion and prosody of the output text based on the emotion and prosody analysis data includes: When the text to be output includes onomatopoeia, the timbre text corresponding to the onomatopoeia is determined from a preset timbre text set based on the sentiment tag of the onomatopoeia, and the onomatopoeia is replaced with text based on the timbre text; the timbre text is used to indicate the timbre of the onomatopoeia output. Based on the sentiment and prosody analysis data after text replacement, voice data and visual data that match the sentiment and prosody of the output text are generated.
[0013] According to a method for generating a digital human video provided by the present invention, the step of generating a digital human video that matches the emotion and rhythm of the output text based on the speech data and the visual data includes: Based on the emotion and rhythm analysis data, special effects data that match the emotion and rhythm of the output text are determined from a preset set of special effects. Based on the voice data, the visual data, and the special effects data, a digital human video is generated that matches the emotion and rhythm of the output text.
[0014] According to a method for generating a digital human video provided by the present invention, the step of obtaining the output text to be output by the digital human includes: Acquire user-input voice data and user visual data; the user visual data includes user facial movement data and / or user body movement data. Based on the user's voice data and visual data, determine the user's emotional state data; The emotional state data and dialogue context are input into the dialogue generation model to obtain the output text output by the dialogue generation model; the dialogue context is the dialogue between the user and the digital human, and the dialogue generation model is built based on a large model.
[0015] According to a method for generating digital human videos provided by the present invention, determining the user's emotional state data based on the user's voice data and user visual data includes: Emotion and prosodic analysis is performed on the acoustic features of the user's voice data to generate first emotional state data; Sentiment analysis is performed on the visual features of the user visual data to generate second emotional state data; the user visual data also includes background environment data and user clothing data, and the body movement data includes body tilt data. Based on the first emotional state data and the second emotional state data, the user's emotional state data is determined.
[0016] According to a method for generating digital human videos provided by the present invention, determining the user's emotional state data based on the first emotional state data and the second emotional state data includes: Emotion and prosody analysis is performed on the input text of the user's voice data to generate third emotional state data; Based on the first emotional state data, the second emotional state data, and the third emotional state data, the user's emotional state data is determined.
[0017] According to a method for generating digital human videos provided by the present invention, the step of inputting the emotional state data and dialogue context into a dialogue generation model to obtain the output text output by the dialogue generation model includes: The visual features of the user's visual data are analyzed to generate the user's age data. The emotional state data, dialogue context, and age data are input into the dialogue generation model to obtain the output text output by the dialogue generation model.
[0018] The present invention also provides an apparatus for generating digital human videos, comprising: The text acquisition module is used to acquire the output text to be output by the digital human. The sentiment analysis module is used to perform sentiment and prosody analysis on the output text and generate sentiment and prosody analysis data; the sentiment and prosody analysis includes the analysis of the tone and / or parallelism of the output text; The data generation module is used to generate voice data and visual data that match the emotion and rhythm of the output text based on the emotion and rhythm analysis data; the visual data includes facial motion data and / or body motion data. The video generation module is used to generate a digital human video that matches the emotion and rhythm of the output text based on the voice data and the visual data.
[0019] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the digital human video generation method as described above.
[0020] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating digital human videos as described above.
[0021] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the digital human video generation method as described above.
[0022] The present invention provides a method, apparatus, device, medium, and program product for generating digital human videos. It acquires the output text to be output by the digital human, performs emotion and rhythm analysis on the output text, and generates emotion and rhythm analysis data. This emotion and rhythm analysis includes analysis of the tonal patterns and / or parallelism of the output text, thus going beyond the analysis of emotional words to conduct a deeper analysis of the text, particularly analyzing the rhythmic features unique to language such as tonal patterns and parallelism. This aims to generate digital human videos with richer emotional expression and cultural connotations, thereby improving the generation effect of digital human videos and ultimately enhancing the user's interactive experience. Furthermore, the emotion and rhythm analysis of the output text generates voice and visual data that match the emotion and rhythm of the output text. The visual data includes facial movement data and / or body movement data, thus avoiding mismatches between the digital human's output voice and movements and the output text, as well as avoiding mismatches and asynchronies between voice and movements. Based on the voice and visual data, it generates an accurately synchronized digital human video that matches the emotion and rhythm of the output text, ultimately improving the user experience. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0024] Figure 1 This is one of the flowcharts illustrating the digital human video generation method provided by the present invention.
[0025] Figure 2 This is the second flowchart illustrating the digital human video generation method provided by the present invention.
[0026] Figure 3This is a schematic diagram of the structure of the digital human video generation device provided by the present invention.
[0027] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0029] Currently, emotional interactions that rely on pre-set scripts and keywords will result in digital humans lacking accurate emotional expression, which in turn will reduce the user interaction experience.
[0030] Based on the problems existing in the aforementioned digital human video generation schemes, this study aimed to generate corresponding digital human videos based on the output text to be produced. Specifically, Natural Language Processing (NLP) technology was used to perform sentiment analysis on the output text, identifying basic emotional categories such as joy, anger, sorrow, and happiness. Subsequently, based on the analyzed emotional categories, a speech synthesis module was driven to generate speech with corresponding emotional tone, and a virtual avatar model was driven to generate matching facial expressions and body movements. Finally, the speech and visual images were synthesized into a digital human video for output.
[0031] However, in-depth research has revealed significant technical flaws in the practical application of this approach. First, its emotional expression often remains superficial, merely focusing on the recognition of emotional words and the matching of basic emotional categories. This results in the digital human's voice output resembling a mechanical reading of text rather than an engaging emotional expression. Furthermore, its facial expressions and body movements often appear rigid and unnatural, failing to meet the requirements for authenticity and immersion in high-quality language teaching.
[0032] To address the shortcomings of the aforementioned approach, further research revealed that its text analysis is rather simplistic, neglecting the deep emotional and cultural connotations inherent in the rhythmic structure of language. For instance, in the rich and profound Chinese linguistic context, the tonal patterns and parallel structures of a text are not only crucial for the phonetic beauty of the language but also reveal the author's subtle emotional fluctuations, emphasis, and cultural depth. Failure to effectively analyze and utilize these deep-seated linguistic rhythmic features will result in generated speech and actions that fail to reflect the unique rhythm and cadence of Chinese. This is especially true when dealing with poetry, idioms, or literary and beautiful phrases, where the expressive effect falls far short of natural and fluent expression, failing to accurately convey the "warmth" and "expression" of the language to learners.
[0033] Therefore, this invention ultimately proposes a method for generating digital human videos. This method acquires the output text to be produced by the digital human, performs emotion and rhythm analysis on the output text, and generates emotion and rhythm analysis data. The emotion and rhythm analysis includes analysis of the tone and / or parallelism of the output text, thus going beyond the analysis of emotional words to conduct a deeper analysis of the text, especially analyzing the rhythmic features unique to language such as tone and parallelism, in order to generate digital human videos with richer emotional expression and richer cultural connotations, thereby improving the generation effect of digital human videos and ultimately enhancing the user interaction experience. Furthermore, the emotion and rhythm analysis of the output text to be produced by the digital human generates voice data and visual data that match the emotion and rhythm of the output text. The visual data includes facial movement data and / or body movement data, thereby avoiding mismatches between the voice and movements of the digital human outputting the text, as well as avoiding mismatches and asynchrony between voice and movements. Based on the voice data and visual data, an accurate and synchronized digital human video that matches the emotion and rhythm of the output text is generated, ultimately improving the user experience.
[0034] The method for generating digital human videos provided by the present invention will be described below through various embodiments. Figures 1-2 This invention describes a method for generating digital human videos. This method can be applied to electronic devices, such as servers, user devices, or cloud computing platforms.
[0035] Figure 1 This is one of the flowcharts illustrating the digital human video generation method provided by the present invention, such as... Figure 1 As shown, the method for generating the digital human video includes the following steps 110, 120 and 130.
[0036] Step 110: Obtain the output text to be output by the digital human.
[0037] The application scenario of this invention is a dialogue scenario (interactive scenario). For ease of explanation, this invention will be illustrated using the "international Chinese education and teaching" scenario as an example.
[0038] Here, digital human is also called virtual human or virtual avatar, and the digital human video generated later will showcase the digital human.
[0039] Here, output text refers to the text content that the digital human will express through voice and actions. The source of this output text can be diverse; for example, it could be response text generated by the dialogue system based on the current dialogue context, pre-set dialogue text in a teaching script, or text fragments extracted from an external knowledge base. For instance, in an international Chinese education scenario, the output text could be "Spring slumber unaware of dawn, everywhere the birdsong is heard."
[0040] In one embodiment, the dialogue context is input into the dialogue generation model to obtain the output text from the dialogue generation model. The dialogue context is a dialogue between the user and the digital human, and the dialogue generation model is built based on a large model.
[0041] Step 120: Perform sentiment and prosody analysis on the output text to generate sentiment and prosody analysis data.
[0042] The emotion and rhythm analysis includes the analysis of the tone and / or parallelism of the output text.
[0043] It should be noted that this sentiment and prosody analysis also includes literal sentiment analysis of the output text, i.e., analysis of sentiment words. For example, it analyzes the sentiment words contained in the output text (such as "happy" and "sad"), classifies the sentiment category of the entire output text (such as positive, negative, and neutral), and scores the dimensions of sentiment, such as valence score, arousal score, and dominance score.
[0044] Here, sentiment and prosodic analysis refers to conducting in-depth sentiment analysis on the output text to understand its literal meaning, emotional tone, and inherent prosodic structure.
[0045] In one specific embodiment, input prompts are generated based on the output text and a preset prompt template. These prompts are then fed into a sentiment and prosody analysis model to obtain sentiment and prosody analysis data output by the model. The sentiment and prosody analysis model is obtained by fine-tuning a pre-trained large model and is used to perform sentiment and prosody analysis on the output text. Further, the sentiment and prosody analysis model first analyzes the role corresponding to the output text (whether it is a question, statement, or exclamation), then analyzes the role's identity and personality, and finally analyzes the role's emotional tone.
[0046] Here, the level and oblique tone analysis refers to, for the output text, calling a preset Chinese character level and oblique tone database or passing through a pre-trained large language model to judge the tone of each Chinese character in the text and mark whether it belongs to the level tone or the oblique tone (the oblique tone usually includes the rising tone, the falling tone, and the entering tone). For example, for the output text "Spring slumber goes unaware of dawn", its level and oblique tone analysis result may be "[level][level][oblique][oblique][oblique]". This analysis helps to understand the phonetic rhythm of the sentence and guides the subtle changes in pitch and duration during subsequent speech synthesis to produce a sense of rhythm that conforms to the chanting habits of classical poetry, thereby improving the generation effect of the digital human video.
[0047] Here, the antithesis analysis refers to analyzing whether there is an antithesis structure in the output text, that is, two sentences with equal number of characters, opposite词性, the same structure, and related meanings. For example, for the output text "Two golden orioles sing amid the willows green; A flock of white egrets fly into the blue sky", it will be analyzed that this is an antithesis sentence. After analyzing the antithesis relationship, when generating visual data, similar or symmetrical body movements and camera angles can be matched for these two sentences to strengthen the structural beauty of the language visually, and when generating voice data, similar or symmetrical voice data can be matched for these two sentences to strengthen the structural beauty of the language aurally.
[0048] Furthermore, the emotion and rhythm analysis also includes the analysis of the number of vowels in the output text. Specifically, analyze the length of the output text and the number of its vowels, so as to consider that if there are more vowels, the voice will be louder, and thus improve the generation effect of the digital human video.
[0049] In one embodiment, the emotion and rhythm analysis also includes the analysis of the emotion category of the output text, as well as the dimension scoring for the emotion category, and the dimension scoring includes at least one of valence scoring, arousal scoring, and dominance scoring.
[0050] The emotion and rhythm analysis data may include but is not limited to at least one of the following: the text to be output by the digital human, the emotion label of each word or each character in the text to be output, the emotion label of each phrase in the text to be output, stress position suggestions, pause marks, sentence-breaking suggestions, level and oblique tone annotation sequences, and antithesis structure relationships, etc. The text to be output may be the output text or the text optimized from the output text.
[0051] In one embodiment, the emotion and rhythm analysis data is an annotation data (such as an annotation file).
[0052] In one embodiment, the emotion and rhythm analysis data is a structured data file. This data file details the various analysis attributes of the output text and is used to guide the generation of subsequent voice data and visual data.
[0053] Step 130: Based on the emotion and prosody analysis data, generate speech data and visual data that match the emotion and prosody of the output text.
[0054] In one specific embodiment, emotion and prosody analysis data are input into a speech synthesis model to obtain speech data output by the speech synthesis model. The speech synthesis model is built upon a large model. For example, based on emotion tags (such as joy) and arousal scores (such as high) in the emotion and prosody analysis data, the speech synthesis model raises the fundamental frequency (pitch) of the speech as a whole and increases its variation amplitude; according to the tonal marking sequence, it fine-tunes the duration and pitch curve of each character, making oblique tone characters short and forceful, and level tone characters long and smooth. The speech data generated in this way not only matches the emotion in timbre, but its intonation also highly matches the prosodic structure of the text.
[0055] The visual data includes facial motion data and / or body motion data. The facial motion data includes facial expression data.
[0056] In one specific embodiment, emotion and prosody analysis data are input into the motion generation model to obtain visual data output by the motion generation model. The motion generation model is built upon a large model.
[0057] For example, the motion generation model generates corresponding facial motion data based on emotion tags (such as joy) in emotion and rhythm analysis data. This facial motion data is used to drive the digital human to raise the corners of its mouth and bend its eyes.
[0058] For example, motion generation models can generate corresponding body postures and gestures based on the output text and prosodic structure in sentiment and prosodic analysis data. For instance, when the analysis reveals that the text has a parallel structure, the digital human can be driven to make directional gestures with its left and right hands in sequence to visually correspond to the duality of the text.
[0059] It should be noted that both speech data and visual data include multiple frames of data, and the same frames are matched; that is, speech data is speech sequence data, and visual data is visual sequence data.
[0060] Step 140: Based on the voice data and the visual data, generate a digital human video that matches the emotion and rhythm of the output text.
[0061] In one specific embodiment, voice data and visual data are synthesized and rendered to be rendered onto a digital human model to obtain a digital human video.
[0062] Furthermore, the speech and visual data are fused and synchronized, meaning strict timestamp alignment is performed to ensure that every phoneme in the speech data is precisely synchronized with the lip-sync animation of the digital human (included in facial motion data), and that the stress points in the speech occur synchronously with the keyframes of facial expressions (such as raising eyebrows or nodding). Finally, the rendering engine outputs a complete digital human video containing sound (tone and timbre), lip movements, facial expressions, and movements. This multimodal fusion and synchronization generates a digital human video with synchronized audio and visuals and consistent emotions.
[0063] Furthermore, after outputting the digital human video, users can provide real-time feedback during interaction (such as "Your response is too lukewarm"). The digital human can then dynamically adjust its response strategy based on this feedback, achieving personalized emotional interaction calibration. For example, based on user interaction feedback, models such as emotion and prosody analysis models, action generation models, and speech synthesis models can be optimized.
[0064] It should be understood that by introducing the analysis of rhythmic features such as tone and parallelism, the problem of mechanical expression and lack of cultural connotation in digital human is solved. This enables the generation of digital human videos with richer emotional expression and richer cultural connotation, significantly improving the generation effect and realism of digital human videos, thereby bringing users (especially learners in the context of international Chinese education) a better and more immersive interactive experience.
[0065] The digital human video generation method provided in this invention obtains the output text to be output by the digital human, performs emotion and rhythm analysis on the output text, and generates emotion and rhythm analysis data. This emotion and rhythm analysis includes analysis of the tonal patterns and / or parallelism of the output text, thus going beyond the analysis of emotional words to conduct a deeper analysis of the text, particularly analyzing the rhythmic features unique to language such as tonal patterns and parallelism. This aims to generate digital human videos with richer emotional expression and cultural connotations, thereby improving the generation effect of digital human videos and ultimately enhancing the user's interactive experience. Furthermore, the emotion and rhythm analysis of the output text to be output by the digital human generates voice and visual data that match the emotion and rhythm of the output text. The visual data includes facial movement data and / or body movement data, thus avoiding mismatches between the digital human's output voice and movements and the output text, as well as avoiding mismatches and asynchronies between voice and movements. Based on the voice and visual data, an accurate and synchronized digital human video matching the emotion and rhythm of the output text is generated, ultimately improving the user experience.
[0066] Based on any of the above embodiments, in this method, step 120 includes: Based on the output text and the preset prompt word template, an input prompt word is generated; The input prompts are fed into the sentiment and prosody analysis model to obtain the sentiment and prosody analysis data output by the model.
[0067] Here, the preset prompt template can be a text string containing instructions and placeholders. For example, a preset prompt template is shown below: "You are an expert in Chinese linguistics and sentiment analysis. Please perform a comprehensive sentiment and prosodic analysis on the following text: '{text}'."
[0068] Your analysis should include: 1. Overall emotional category of the text (choose from happy, sad, angry, fearful, surprised, disgusted).
[0069] 2. The tonal sequence of the text.
[0070] 3. Does the text contain a parallel structure?
[0071] 4. The number of vowels in each sentence of the text.
[0072] Please output your analysis results in JSON format. When the system receives the output text, such as "Spring sleep is so sweet that one does not notice the dawn, but everywhere one hears birds singing", it will fill the {text} placeholder in the prompt word template to generate the final input prompt word. Then, this complete input prompt word will be input into the sentiment and prosody analysis model.
[0073] It should be noted that, in order for the large model (the sentiment and prosody analysis model) to accurately perform the required sentiment and prosody analysis tasks, the original output text is not directly input into the model; instead, a preset prompt word template is used to embed the output text into it, so as to form a structured and clearly instructive input prompt word, ensuring that it can perform analysis according to the analysis requirements.
[0074] For example, the analysis requirements are: first, analyze the role corresponding to the output text (whether it is a question, statement, or exclamation); second, the role's identity and personality; finally, the role's emotional tone; and third, analyze the length of the output text, the number of vowels (more vowels will make the sound louder), tone patterns, and parallelism, in order to analyze the emotional and cultural connotations of the output text.
[0075] Furthermore, considering that textual characters and their identities and personalities are usually relatively easy to classify and determine, but the determination of emotional color is more difficult; based on this, the analysis requirements also include: sorting out the emotional color of the text in advance and the emotional analysis dimensions. The emotional color of the text usually includes basic emotions and complex emotions. Basic emotions include happiness, sadness, anger, fear, surprise, and disgust, while complex emotions include embarrassment, jealousy, envy, guilt, pride, gratitude, expectation, and confusion.
[0076] Furthermore, the analysis requirements also include: coarse-grained classification, first determining whether it is positive / negative / neutral; second, fine-grained classification, further classifying it into specific categories such as happy, sad, and angry, and providing confidence levels; next is dimensional scoring, predicting the continuous values of the text in valence, arousal, and dominance, such as a pleasure level of 0.8 and an arousal level of 0.9 representing "very excited"; finally, in specific applications, especially when driving digital humans, it is also necessary to analyze more subtle linguistic features and social intentions, such as knowing whether it is a "gratified smile" (low arousal) or "ecstasy" (high arousal), or "sincere praise" or "insincere flattery".
[0077] It should be understood that by adopting cue word engineering, the behavior of large models can be controlled more precisely and stably, ensuring that they can be analyzed according to preset dimensions and formats, thereby improving the accuracy and reliability of sentiment and prosody analysis, and thus improving the generation effect of digital human videos.
[0078] The sentiment and prosody analysis model is obtained by fine-tuning a pre-trained large model. The sentiment and prosody analysis model is used to perform sentiment and prosody analysis on the output text. The sentiment and prosody analysis also includes the analysis of the number of vowels in the output text.
[0079] Here, the pre-trained large model refers to a general-purpose large language model that has been trained on massive and diverse text data, such as iFlytek's Spark large model. This large model already possesses powerful natural language understanding and generation capabilities.
[0080] Here, fine-tuning a pre-trained large model refers to further training and adjusting the model parameters using a relatively small, labeled dataset relevant to a specific task, based on a pre-trained large model. This allows the general-purpose large model to adapt to and become proficient in the specific task required, thereby improving the accuracy and reliability of sentiment and prosody analysis, and ultimately enhancing the generation of digital human videos.
[0081] In one embodiment, the dataset used for fine-tuning contains a large number of text samples, along with corresponding sentiment and prosodic analysis data annotated by human experts or high-precision tools. By fine-tuning on such a dataset, the general-purpose large model is transformed into a sentiment and prosodic analysis model focused on the sentiment and prosodic analysis task.
[0082] Furthermore, to make the language style of the sentiment and prosody analysis model more closely resemble contemporary spoken Chinese, the corpus used in the pre-training and / or fine-tuning phases of the large model can specifically include a large number of spoken language examples obtained from news websites (such as Xinhua News Agency) and new media platforms (such as Xiaohongshu, Weibo, and Douyin). This allows the interactive content generated by the model to be more timely, more everyday, and more human-like, thus making the model's understanding of modern Chinese more grounded in reality.
[0083] It should be understood that domain-specific fine-tuning enables general-purpose large models to perform specific sentiment and prosody analysis tasks professionally and efficiently, resulting in more accurate analysis results. This improves the accuracy and reliability of sentiment and prosody analysis, thereby enhancing the generation effect of digital human videos.
[0084] Here, vowels refer to phonemes in which airflow is unobstructed in the oral cavity during pronunciation. In Chinese Pinyin, they usually refer to a, o, e, i, u, and ü.
[0085] Vowel count analysis refers to the model's calculation of the total number of vowel phonemes contained in each sentence or phrase of the output text. Based on this, considering that in linguistics and acoustics, the number of vowels is usually related to the loudness and fullness of a sound, a text with more vowels tends to sound louder and more expansive, and its emotional expression may be more unrestrained; conversely, it may sound more restrained or abrupt. This adds a new and effective physical dimension to prosodic analysis. This feature can directly guide subsequent speech data generation, adjusting its energy and resonance peaks to match the loudness of the sound with the acoustic potential of the text, thus making the emotional expression more realistic and believable at the physical level. For example, sentences with more vowels can be synthesized to be louder and fuller, thus better expressing excitement or passion. This further enriches the dimensions of emotion and prosodic analysis, improving the generation effect and realism of the final digital human video.
[0086] The digital human video generation method provided in this invention generates input prompts based on output text and preset prompt word templates. This allows for more precise and stable control of the emotion and prosody analysis model, ensuring it can perform analysis according to preset dimensions and formats. This improves the accuracy and reliability of emotion and prosody analysis, thereby enhancing the generation effect of digital human videos and ultimately improving the user's interactive experience. The input prompts are then fed into the emotion and prosody analysis model to obtain the output emotion and prosody analysis data. The emotion and prosody analysis model is obtained by fine-tuning a pre-trained large model. Through domain-specific fine-tuning, the emotion and prosody analysis model can professionally and efficiently perform specific emotion and prosody analysis tasks, resulting in more accurate analysis results. This further improves the accuracy and reliability of emotion and prosody analysis, enhancing the generation effect of digital human videos and ultimately improving the user's interactive experience.
[0087] Based on any of the above embodiments, the sentiment and prosody analysis in this method further includes the analysis of the sentiment category of the output text and the dimensional scoring of the sentiment category, wherein the dimensional scoring includes at least one of valence score, arousal score and dominance score.
[0088] Here, the analysis of the sentiment category of the output text is performed by the aforementioned finely tuned sentiment and prosody analysis model, guided by cue word engineering.
[0089] In one embodiment, the analysis of the sentiment category of the output text is typically multi-layered, achieving a coarse-to-fine sentiment judgment. Specifically, first, the model performs a basic polarity judgment on the output text, classifying it as positive, negative, or neutral. Then, based on the coarse-grained classification, the model further categorizes the text into more specific sentiment categories. These categories can include basic emotions such as happiness, sadness, anger, fear, surprise, and disgust; or more complex compound emotions such as embarrassment, jealousy, envy, guilt, pride, gratitude, anticipation, and confusion. Furthermore, the model typically provides a confidence level during classification, indicating its degree of certainty about the classification result. For example, for the output text "Great! I finally passed the exam!", the model would first classify it coarsely as positive, then finely as happy, and might provide a confidence level of, for example, 98%.
[0090] Considering that analyzing emotion categories alone is insufficient, as it cannot distinguish between different intensities and states within the same emotion category—for example, both "a relieved smile" and "a joyful laugh" belong to the category of "happiness," but their expressions and underlying states differ greatly—a dimensional scoring system is used to classify emotion categories.
[0091] Considering that emotions are not discrete categories, but rather a point in one or more continuous dimensional spaces, in this embodiment of the invention, the dimensional rating includes at least one or more of valence rating, arousal rating, and dominance rating. These ratings are typically standardized continuous values, for example, within the interval [-1, 1] or [0, 1].
[0092] Here, the valence score represents the degree of pleasure or displeasure. For example, its value ranges from -1 to 1, with positive values representing positive emotions (such as happiness and satisfaction) and negative values representing negative emotions (such as sadness and anger). The absolute value of the score represents the intensity of pleasure or displeasure.
[0093] Here, the arousal score represents the degree of emotional excitement or calmness, that is, the level of physiological activation. For example, its value ranges from [0, 1], with higher values indicating high arousal (such as excitement, anger, fear) and lower values indicating low arousal (such as calmness, sadness, satisfaction).
[0094] Here, the dominance score represents an individual's sense of control or power in an emotional state. Its value ranges from [0, 1], with higher values indicating a feeling of power and control (such as pride or anger), and lower values indicating a feeling of weakness and passivity (such as fear or guilt).
[0095] For example, if the output text is "Great! I finally passed the exam!", the sentiment and prosody analysis model, after classifying it into happiness at a fine-grained level, will also assign it a dimensional score. Because the text expresses a very strong positive emotion, its score might be: Valence: 0.9 (indicating very pleasant), Arousal: 0.8 (indicating very excited), and Dominance: 0.7 (indicating confidence and a sense of control). However, for another output text belonging to the same happiness category, such as "Seeing that you are all well makes me feel relieved," its dimensional scores might be: Valence: 0.7 (indicating pleasant), Arousal: 0.3 (indicating calm and relaxed), and Dominance: 0.4 (indicating a mild sense of control). Based on this, a comparison reveals that although both texts express happiness, the difference in dimensional scores clearly reveals a significant distinction between them: the former is a very excited happiness, while the latter is a relieved and calm happiness.
[0096] It should be understood that by combining continuous values such as valence, arousal, and dominance, it is possible to distinguish more subtle linguistic features and social intentions that discrete emotion classification cannot differentiate, as well as to judge rhetorical devices (such as irony and hyperbole). This enables the system to understand whether the text expresses ecstasy or satisfaction, genuine praise or insincere flattery, thus providing a key basis for subsequently generating more accurate and nuanced speech and motion data, thereby improving the refinement of emotional expression and ultimately enhancing the generation effect of digital human videos.
[0097] It should be understood that dimensional scoring provides direct, quantifiable parameters for the generation of voice and visual data. For example, a high arousal score can be directly mapped to a faster speech rate, higher pitch, louder volume, and richer body language; a high valence score corresponds to an upturned mouth and relaxed eyebrows. This quantitative guidance makes the generation process of emotional expression more precise and controllable, significantly improving the realism and emotional impact of the final generated video, thus enhancing the refinement of emotional expression and ultimately improving the generation effect of digital human videos.
[0098] The digital human video generation method provided in this invention, through valence scoring, arousal scoring, and dominance scoring, can distinguish more subtle linguistic features and social intentions that discrete emotion classification cannot differentiate, thereby improving the precision of emotional expression, thus enhancing the generation effect of digital human videos and ultimately improving the user's interactive experience. Moreover, the quantitative guidance of dimensional scoring makes the generation process of emotional expression more precise and controllable, significantly improving the authenticity and appeal of the final generated video, that is, improving the precision of emotional expression, thereby enhancing the generation effect of digital human videos and ultimately improving the user's interactive experience.
[0099] Based on any of the above embodiments, in this method, the visual data is generated in the following manner: The emotion and rhythm analysis data are input into the action generation model to obtain the visual data output by the action generation model.
[0100] The sentiment and prosody analysis data includes the digital human's output text and the sentiment tag for each word in the output text.
[0101] The action generation model is built based on a large model; the action generation model is used to synchronize facial expressions and lip movements based on the text to be output and each of the emotion tags, so as to ensure that the facial expressions and lip movements of the same frame represented by the facial action data are synchronized.
[0102] Here, the action generation model is a model that transforms text-level analysis results into dynamic visual representations. It should be noted that this action generation model is built upon a large-scale model, thus leveraging the powerful sequence-to-sequence conversion capabilities and contextual understanding of the pre-trained large-scale model to improve the accuracy of visual data generation.
[0103] In one embodiment, similar to the aforementioned sentiment and prosody analysis model, the action generation model can also be obtained by fine-tuning a general large model on a specific dataset, whose training data contains a large number of text-action pairing samples. Specifically, the action generation model is obtained by fine-tuning the large model based on the sample sentiment and prosody analysis data and the corresponding visual data labels.
[0104] In one embodiment, visual data tags can be acquired as follows. Specifically, professional actors are hired and fitted with high-precision facial and full-body motion capture equipment in a professional motion capture environment. The actors are required to read a large amount of text covering various emotions (such as happiness, sadness, anger, etc.) and rhythms (such as calm, passionate, and abrupt). In this way, a large amount of natural, high-quality dynamic facial expression and movement data that is highly synchronized with the speech and text content can be collected. This is considered the gold standard for acquiring natural facial expression data.
[0105] Furthermore, in addition to regular emotional expressions, to make the actions of digital humans more culturally meaningful, actions with special meanings in specific cultural contexts will be specifically captured. For example, by analyzing film and television clips or publicly available video materials, gestures and movements that have evolved into cultural symbols will be captured and incorporated into the dataset.
[0106] In addition, existing or newly acquired visual data labels are manually annotated frame by frame or keyframe by frame. Annotation typically employs a facial motion coding system, breaking down complex expressions into basic Action Units (AUs). For example, "pulling up the corners of the mouth" is labeled as a sign of "smiling," while "drooping and furrowing eyebrows" is labeled as a sign of "anger" or "focus." In this way, a correspondence is established between concrete muscle movements and abstract emotional labels.
[0107] It should be noted that the goal of fine-tuning is to enable the large model to predict the corresponding motion sequence data based on the input text content (output text) and emotional rhythm (emotional labels). Specifically, the training data pairs collected and processed in the previous step, i.e., the sample emotion and prosody analysis data, are used as input, and their corresponding visual data labels are used as the expected output to fine-tune the motion generation model based on the large model.
[0108] Furthermore, during training, the model learns the mapping relationship between eye contact and emotion. Based on this, considering that eye contact is a key window for conveying emotion, the model specifically learns the correlation between eye movement, blinking frequency, and gaze direction with current emotions and thinking states (such as upward gaze when confused, and fixed gaze when focused), thereby improving the accuracy of visual data generation and thus improving the generation effect of digital human videos.
[0109] To achieve accurate visual data generation, the input data fed into the motion generation model must be detailed and structured. In this embodiment of the invention, the input sentiment and prosody analysis data includes at least: the digital human's output text (i.e., the complete text content) and the sentiment tags for each word in the output text (these are detailed annotations generated during the sentiment and prosody analysis stage). For example, for the output text "I am sad, but I am happy for you," the sentiment tags might be: "I (neutral), although (neutral), very (neutral), sad (negative - sadness), but (neutral), still (neutral), for (neutral), you (neutral), feel (neutral), happy (positive - joy)."
[0110] It should be noted that lip-sync generation, i.e., lip-sync animation, is primarily driven by the speech characteristics of the text to be output. In one embodiment, the action generation model converts the text into a phoneme sequence and, based on the standard phoneme-lip-sync correspondence, generates a sequence of lip-sync parameters that precisely matches the pronunciation of each phoneme on the timeline.
[0111] Facial expression generation, i.e., facial expressions, is mainly driven by the sentiment label and / or sentiment dimension score (such as valence, arousal) of each word. The model predicts the corresponding facial muscle movement parameters based on the sentiment label of the word. For example, when the model processes a word with the "happy" label, it generates parameters that drive the corners of the mouth to rise and the orbicularis oculi muscle to contract (forming a smiling eye); when it processes a word with the "sad" label, it generates parameters that drive the inner corners of the eyebrows to rise and the corners of the mouth to fall.
[0112] Because the motion generation model is built on a large model, its internal attention mechanism allows it to simultaneously attend to text content (for lip movements) and emotion tags (for facial expressions). When generating facial motion data for a given moment (or a specific frame), the model collaboratively processes both types of information, ensuring that the facial expressions and lip movements represented by the facial motion data for the same frame are synchronized.
[0113] It should be understood that by having an action generation model based on a large model simultaneously process the text to be output (driving lip movements) and the sentiment label of each word (driving facial expressions), the problem of asynchrony between facial expressions and lip movements is fundamentally solved. It ensures that the expression of emotion does not cover the entire sentence in a general way, but can be accurate to the word level, realizing real-time and accurate synchronization between facial expressions and semantic content. This effectively avoids the occurrence of inconsistencies such as smiling while reading sad words, and achieves deep binding and synchronization between facial expressions and content.
[0114] It should be understood that the powerful sequence generation capabilities of large models allow transitions between expressions (such as from sadness to happiness) to be handled very smoothly and naturally, rather than abruptly. This greatly improves the quality of the final generated visual data, making the facial expressions of digital humans more closely resemble the complex dynamics of real people, thereby improving the generation effect of digital human videos, significantly enhancing the user's interactive immersion, and ultimately improving the user's interactive experience.
[0115] The digital human video generation method provided in this invention ensures the synchronization of facial expressions and lip movements in the same frame represented by facial motion data through the above-described method, realizing real-time and accurate synchronization of facial expressions and semantic content, thereby improving the generation effect of digital human videos and ultimately enhancing the user's interactive experience.
[0116] Based on any of the above embodiments, in this method, the action generation model is generated in the following manner: Based on the sample sentiment and prosody analysis data, and the corresponding visual data labels, the large model is fine-tuned to obtain the action generation model.
[0117] The visual data tags include action data that represents cultural symbols.
[0118] Here, the sample sentiment and prosody analysis data constitute the input of the training data. Each sample sentiment and prosody analysis data has the same format as the sentiment and prosody analysis data described above, and will not be repeated here.
[0119] Here, visual data labels constitute the expected output of the training data. Each visual data label corresponds one-to-one with a sample emotion and prosody analysis data point, precisely describing the visual performance that the digital human should present when representing that sample data. These labels are typically structured data, such as sequences of facial motion unit parameters for each frame, and sequences of spatial coordinates of key body skeletal nodes.
[0120] For example, in international Chinese language education, simply expressing universal emotions like joy, anger, sorrow, and happiness is insufficient. To enable learners to better understand the deeper cultural connotations behind texts, digital human actions need to go beyond basic emotional expression and incorporate elements with specific cultural orientations. Based on this, action data representing cultural symbols is added to the visual data labels.
[0121] Here, the action data representing cultural symbols refers to body postures or gestures that have become conventional and have specific symbolic or referential meanings within a specific cultural context (specifically Chinese culture). These actions are often not purely expressions of emotion, but rather carry specific information, allusions, or social etiquette.
[0122] For example, when constructing the training dataset for fine-tuning, in addition to collecting action data of actors expressing common emotions, the following types of cultural symbolic action data will be specifically collected and labeled. For example, when processing text related to gratitude or requests, the visual data labels may include action data of hands clasped in a gesture of respect. For example, when expressing victory or "great!", in addition to the common thumbs-up, it may also include the heart gesture commonly used by young people in internet culture. For example, when expressing text related to contemplating ancient poems, the visual data labels may include a contemplative posture with one hand behind the back, the other gently touching the chin, and the body slightly leaning forward. For example, when the text mentions "pointing out the flaws in the world," the visual data labels may include a powerful waving gesture pointing into the distance.
[0123] It should be understood that by introducing motion data representing cultural symbols into the training data, the facial and body movements of digital humans are no longer limited to the basic emotional realm. They can perform actions with specific cultural connotations, upgrading the expression of digital humans from merely emotional to culturally rich and expressive. This has immeasurable value for cultural transmission in language teaching, greatly enriching the connotation and depth of digital human expression.
[0124] It should be understood that many cultural concepts are abstract for international Chinese learners. Presenting these concepts through digital avatars, using concrete actions that align with cultural habits, helps learners establish a strong connection between text and visual imagery. This allows for a more intuitive and profound understanding of the text's meaning, effectively lowering the barrier to cross-cultural learning. For example, through a gesture of clasped hands, learners can immediately understand a subtle yet formal expression of gratitude in the Chinese context, which is far more vivid than a simple textual explanation. This enhances the learner's (user's) understanding and memory.
[0125] The digital human video generation method provided in this embodiment of the invention greatly enriches the connotation and depth of digital human expression through the above-mentioned method, thereby improving the generation effect of digital human videos, enhancing users' understanding and memory of the digital human's output text, and ultimately improving users' interactive experience.
[0126] Based on any of the above embodiments, in this method, the voice data is generated in the following manner: The emotion and prosody analysis data are input into the speech synthesis model to obtain the speech data output by the speech synthesis model.
[0127] The speech synthesis model is built based on a large model and is used to control at least one parameter among pitch, volume, speech rate and timbre based on the emotion and prosody analysis data.
[0128] Here, the speech synthesis model is responsible for converting text information into speech waveforms that are audible to humans.
[0129] It should be noted that the speech synthesis model is built on a large model, thereby leveraging the powerful sequence generation capabilities and deep prosodic understanding of the pre-trained large model to improve the accuracy of speech data generation.
[0130] In one embodiment, the speech synthesis model is obtained by training or fine-tuning on massive text-speech data pairs, enabling it to generate highly natural, fluent, and expressive speech data.
[0131] In one embodiment, the dataset used to train the speech synthesis model typically requires a large amount of speech data recorded by professional voice actors in different emotional states (such as happiness, sadness, anger, etc.) to ensure that the speech synthesis model can learn the mapping relationship between different emotions and acoustic features.
[0132] Here, pitch, volume, speech rate, and timbre are the key acoustic features that constitute speech expressiveness. By learning from massive amounts of data, the speech synthesis model has mastered the complex mapping relationship between various indicators in emotion and prosody analysis data (such as emotion category, dimension score, tone sequence, number of vowels, etc.) and these acoustic features.
[0133] Here, pitch refers to the highness or lowness of sound frequency, which is mainly determined by the vibration frequency of the vocal cords. For example, the speech synthesis model controls the overall level and range of pitch variation based on emotion category and arousal score; for text with an emotion label of happiness or a high arousal score, the model generates a pitch curve that is generally higher and has greater fluctuations; while for sad text, it generates a pitch curve that is generally lower and smoother; in addition, it also finely adjusts the pitch contour of each character based on tonal analysis data to reflect the beauty of Chinese tones.
[0134] Here, volume (energy) refers to the intensity of sound, which is related to the amplitude of the sound wave. For example, speech synthesis models may amplify the energy of specific words based on emotion category (such as anger) or accent marks; the voice will be louder when expressing anger, and the energy will be significantly reduced when expressing whispers or sadness.
[0135] Here, speech rate refers to the number of syllables or words spoken per unit of time. For example, speech synthesis models adjust speech rate based on emotional arousal; the speech rate speeds up when expressing excitement or nervousness, and slows down when expressing hesitation or sadness. Furthermore, the model inserts natural silent pauses at appropriate locations based on pause suggestions from prosodic analysis data.
[0136] Here, timbre refers to the distinctive characteristics of a sound, which determines our ability to distinguish different human voices or instrumental sounds. In this embodiment of the invention, it can also refer to specific vocalization methods. For example, by including samples with special timbres in the training data, the model can learn to generate special timbres that go beyond regular speech. For instance, for texts with strong emotions, the model can synthesize speech with special effects such as crying, laughing, and panting sounds, greatly enhancing the realism of the emotions.
[0137] It should be understood that by directly linking emotion and prosody analysis data with key acoustic parameters such as pitch, volume, speech rate, and timbre, the emotional expression in speech is no longer vague and uncontrollable. Based on precise analysis results, it can make meticulous adjustments to every physical dimension of the sound, thereby achieving precise and quantitative control over emotional expression.
[0138] It should be understood that because speech synthesis models are built on large models and can comprehensively control multiple acoustic parameters, the speech they generate far surpasses traditional methods in fluency, rhythm, and emotional realism. In particular, the ability to synthesize special timbres (such as crying and laughing) enables digital humans to express a wider range of emotions, greatly enhancing the realism and immersion of the interaction.
[0139] The digital human video generation method provided in this embodiment of the invention makes meticulous adjustments to each physical dimension of sound through the above-described method, thereby achieving precise and quantitative control of emotional expression, improving the generation effect of voice data, thus improving the generation effect of digital human video, and ultimately enhancing the user's interactive experience.
[0140] Based on any of the above embodiments, in this method, the emotion and prosody analysis model, the action generation model, and the speech synthesis model can be jointly trained. That is, a multi-task learning model is implemented, which shares a text encoder (emotion and prosody analysis model) but has different decoders (action generation model and speech synthesis model) for generating speech and visual data. Based on this, the multi-task learning model can internally learn the intrinsic connection between speech and action; for example, when a loud laugh is produced, it necessarily corresponds to a laughing facial expression. Based on this, multimodal fusion and collaborative training achieve a 1+1+1>3 effect, ensuring that facial expressions, voices, and tones are synchronized and consistent, rather than operating independently.
[0141] To facilitate understanding of the above embodiments, a specific embodiment will be described here. Figure 2 As shown, the process is divided into four stages: Stage 1, the output text is input into the emotion and prosody analysis model to obtain the emotion and prosody analysis data output by the model; Stage 2, the emotion and prosody analysis data is input into the action generation model to obtain the visual data output by the model; Stage 3, the emotion and prosody analysis data is input into the speech synthesis model to obtain the speech data output by the model; Stage 4, the speech data and visual data are input into the multimodal fusion and synchronization module to obtain a digital human video with synchronized audio and video and consistent emotion.
[0142] Based on any of the above embodiments, in this method, the sentiment and prosody analysis data includes the digital human's output text and the sentiment tag of each word in the output text; correspondingly, step 130 includes: When the text to be output includes onomatopoeia, the timbre text corresponding to the onomatopoeia is determined from a preset timbre text set based on the sentiment tag of the onomatopoeia, and the onomatopoeia is replaced with text based on the timbre text; the timbre text is used to indicate the timbre of the onomatopoeia output. Based on the sentiment and prosody analysis data after text replacement, voice data and visual data that match the sentiment and prosody of the output text are generated.
[0143] Considering that in real human communication, the expression of onomatopoeia (such as "haha," "uh-huh," and "ah") is not static, and its specific pronunciation and the emotions it carries are highly dependent on its context, traditional text-to-speech conversion often mechanically and verbatim pronounces onomatopoeia according to its literal pronunciation (for example, pronouncing "hahahaha" as four separate "ha" sounds). This sounds very unnatural and seriously undermines the realism of the interaction. Therefore, this invention automates the processing of onomatopoeia.
[0144] Here, the preset timbre text set includes multiple preset timbre texts. Based on this, it is necessary to pre-store the mapping relationships between various pre-recorded or synthesized non-verbal sound segments (i.e., timbres) representing different emotional tones and specific text tags (preset timbre texts). These timbres can be real audio files such as laughter, crying, and sighing.
[0145] Here, timbre text is a special text tag or instruction that is not used for direct pronunciation, but rather to instruct the subsequent speech synthesis model which specific timbre audio file to call.
[0146] Specifically, in the output text, the original onomatopoeic words are replaced with determined timbre text, thereby generating a sentiment and prosody analysis data after text replacement.
[0147] In one specific embodiment, upon detecting an onomatopoeic word, it is not immediately used for subsequent generation. Instead, the sentiment tag obtained in the sentiment and prosody analysis stage is used to analyze the context of the onomatopoeic word to determine its true social intention and emotional connotation. This is because the same onomatopoeic word may represent completely different emotions in different contexts.
[0148] For example, the onomatopoeic word "haha" might be labeled as "positive - laughing" in "This joke is so funny, haha"; as "Oh, really? Haha" it might be labeled as "neutral - social agreement"; and as "You think you can win? Haha" it might be labeled as "negative - sarcastic".
[0149] In one specific embodiment, when the speech synthesis model parses the timbre text, it does not attempt to read the text aloud, but instead directly retrieves the audio file corresponding to the timbre text from the audio library and seamlessly splices it into the final speech data.
[0150] In one specific embodiment, when the motion generation model parses the timbre text, it generates facial and body movements that match the timbre text.
[0151] It should be noted that the speech synthesis method for generating speech data supports the fusion of multiple timbres, which means that multi-track timbres can be recorded synchronously.
[0152] It should be understood that by using onomatopoeic sentiment analysis and text replacement, the simple pronunciation processing of onomatopoeia has been fundamentally changed. It transforms the processing of onomatopoeia from mechanical reading aloud into the direct invocation and playback of authentic emotional sounds (such as laughter and crying), greatly enhancing the realism and naturalness of the interaction.
[0153] The digital human video generation method provided in this embodiment of the invention avoids the mechanical reading of onomatopoeia through the above-mentioned method, thereby improving the generation effect of voice data, thus improving the generation effect of digital human video and ultimately enhancing the user's interactive experience; and can also generate visual data that better matches the real emotions of onomatopoeia, thereby improving the generation effect of digital human video and ultimately enhancing the user's interactive experience.
[0154] Based on any of the above embodiments, in this method, step 140 includes: Based on the emotion and rhythm analysis data, special effects data that match the emotion and rhythm of the output text are determined from a preset set of special effects. Based on the voice data, the visual data, and the special effects data, a digital human video is generated that matches the emotion and rhythm of the output text.
[0155] Here, the preset effects set includes a variety of preset effects data. For example, it is a database containing various visual effects resources. These effects can be 2D or 3D animations, particle effects, lighting changes, or background elements. Each effect is pre-associated with one or more specific emotional tags, emotional dimension scores, or cultural concepts. Specifically, in the context of international Chinese language teaching, this effects set can be a library of effects reflecting teaching preferences, where the styles and content of the effects are carefully selected to ensure they meet the needs of the teaching environment and avoid overly exaggerated or entertainment-oriented elements that could interfere with learning.
[0156] In one embodiment, the special effects data includes the identifier of the special effects resource, the start and end times of playback, the position, size, transparency, and other parameters required for rendering.
[0157] For example, information from sentiment and prosody analysis data (e.g., sentiment category, arousal score, etc.) is used to search and match within a preset set of effects. For instance, if the sentiment tag is "happy" and the arousal score is high, an effect might be matched that causes small fireworks to bloom around the character or the background to turn warmer and show twinkling stars. For instance, if the sentiment tag is "sad," an effect might be matched that causes the overall saturation of the image to decrease or rain clouds to appear above the digital character's head. For instance, if the text content involves thinking or wisdom, an effect might be matched that causes a glowing light bulb to appear above the digital character's head.
[0158] It should be understood that by introducing visual effects synchronized with emotions, abstract emotions are made concrete and visual. This makes the transmission of emotions no longer solely dependent on the digital human's own expressions and movements, but rather amplified and emphasized through additional visual symbols. This allows users (especially learners whose language comprehension abilities are still developing) to understand the emotions and intentions that the digital human is trying to express more quickly and intuitively, thereby improving the user's interactive experience.
[0159] It should be understood that the addition of special effects breaks the monotony of traditional digital human videos, which typically feature only one person per scene, adding visual interest and dynamic changes. In educational settings, appropriate and content-matched special effects can effectively attract learners' attention, stimulate their interest, and make what might otherwise be a tedious language practice process more vivid and engaging, thus giving Chinese a visual warmth and expression. This enriches the expressiveness and fun of the visuals, thereby improving the overall quality of digital human videos.
[0160] It should be understood that in certain scenarios, a simple visual symbol (such as a thumbs-up or a light bulb) can convey far more information and at a faster pace than complex verbal descriptions. By combining these symbolic effects with linguistic content, multi-channel information transmission is achieved, improving the efficiency of communication and teaching. In other words, improving the efficiency of information transmission enhances the user's interactive experience.
[0161] The digital human video generation method provided in this invention, through the above-described method, by adding special effects that match emotions, enables a faster and more intuitive understanding of the emotions and intentions that the digital human wants to express, thereby improving the user's interactive experience; it also enriches the expressiveness and fun of the picture, thereby improving the generation effect of the digital human video, and further improving the user's interactive experience; and by improving the efficiency of information transmission through special effects, it also improves the user's interactive experience.
[0162] Based on any of the above embodiments, in this method, step 110 includes: Acquire user-input voice data and user visual data; the user visual data includes user facial movement data and / or user body movement data. Based on the user's voice data and visual data, determine the user's emotional state data; The emotional state data and dialogue context are input into the dialogue generation model to obtain the output text output by the dialogue generation model; the dialogue context is the dialogue between the user and the digital human, and the dialogue generation model is built based on a large model.
[0163] Here, user voice data refers to raw audio data containing the user's voice, collected through an audio acquisition device (such as a microphone).
[0164] Here, user visual data refers to video stream data captured by a camera that includes real-time images of the user. This user visual data can be further processed to extract user facial movement data (e.g., changes in eyebrows, eyes, and mouth) and / or user body movement data (e.g., gestures, shrugging, nodding, shaking the head, etc.).
[0165] In one embodiment, the user visual data further includes background environment data. In one embodiment, the user visual data further includes user clothing data. In one embodiment, body movement data includes body tilt data.
[0166] Specifically, multimodal data (user voice data and user visual data) are comprehensively analyzed to determine the user's true emotional state while speaking, thereby identifying the user's emotional state data.
[0167] In one embodiment, voice sentiment analysis is performed on user voice data to generate first sentiment state data, and visual sentiment analysis is performed on user visual data to generate second sentiment state data. Based on the first and second sentiment state data, the user's sentiment state data is determined. Specifically, the acoustic features of the user voice data (such as pitch, intensity, speech rate, spectrum, etc.) are analyzed to identify the emotions contained in the voice; the user visual data (including facial movements and body movements) are analyzed to identify the emotional information conveyed by the user's expressions and postures; the emotional information obtained from the above two modal analyses (voice and vision) is fused. The purpose of fusion is to obtain a more accurate and robust judgment result than any single modal analysis. For example, when a user says "I'm fine," but the tone of voice is low and the visual expression shows a frown and downturned mouth, fusion can determine that the user's true emotion is sadness or frustration, rather than the neutrality of the text.
[0168] Furthermore, the input text of the user's voice data is subjected to sentiment and prosody analysis to generate third sentiment state data; based on the first sentiment state data, the second sentiment state data, and the third sentiment state data, the user's sentiment state data is determined.
[0169] The emotional state data may include, but is not limited to, at least one of the following: the user's current emotional category (e.g., happy, sad), dimensional ratings (valence, arousal, dominance, etc.), and confidence levels, etc. In one embodiment, the emotional state data is structured data.
[0170] In one embodiment, the dialogue generation model is obtained by fine-tuning a large model on a massive dialogue corpus, thereby giving it powerful language understanding and generation capabilities to improve the accuracy of the generated output text.
[0171] Here, the dialogue context refers to the history of conversations between the user and the digital human so far. It provides the dialogue generation model with the necessary information to understand the current topic and context of the conversation.
[0172] For example, suppose in a Chinese teaching scenario, a user (learner) tries to construct a sentence but stutters and finally sighs, saying, "Oh, I'm so stupid." The system acquires the user's voice recording of "Oh, I'm so stupid" and visually observes the user frowning and looking down. After comprehensive analysis, the user's emotional state is determined to be "negative-frustrated-dejected." The dialogue context is: "The user is practicing sentence construction." Traditional dialogue generation models might only generate a bland response like "Please don't say that" based on the literal meaning of "I'm so stupid." However, the dialogue generation model in this embodiment of the invention, upon receiving the user's "frustrated" emotion, generates a warmer and more encouraging response. Its output text might be: "It's okay, everyone encounters difficulties at the beginning of learning, that's normal. Your pronunciation was already very standard, shall we try again?" Based on this generated output text, a digital human video with encouraging emoticons and tone can be subsequently generated.
[0173] It should be understood that the embodiments of the present invention transform the digital human from a passive machine that merely responds to the literal meaning of text. By introducing the perception of the user's multimodal emotional state, the digital human's responses can take the user's feelings into account, thus providing more empathetic interactive feedback. This empathic ability is key to establishing deep and effective human-computer interaction, significantly enhancing the user's trust and willingness to communicate, thereby improving the user's interactive experience.
[0174] It should be understood that by using the user's emotional state as a direct input to the dialogue generation model, the digital human's language strategies (whether to comfort, encourage, or share joy) can be dynamically adjusted according to the user's real-time emotions, making the entire dialogue process more like real interpersonal communication, more intelligent and humanized, thus improving the level of intelligence and humanization of the interaction, thereby enhancing the user's interactive experience.
[0175] It should be understood that connecting the two stages of perceiving user emotions and expressing one's own emotions forms a complete emotional interaction loop of perception-understanding-response-expression. This enables digital humans not only to understand users' moods but also to make considerate responses, thereby providing users (especially in educational scenarios that require emotional support) with a high-quality immersive interactive experience and enhancing the user's overall interactive experience.
[0176] The digital human video generation method provided in this embodiment of the invention enables the digital human's response to take into account the user's emotions, thereby providing warmer interactive feedback and improving the user's interactive experience; it also improves the intelligence and humanization of the interaction, thereby enhancing the user's interactive experience.
[0177] Based on any of the above embodiments, in this method, determining the user's emotional state data based on the user's voice data and user visual data includes: Emotion and prosodic analysis is performed on the acoustic features of the user's voice data to generate first emotional state data; Sentiment analysis is performed on the visual features of the user visual data to generate second emotional state data; the user visual data also includes background environment data and user clothing data, and the body movement data includes body tilt data. Based on the first emotional state data and the second emotional state data, the user's emotional state data is determined.
[0178] Here, acoustic features may include, but are not limited to, at least one of the following: pitch, intensity (energy), speech rate, spectral features, etc.
[0179] In one specific embodiment, a pre-trained speech emotion recognition model is used to analyze the acoustic features of user speech data to determine the emotions contained in the user speech data. The output is a first emotional state data based on the speech modality, which may include, but is not limited to, at least one of the following: the category of the user's current emotion (such as happiness, sadness), dimensional scores (valence, arousal, dominance, etc.), and confidence, etc.
[0180] Here, visual features include features extracted based on user facial motion data and features extracted based on user body motion data. For example, by analyzing user facial motion data, facial features are obtained, such as raising or frowning eyebrows, contracting corners of the eyes, and raising or lowering corners of the mouth; the rotation angle and direction of the head are analyzed, such as nodding, shaking, and tilting the head.
[0181] In one specific embodiment, background environment data can be obtained through image segmentation and scene recognition technology, thereby analyzing the characteristics of the user's environment, such as the cleanliness of the environment, the orderliness of the items, and the brightness of the light.
[0182] Considering that an individual's external environment is, to some extent, an extension of their inner state and long-term habits—for example, a neat and orderly learning environment may indicate that the user has high levels of organization and self-discipline—this background environment is considered an important visual feature to aid in judging the user's emotional state, thereby improving the accuracy of determining secondary emotional state data.
[0183] In one specific embodiment, object detection and color analysis technologies are used to identify user clothing data (such as color, style, material, etc.) of the clothing currently worn by the user.
[0184] Considering that clothing choices often reflect a person's current mood or stable personality traits—for example, bright colors may be associated with positive, extroverted emotions, while muted colors may be associated with low, introverted emotions—this clothing data is used as an important visual feature to help determine a user's emotional state, thereby improving the accuracy of secondary emotional state data.
[0185] In one specific embodiment, a human posture estimation algorithm is used to calculate the body tilt data and detect the tilt angle and direction of the user's upper body relative to the vertical direction.
[0186] In interpersonal psychology, body orientation and angle often reflect an individual's level of engagement and psychological distance. For example, leaning forward typically indicates high interest, focus, and participation in the current topic, while leaning back or to one side may suggest relaxation, doubt, disinterest, or even psychological resistance. Therefore, this body tilt is considered an important visual feature to help determine a user's learning interest and engagement, thereby improving the accuracy of secondary affective state data.
[0187] In one specific embodiment, visual features are input into a pre-trained visual emotion recognition model to generate a second emotion state data based on the visual modality.
[0188] In one specific embodiment, first emotional state data and second emotional state data are fused using a multimodal method to obtain first emotional state data. For example, a fusion model (e.g., weighted fusion or a deep fusion network based on an attention mechanism) is used to integrate the first emotional state data obtained from the speech modality and the second emotional state data obtained from the visual modality. Through fusion, more accurate judgments can be made than with a single modality. For example, even if the user's voice and facial expressions appear relatively calm (first emotional state data and second emotional state data), if the system detects that the user's body is continuously leaning forward (second emotional state data), the fusion model can determine that the user is actually in a calm but highly focused state.
[0189] It should be understood that by supplementing visual analysis with data on body tilt, background environment, and clothing, the information sources for emotion recognition are greatly expanded. This extends the analysis from the user's instantaneous, highly controllable facial expressions to body postures and environmental features that better reflect their subconscious state and long-term habits. This multi-dimensional information input makes the judgment of the user's emotional state more comprehensive and three-dimensional, thereby improving the accuracy of the second emotional state data and the final fused emotional state data.
[0190] It should be understood that features such as body tilt, environment, and clothing can not only reflect emotions but also reveal, to some extent, a user's learning interests, level of concentration, personality habits, and values. By analyzing these features, digital humans can gain deeper insights into users, enabling subsequent interactions to go beyond emotional empathy to include strategic personalization (e.g., providing more challenging content to highly focused users).
[0191] It should be understood that traditional visual sentiment analysis methods may fail in certain scenarios (such as poor lighting making facial expressions difficult to recognize, or when users intentionally hide their expressions). Features such as body tilt are relatively difficult to fake and have lower environmental requirements. Introducing these new features can serve as an effective supplement, significantly enhancing the robustness and reliability of the entire sentiment recognition process in complex environments, thereby improving the accuracy of determining the second sentiment state data and the final fused sentiment state data.
[0192] The digital human video generation method provided in this embodiment of the invention, through the above-described method and multi-dimensional information input, makes the judgment of the user's emotional state more comprehensive and three-dimensional, thereby improving the accuracy of the determination of the second emotional state data and the final fused emotional state data, thereby improving the accuracy of the output text generation, and ultimately improving the generation effect of the digital human video.
[0193] Based on any of the above embodiments, in this method, determining the user's emotional state data based on the first emotional state data and the second emotional state data includes: Emotion and prosody analysis is performed on the input text of the user's voice data to generate third emotional state data; Based on the first emotional state data, the second emotional state data, and the third emotional state data, the user's emotional state data is determined.
[0194] Here, the input text is the transcribed text obtained by speech recognition of the user's voice data, that is, the user's voice signal is converted into a computer-readable text string.
[0195] Here, the sentiment and prosodic analysis of the input text is essentially similar to the sentiment and prosodic analysis of the output text described above, and will not be repeated here. For example, it involves identifying words with positive or negative sentiment in the text; understanding the syntactic structure and deep semantics of the entire sentence to determine its true intention, especially for texts containing complex linguistic phenomena such as irony and metaphor; and analyzing the text's tones, parallelism, and vowel distribution to uncover its linguistic emotional potential.
[0196] Here, the third sentiment state data based on text modality is basically similar to the sentiment and prosody analysis data mentioned above, and will not be repeated here. This third sentiment state data reveals the emotions that users want to express from the perspective of language content.
[0197] In one specific embodiment, a multimodal fusion model is employed to comprehensively process three sentiment state data. This multimodal fusion model is able to learn the correlation and complementarity between different modal information and intelligently weight them.
[0198] It should be understood that by further introducing sentiment and prosodic analysis of the input text after user speech conversion, and combining textual (verbal) information with speech and visual (non-verbal) information, a comprehensive multimodal sentiment recognition system integrating speech, vision, and text is constructed. This makes the basis for sentiment judgment more comprehensive.
[0199] It should be understood that human emotional expression is often complex, and verbal and nonverbal cues can be inconsistent, such as irony or insincerity. Relying on any single modality alone is difficult to identify accurately. By cross-validating data from all three modalities, the accuracy of identifying such complex and inconsistent emotional expressions can be greatly improved, giving digital humans a stronger ability to read between the lines.
[0200] It should be understood that by integrating data from three independent information channels, the system can effectively overcome the noise, ambiguity, or information gaps that may exist in a single modality (for example, unreliable voice information in noisy environments or unreliable visual information in backlit environments). This multimodal complementarity makes the final determined user emotional state data more accurate and robust, laying the most solid foundation for generating truly empathetic interactive responses.
[0201] The digital human video generation method provided in this embodiment of the invention improves the accuracy of determining the user's emotional state data through the above-mentioned methods, thereby improving the accuracy of output text generation, thus improving the generation effect of digital human videos, and ultimately enhancing the user's interactive experience.
[0202] Based on any of the above embodiments, in this method, the step of inputting the emotional state data and the dialogue context into the dialogue generation model to obtain the output text output by the dialogue generation model includes: The visual features of the user's visual data are analyzed to generate the user's age data. The emotional state data, dialogue context, and age data are input into the dialogue generation model to obtain the output text output by the dialogue generation model.
[0203] In one embodiment, the visual features used for age analysis are primarily derived from the user's facial images, including but not limited to skin texture (such as the depth and distribution of wrinkles), fullness of facial contours, hair color, and other biometric features related to aging.
[0204] In one specific embodiment, a user age analysis model is used to analyze a user's visual features to estimate the user's age. This user age analysis model is a deep learning-based computer vision model. In one embodiment, the user age analysis model is trained on a large dataset of age-labeled face images, learning to extract age-related patterns from the visual features of faces.
[0205] Here, age data can be a specific age estimate (e.g., 25 years old) or an age group division (e.g., teenagers, young adults, middle-aged, elderly, etc.).
[0206] It should be understood that using age data as input allows dialogue generation models to adjust the language style, vocabulary habits, example methods, and knowledge boundaries of the generated content, making the output text more easily understood and accepted by users of specific age groups, and more likely to evoke resonance.
[0207] It should be understood that by using visual analysis to determine the user's age and using it as a key basis for generating responses, the interaction of digital humans can adjust its communication strategies according to the user's age, thereby achieving a high degree of personalization and intergenerational adaptability.
[0208] It should be understood that using language styles and knowledge backgrounds appropriate for the user's age group is a fundamental principle of effective human communication. Applying this principle to digital humans makes their output text more relatable to users of the same age. This sense of resonance can quickly bridge the psychological distance between users and the digital human, reduce communication barriers, and increase user acceptance of the digital human and its conveyed information (such as educational content).
[0209] It should be understood that a digital human capable of adjusting its speech based on the person it is interacting with has a behavior pattern closer to that of a real human, thus appearing more intelligent and trustworthy. This sense of realism in detail is crucial for building long-term, stable, and effective human-machine relationships (especially in fields such as education and companionship), enhancing the authenticity and credibility of the digital human role and greatly improving the user experience.
[0210] The digital human video generation method provided in this embodiment of the invention makes the output text easier for users of specific age groups to understand and accept, and easier to evoke resonance, thereby improving the user's interactive experience.
[0211] Based on the above embodiments, the digital human is systematically trained from several dimensions, including voice, tone, facial expression, and movement. This significantly improves the emotional connection between the digital human and the text, enabling the digital human to provide emotional feedback through tone of voice, facial expression, movement, special effects, and language content, making the interaction more realistic. Therefore, AI-powered Chinese education products will place greater emphasis on the authenticity of emotional interaction and the accuracy of cultural transmission. The digital human will not only be able to understand and express basic emotions but also capture more subtle cultural contextual differences, providing international Chinese learners with a more in-depth and engaging learning experience.
[0212] The apparatus for generating digital human videos provided by the present invention will be described below. The apparatus for generating digital human videos described below can be referred to in correspondence with the method for generating digital human videos described above.
[0213] Figure 3 This is a schematic diagram of the structure of the digital human video generation device provided by the present invention, as shown below. Figure 3 As shown, the digital human video generation device includes: a text acquisition module 310, a sentiment analysis module 320, a data generation module 330, and a video generation module 340.
[0214] The text acquisition module 310 is used to acquire the output text to be output by the digital human.
[0215] The sentiment analysis module 320 is used to perform sentiment and rhythm analysis on the output text and generate sentiment and rhythm analysis data; the sentiment and rhythm analysis includes the analysis of the tone and / or parallelism of the output text.
[0216] The data generation module 330 is used to generate voice data and visual data that match the emotion and rhythm of the output text based on the emotion and rhythm analysis data; the visual data includes facial motion data and / or body motion data.
[0217] The video generation module 340 is used to generate a digital human video that matches the emotion and rhythm of the output text based on the voice data and the visual data.
[0218] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a method for generating a digital human video. This method includes: acquiring the output text to be output by the digital human; performing emotion and prosody analysis on the output text to generate emotion and prosody analysis data; the emotion and prosody analysis includes analysis of the tonal patterns and / or parallelism of the output text; based on the emotion and prosody analysis data, generating speech data and visual data that match the emotion and prosody of the output text; the visual data includes facial movement data and / or body movement data; and based on the speech data and the visual data, generating a digital human video that matches the emotion and prosody of the output text.
[0219] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0220] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the digital human video generation method provided by the above methods. The method includes: acquiring output text to be output by the digital human; performing emotion and rhythm analysis on the output text to generate emotion and rhythm analysis data; the emotion and rhythm analysis includes analysis of the tonal patterns and / or parallelism of the output text; generating voice data and visual data that match the emotion and rhythm of the output text based on the emotion and rhythm analysis data; the visual data includes facial movement data and / or body movement data; and generating a digital human video that matches the emotion and rhythm of the output text based on the voice data and the visual data.
[0221] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for generating a digital human video provided by the methods described above. This method includes: acquiring output text to be output by the digital human; performing emotion and prosody analysis on the output text to generate emotion and prosody analysis data; the emotion and prosody analysis including analysis of the tonal patterns and / or parallelism of the output text; generating, based on the emotion and prosody analysis data, speech data and visual data matching the emotion and prosody of the output text; the visual data including facial movement data and / or body movement data; and generating a digital human video matching the emotion and prosody of the output text based on the speech data and the visual data.
[0222] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0223] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating digital human videos, characterized in that, include: Obtain the output text to be output by the digital human; Sentiment and prosody analysis is performed on the output text to generate sentiment and prosody analysis data; The emotion and rhythm analysis includes the analysis of the tone and / or parallelism of the output text; Based on the emotion and prosody analysis data, voice data and visual data that match the emotion and prosody of the output text are generated; the visual data includes facial motion data and / or body motion data. Based on the voice data and the visual data, a digital human video is generated that matches the emotion and rhythm of the output text.
2. The method for generating digital human videos according to claim 1, characterized in that, The process involves performing sentiment and prosody analysis on the output text to generate sentiment and prosody analysis data, including... Based on the output text and the preset prompt word template, an input prompt word is generated; The input prompt words are input into the sentiment and prosody analysis model to obtain the sentiment and prosody analysis data output by the sentiment and prosody analysis model; The sentiment and prosody analysis model is obtained by fine-tuning a pre-trained large model. The sentiment and prosody analysis model is used to perform sentiment and prosody analysis on the output text. The sentiment and prosody analysis also includes the analysis of the number of vowels in the output text.
3. The method for generating digital human videos according to claim 2, characterized in that, The sentiment and prosody analysis also includes the analysis of the sentiment category of the output text, and the dimensional scoring of the sentiment category, wherein the dimensional scoring includes at least one of valence score, arousal score and dominance score.
4. The method for generating digital human videos according to claim 1, characterized in that, The visual data is generated based on the following method: The emotion and prosody analysis data is input into the action generation model to obtain the visual data output by the action generation model; the emotion and prosody analysis data includes the digital human's output text and the emotion tag of each word in the output text; The action generation model is built based on a large model; the action generation model is used to synchronize facial expressions and lip movements based on the text to be output and each of the emotion tags, so as to ensure that the facial expressions and lip movements of the same frame represented by the facial action data are synchronized.
5. The method for generating digital human videos according to claim 4, characterized in that, The action generation model is generated based on the following method: Based on the sample sentiment and prosody analysis data, and the visual data labels corresponding to the sample sentiment and prosody analysis data, the large model is fine-tuned to obtain the action generation model. The visual data tags include action data that represents cultural symbols.
6. The method for generating digital human videos according to claim 1, characterized in that, The voice data is generated based on the following method: The emotion and prosody analysis data are input into the speech synthesis model to obtain the speech data output by the speech synthesis model; The speech synthesis model is built based on a large model and is used to control at least one parameter among pitch, volume, speech rate and timbre based on the emotion and prosody analysis data.
7. The method for generating digital human videos according to claim 1, characterized in that, The sentiment and prosody analysis data includes the digital human's output text and the sentiment tag for each word in the output text; The step of generating speech and visual data that match the emotion and prosody of the output text based on the emotion and prosody analysis data includes: When the text to be output includes onomatopoeia, the timbre text corresponding to the onomatopoeia is determined from a preset timbre text set based on the sentiment tag of the onomatopoeia, and the onomatopoeia is replaced with text based on the timbre text; the timbre text is used to indicate the timbre of the onomatopoeia output. Based on the sentiment and prosody analysis data after text replacement, voice data and visual data that match the sentiment and prosody of the output text are generated.
8. The method for generating digital human videos according to claim 1, characterized in that, The step of generating a digital human video that matches the emotion and rhythm of the output text based on the speech data and the visual data includes: Based on the emotion and rhythm analysis data, special effects data that match the emotion and rhythm of the output text are determined from a preset set of special effects. Based on the voice data, the visual data, and the special effects data, a digital human video is generated that matches the emotion and rhythm of the output text.
9. The method for generating digital human videos according to any one of claims 1 to 8, characterized in that, The process of obtaining the output text to be output by the digital human includes: Acquire user-input voice data and user visual data; the user visual data includes user facial movement data and / or user body movement data. Based on the user's voice data and visual data, determine the user's emotional state data; The emotional state data and dialogue context are input into the dialogue generation model to obtain the output text output by the dialogue generation model; the dialogue context is the dialogue between the user and the digital human, and the dialogue generation model is built based on a large model.
10. The method for generating digital human videos according to claim 9, characterized in that, The process of determining the user's emotional state data based on the user's voice data and visual data includes: Emotion and prosodic analysis is performed on the acoustic features of the user's voice data to generate first emotional state data; Sentiment analysis is performed on the visual features of the user visual data to generate second emotional state data; the user visual data also includes background environment data and user clothing data, and the body movement data includes body tilt data. Based on the first emotional state data and the second emotional state data, the user's emotional state data is determined.
11. The method for generating digital human videos according to claim 10, characterized in that, The step of determining the user's emotional state data based on the first emotional state data and the second emotional state data includes: Emotion and prosody analysis is performed on the input text of the user's voice data to generate third emotional state data; Based on the first emotional state data, the second emotional state data, and the third emotional state data, the user's emotional state data is determined.
12. The method for generating digital human videos according to claim 9, characterized in that, The step of inputting the emotional state data and dialogue context into the dialogue generation model to obtain the output text output by the dialogue generation model includes: The visual features of the user's visual data are analyzed to generate the user's age data. The emotional state data, dialogue context, and age data are input into the dialogue generation model to obtain the output text output by the dialogue generation model.
13. A device for generating digital human videos, characterized in that, include: The text acquisition module is used to acquire the output text to be output by the digital human. The sentiment analysis module is used to perform sentiment and prosody analysis on the output text and generate sentiment and prosody analysis data. The emotion and rhythm analysis includes the analysis of the tone and / or parallelism of the output text; The data generation module is used to generate voice data and visual data that match the emotion and rhythm of the output text based on the emotion and rhythm analysis data; the visual data includes facial motion data and / or body motion data. The video generation module is used to generate a digital human video that matches the emotion and rhythm of the output text based on the voice data and the visual data.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for generating digital human videos as described in any one of claims 1 to 12.
15. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for generating digital human videos as described in any one of claims 1 to 12.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for generating digital human videos as described in any one of claims 1 to 12.