Digital human video generation method, electronic device and program product
By receiving the text to be played, recognizing and processing the motion information in blocks, and obtaining the matching digital human animation files and audio data, the problems of high cost, low efficiency and insufficient motion continuity in the existing technology are solved, and efficient and natural digital human video generation is achieved.
Patent Information
- Application Number
- CN202511893842.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-01-16
AI Technical Summary
Existing methods for generating digital human videos are costly, inefficient, and lack sufficient motion continuity and audiovisual synchronization, failing to effectively avoid lip misalignment or screen stuttering caused by premature animation termination.
By receiving the text to be played, recognizing and processing the motion information in blocks, and obtaining the matching digital human animation file and audio data, lip-syncing is achieved. When the audio duration exceeds the animation duration, a general transition animation is automatically spliced to ensure motion continuity and lip-syncing.
It achieves precise matching between text content and digital human movements, avoiding lip misalignment and screen stuttering, and improving the smoothness, realism, and user experience of digital human broadcasting.
Smart Images

Figure CN121353484A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of digital human, in particular to a digital human video generation method, an electronic device and a program product. BACKGROUND
[0002] In recent years, digital human technology is increasingly widely applied in the fields of virtual reality (VR), film and television production and intelligent human-computer interaction, and the core thereof is to replace real people with virtual images to complete information transmission, content creation or interactive services. However, the existing digital human video generation mainly relies on two traditional methods: one is to manually draw key frames by professional animators, and to adjust actions, expressions and scenes frame by frame, which is tedious (actions and expressions need to be designed frame by frame), time-consuming (a single video production takes several hours to several days) and highly dependent on professional skills (need to master animation software); the other is to use high-cost motion capture equipment (such as optical motion capture system, inertial sensor, etc.) to collect real human action data, and then map it to a digital human model through post-processing, which improves the authenticity of actions, but has the defects of expensive equipment, strict environmental requirements and complex post-processing. The above methods generally have the common problems of high cost, low efficiency and high technical threshold. Therefore, there is an urgent need for a digital human video solution that is highly automated, low-cost, can generate coherent actions and is synchronized in vision and hearing. SUMMARY
[0003] The present disclosure provides a digital human video generation method, an electronic device and a program product.
[0004] According to one aspect of the present disclosure, a digital human video generation method is provided, comprising: receiving a text to be broadcast; performing action information recognition on the text to be broadcast to obtain action information contained in the text to be broadcast; performing block processing on the text to be broadcast based on the action information to obtain a plurality of text blocks, wherein each text block corresponds to an action information, and adjacent text blocks correspond to different action information; obtaining a first digital human animation file corresponding to each text block based on the action information corresponding to each text block, wherein the digital human in the first digital human animation file has an action corresponding to the action information; obtaining first audio data corresponding to each text block based on the content of each text block; performing lip driving on the first digital human animation file corresponding to each text block based on the first audio data corresponding to each text block to obtain a second digital human animation file corresponding to each text block, wherein the lip shape of the digital human in the second digital human animation file is aligned with the first audio data on a time axis; and synthesizing the first audio data and the second digital human animation file corresponding to each text block to generate the digital human video. In the process of performing lip driving on the first digital human animation file corresponding to each text block based on the first audio data corresponding to each text block to obtain the second digital human animation file corresponding to each text block: it is determined whether the time length of the first audio data corresponding to each text block is consistent with the time length of the first digital human animation file; in the case where the time length of the first audio data is greater than the time length of the first digital human animation file, a general transition digital human animation file is spliced at the end position of the first digital human animation file, and the spliced result is taken as a third digital human animation file corresponding to each text block; the third digital human animation file corresponding to each text block is driven based on the first audio data corresponding to each text block to obtain the second digital human animation file corresponding to each text block.
[0005] According to the technical solution of one aspect, by performing action information recognition on the received text to be broadcast and performing block processing accordingly, accurate matching of text content and digital human action can be achieved. By matching a first digital human animation file corresponding to the action for each text block and combining it with the audio data for lip driving, a digital human video with highly synchronized lip shape and voice and natural and coherent action can be generated. By automatically splicing a general transition animation when the audio duration exceeds the animation duration, lip misalignment or screen freezing caused by early ending of the animation can be effectively avoided, thereby improving the smoothness, realism and user experience of digital human broadcasting.
[0006] According to at least one embodiment of this disclosure, the action information recognition of the text to be broadcast to obtain the action information contained in the text to be broadcast includes: segmenting the text to be broadcast into semantic units to obtain multiple semantic units; matching action information based on the semantic units to obtain the action information corresponding to each semantic unit; merging adjacent semantic units with the same action information; and using the merged action information as the action information contained in the text to be broadcast.
[0007] According to the technical solution of this embodiment, by performing semantic unit segmentation, action information matching, and merging of adjacent identical action units on the text to be broadcast, natural body action instructions that match the semantics of the text can be accurately extracted, effectively avoiding frequent switching or redundant triggering of actions, thereby generating a digital human action sequence that is rhythmically coordinated, semantically consistent, and conforms to human expression habits, significantly improving the expressiveness, fluency, and immersion of digital human videos.
[0008] According to at least one embodiment of this disclosure, obtaining a first digital human animation file corresponding to each text block based on the action information corresponding to each text block includes: constructing a mapping table for recording the mapping relationship between the action information and the file name of the first digital human animation file; searching for the file name corresponding to the action information from the mapping table based on the action information corresponding to each text block; and obtaining a first digital human animation file having the file name based on the searched file name.
[0009] According to the technical solution of this embodiment, by constructing a mapping table between action information and the first digital human animation file name, and quickly looking up the corresponding animation file based on the action information of each text block, semantically driven limb movements can be matched efficiently and accurately, avoiding hard coding or repeated generation, and significantly improving the response speed and system maintainability of digital human animation synthesis.
[0010] According to at least one embodiment of this disclosure, before obtaining the first digital human animation file corresponding to each text block, the method further includes: creating and storing the first digital human animation file corresponding to different action information, wherein the first digital human animation file satisfies the following condition: the first frame image in the first digital human animation file satisfies a natural transition condition between the first frame image in the first digital human animation file and the last frame image of any other first digital human animation file.
[0011] According to the technical solution of this embodiment, when splicing animations by text blocks in the future, it is possible to ensure seamless transition between different action segments, avoid sudden changes in posture or action stuttering, thereby significantly improving the smoothness, naturalness and visual coherence of the synthesized digital human video.
[0012] According to at least one embodiment of this disclosure, the natural transition condition includes: the change in joint angle is less than a first threshold, and the displacement of the joint horizontal displacement is less than a second threshold.
[0013] According to the technical solution of this embodiment, the range of posture changes during splicing of different digital human animation segments can be effectively constrained, avoiding abrupt limb shaking, jumping or uncoordinated movements, thereby ensuring that the overall action sequence conforms to the laws of human movement.
[0014] According to at least one embodiment of this disclosure, before splicing a general transition digital human animation file at the end position of the first digital human animation file, the method further includes: creating the general transition digital human animation file, wherein the general transition digital human animation file satisfies the following: the action information corresponding to the general transition digital human animation file is standing; any two frames in the general transition digital human animation file satisfy the natural transition condition; and any frame in the general transition digital human animation file satisfies the natural transition condition with the first frame and the last frame in any other first digital human animation file.
[0015] According to the technical solution of this embodiment, a universal transition segment can be seamlessly spliced after any main animation ends, effectively eliminating sudden changes or discontinuities in posture during action switching.
[0016] According to at least one embodiment of this disclosure, based on the first audio data corresponding to each text block, lip-syncing is performed on the first digital human animation file corresponding to each text block to obtain a second digital human animation file corresponding to each text block, including: determining whether the duration of the first audio data corresponding to each text block is consistent with the duration of the first digital human animation file; if the duration of the first audio data is less than the duration of the first digital human animation file, truncating the first digital human animation file and using the truncated result as a third digital human animation file corresponding to each text block; and based on the first audio data corresponding to each text block, lip-syncing is performed on the third digital human animation file corresponding to each text block to obtain a second digital human animation file corresponding to each text block.
[0017] According to the technical solution of this embodiment, it is possible to ensure that lip movements are strictly synchronized with speech content, avoiding the phenomenon of silent lip movements or movement trailing caused by excessively long animations, thereby significantly improving the audio-visual consistency, expression accuracy and audio-visual naturalness of digital human broadcasting.
[0018] According to at least one embodiment of this disclosure, the digital human video is generated by synthesizing the first audio data corresponding to each text block and the second digital human animation file. This includes: splicing the first audio data corresponding to each text block according to the order of each text block in the text to be played to obtain second audio data; splicing the second digital human animation file corresponding to each text block according to the order of each text block in the text to be played to obtain a fourth digital human animation file; and synthesizing the second audio data and the fourth digital human animation file to obtain the digital human video.
[0019] According to the technical solution of this embodiment, it is possible to ensure that the entire digital human video is highly consistent in terms of semantics, speech and action, effectively avoiding problems such as audio-visual misalignment, action jumps or semantic breaks, thereby generating high-quality digital human videos with natural rhythm, smooth expression and strong immersion.
[0020] According to at least one embodiment of this disclosure, synthesizing the second audio data and the fourth digital human animation file to obtain a digital human video includes: determining whether the splicing between two adjacent second digital human animation files constituting the fourth digital human animation file satisfies a natural transition condition; if the natural transition condition is not satisfied, adding transition effects to the frame images before and after the splicing position to obtain a fifth digital human animation file; and synthesizing the second audio data and the fifth digital human animation file to obtain a digital human video.
[0021] According to the technical solution of this embodiment, the problem of sudden action changes or inconsistent postures can be effectively masked, and the audience can avoid perceiving abrupt transitions. Thus, while maintaining semantic and audio synchronization, the overall smoothness, visual naturalness and professional presentation of digital human videos can be further improved.
[0022] According to at least one embodiment of this disclosure, obtaining first audio data corresponding to each text block based on the content of each text block includes: inputting the content of each text block into the CosyVoice model to obtain the first audio data corresponding to each text block.
[0023] According to the technical solution of this embodiment, natural, fluent, expressive speech that is highly matched with the semantics of the text can be synthesized efficiently, significantly improving the speech quality, emotional fit and auditory realism of digital human broadcasts, and providing a high-precision audio foundation for subsequent audio-visual synchronization and motion-driven processes.
[0024] According to at least one embodiment of this disclosure, based on the first audio data corresponding to each text block, lip-syncing is performed on the first digital human animation file corresponding to each text block to obtain the second digital human animation file corresponding to each text block, including: simultaneously inputting the first audio data corresponding to the text block and the first digital human animation file into a lip-syncing timing generation model constructed based on the MuseTalk model, and using the output of the lip-syncing timing generation model as the second digital human animation file corresponding to the text block.
[0025] According to the technical solution of this embodiment, it is possible to accurately drive the key points of digital human face, generate lip animation that is highly aligned with speech, consistent in timing and natural in detail, effectively improving the accuracy of lip synchronization and the realism of facial expression.
[0026] According to at least one embodiment of this disclosure, the digital human video is generated by synthesizing the first audio data and the second digital human animation file corresponding to each text block, including: synthesizing the first audio data and the second digital human animation file corresponding to each text block to obtain a synthesized video; merging the synthesized video with background material to obtain a merged video; and performing visual enhancement processing on the merged video to generate the digital human video.
[0027] According to the technical solution of this embodiment, it is possible to effectively achieve precise synchronization of audio and video, natural fusion of foreground and scene, and optimization of picture quality, thereby generating digital human videos with high realism, strong immersion and professional visual expression.
[0028] According to another aspect of this disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, such that the processor performs a digital human video generation method according to any embodiment of this disclosure.
[0029] According to another aspect of this disclosure, a readable storage medium is provided, wherein executable instructions are stored therein, which, when executed by a processor, are used to implement the digital human video generation method of any embodiment of this disclosure.
[0030] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a digital human video generation method according to any embodiment of this disclosure.
[0031] The above-mentioned technical solution not only achieves accurate matching between text content and digital human movements, but also effectively avoids lip misalignment or screen stuttering caused by premature animation termination, thereby improving the smoothness, realism and user experience of digital human broadcasting. Attached Figure Description
[0032] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0033] Figure 1 This is a flowchart illustrating a digital human video generation method according to one embodiment of the present disclosure.
[0034] Figure 2 This is a flowchart illustrating step S160 according to one embodiment of the present disclosure.
[0035] Figure 3 This is a flowchart illustrating step S120 according to one embodiment of the present disclosure.
[0036] Figure 4 This is a flowchart illustrating step S160 according to yet another embodiment of the present disclosure.
[0037] Figure 5 This is a flowchart illustrating step S160 according to another embodiment of the present disclosure.
[0038] Figure 6 This is a schematic block diagram of a digital human video generation apparatus according to one embodiment of the present disclosure.
[0039] Figure 7 This is a schematic structural block diagram of an electronic device employing a processor-based hardware implementation according to one embodiment of the present disclosure. Detailed Implementation
[0040] The present disclosure will now be described in further detail with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.
[0041] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0042] Existing methods for creating digital human videos generally suffer from technical problems such as high cost, low efficiency, and unnatural synthesis effects. Although there are currently digital human video generation solutions based on speech-driven lip-syncing, which generate speech through text-to-speech (TTS) technology and drive the digital human's lip movements to match the speech using lip-sync algorithms, these solutions generally have the following drawbacks: First, they lack sufficient control over complete action sequences, mostly focusing only on lip movements and failing to generate coherent body movements (such as walking or gesture demonstrations). Second, they lack intelligent processing of action transitions; when the action sequence and speech duration do not match (e.g., the speech has not ended but the action has finished playing), the overall continuity of the performance cannot be guaranteed.
[0043] To address this issue, this disclosure proposes the following technical solution. In this solution, the text to be played is received, and the action information contained within it is identified. Based on this, the text is divided into blocks, with each text block corresponding to specific action information. Then, a first digital human animation file matching the action information and corresponding first audio data are obtained for each text block. Next, lip-syncing technology is used to obtain a second digital human animation file corresponding to each text block, ensuring precise alignment of the digital human's lip movements in the second digital human animation file with the first audio data on the timeline. To address the issue of inconsistent durations between the first audio data and the first digital human animation file, especially when the duration of the first audio data is longer than that of the first digital human animation file, a general transitional digital human animation file is automatically spliced to ensure action continuity, ultimately synthesizing a complete digital human video. This method not only achieves precise matching between text content and digital human actions but also effectively avoids lip misalignment or screen stuttering caused by premature animation termination, thereby improving the smoothness, realism, and user experience of digital human playback.
[0044] To facilitate description and make the technical solutions of this disclosure easier to understand, the terminology of this disclosure will be explained before describing the technical solutions of this disclosure.
[0045] The text to be read refers to the original text content that needs to be read aloud and acted out by a digital human.
[0046] Lip-syncing refers to the technical process of automatically controlling the movement of the lips, jaw, facial muscles, and other areas of a digital human or virtual character based on the input speech signal (first audio data), so that the changes in mouth shape are highly consistent with the speech content in terms of time and pronunciation posture.
[0047] This disclosure can be applied to the following scenarios: (1) Real-time virtual anchor broadcasting: The news text is used as the text to be broadcast, and the corresponding unique action information is "standing broadcasting". It can automatically generate virtual anchor videos with lip-sync and natural movements.
[0048] (2) Online education course recording: The teaching text is used as the text to be broadcast, and the corresponding unique action information is "gesture demonstration". It can quickly generate highly interactive teaching videos.
[0049] (3) Short video content creation: The text of the short video content is used as the text to be broadcast, corresponding to multiple action information. Marketing and promotion short videos can be automatically produced.
[0050] (4) Film and television script preview: The script lines are used as the text to be broadcast, corresponding to multiple action information. Visual preview videos can be generated in advance to assist the director's decision-making.
[0051] (5) Metaverse Virtual Interaction: The user's voice dialogue is used as the text to be played, corresponding to multiple action information. Actions and expressions that match the user's voice dialogue can be generated for the virtual character, enhancing the realism of the interaction.
[0052] Figure 1 A schematic diagram illustrating the overall flow of a digital human video generation method according to one embodiment of this disclosure is shown. Figure 1 The method shown includes steps S110 to S170.
[0053] In step S110, the text to be broadcast is received.
[0054] The text to be broadcast is the original text content that needs to be read aloud and acted out by the digital human. Its content format can be: press releases, product introductions, customer service replies, teaching explanations, notices and announcements, etc., any text that needs to be "spoken" by the digital human.
[0055] In step S120, action information recognition is performed on the text to be broadcast to obtain the action information contained in the text to be broadcast.
[0056] As one possible implementation, in recognizing action information from the text to be broadcast, Natural Language Understanding (NLU) technology can be used to automatically extract implicit or explicit semantic information about human actions (i.e., action information) from the text. This determines the corresponding body movements, gestures, facial expressions, or postures that the digital human should perform when broadcasting the text. For example, explicit action words include verbs that directly describe actions, such as "wave," "nod," "point," and "clap." Implicit semantic actions include "welcome" which might correspond to a "smile + wave," and "emphasize" which might correspond to an "increased gesture" or "leaning forward." Emotional and tone cues include "very happy" which might correspond to an "upward gesture," and "warn" which might correspond to a "stopping gesture." For example, if the text to be broadcast is "Please look this way, I will point to the data chart on the screen.", the recognized action information is a "pointing gesture." If the text to be broadcast is "Welcome everyone to our new product launch!", the recognized action information is a "smile + wave."
[0057] In step S130, the text to be broadcast is divided into blocks based on the action information to obtain multiple text blocks, where each text block corresponds to one action information and adjacent text blocks correspond to different action information.
[0058] The purpose of segmentation is to rationally divide the text to be broadcast according to the semantic boundaries of actions, so that each text segment (i.e., "text block") can be matched with a coherent and single action during broadcast, thereby ensuring natural action switching and synchronization between voice and action in the digital human. As one possible implementation, to ensure that each text block corresponds to one action information, the action information is used as a label for each text block. For example, the text to be broadcast is: "Hello everyone! (waves) Today I will introduce our new product. (explains gesture) It has three core functions." Through action information recognition, it is found that the text to be broadcast contains three action information: "Hello everyone!" corresponds to the action information "waves"; "Today I will introduce our new product." corresponds to the action information "explains gesture"; and "It has three core functions." corresponds to the action information "default posture." Therefore, the text to be broadcast is divided into three text blocks, corresponding to the action information "waves," "explains gesture," and "default posture," respectively.
[0059] In step S140, based on the action information corresponding to each text block, the first digital human animation file corresponding to each text block is obtained, wherein the digital human in the first digital human animation file has actions corresponding to the action information.
[0060] For example, the first digital human animation file can be a digital human motion GIF (Graphics InterchangeFormat) file, which has a standardized sequence of motions (such as standing and explaining, gesture demonstration, etc.).
[0061] As one possible implementation, based on the action information corresponding to each text block, the first digital human animation file corresponding to each text block is obtained, including: constructing a mapping table to record the mapping relationship between action information and the filenames of the first digital human animation files; searching for the filename corresponding to the action information in the mapping table based on the action information corresponding to each text block; and obtaining the first digital human animation file with the found filename based on the found filename. By constructing a mapping table, the corresponding animation file can be quickly found and obtained, significantly improving the efficiency and accuracy of digital human animation matching.
[0062] As another possible implementation, the motion information is used as the filename of the first digital human animation file. After storing all the first digital human animation files in the same database, when it is necessary to retrieve the first digital human animation file corresponding to each text block, the filename corresponding to the motion information is used to find the corresponding first digital human animation file.
[0063] As a further implementation, before obtaining the first digital human animation file corresponding to each text block, the method further includes: creating and storing first digital human animation files corresponding to different motion information. The first digital human animation file satisfies the following condition: the first frame image in the first digital human animation file meets a natural transition condition with the last frame image of any other first digital human animation file. When creating the first digital human animation file, the user can directly log in and create it, or the user can pre-create it and then upload it. By setting a natural transition condition, it can be ensured that the digital human's posture, position, and movements can be smoothly connected when switching between different first digital human animation files, avoiding abrupt jumps or visual breaks.
[0064] In one example, the natural transition conditions include: the change in joint angle is less than a first threshold, and the displacement of the joint's horizontal displacement is less than a second threshold. The joint can be any joint related to movement, such as the shoulder, elbow, or hip. The values of the first and second thresholds can be determined based on the actual application and are not limited here. For example, when the joint is the shoulder joint, the change in joint angle is the difference in shoulder joint rotation angle, and the corresponding first threshold is 10° per frame. When the joint is the hip joint, the displacement of the joint's horizontal displacement is the same as the displacement of the hip joint's horizontal displacement, and the corresponding second threshold is 5 pixels per frame. It should be noted that the joint in the first condition "change in joint angle" and the joint in the second condition "displacement of joint's horizontal displacement" can be the same or different. Each condition can have one or more joints. When there are multiple joints, all joints must meet the corresponding conditions. By limiting the range of change in joint angle and horizontal displacement between adjacent frames, abrupt jumps, jerks, or disjointed limb movements in digital human actions can be effectively avoided, resulting in a smooth overall posture transition that conforms to the laws of human movement and significantly enhances visual realism.
[0065] In step S150, the first audio data corresponding to each text block is obtained based on the content of each text block.
[0066] As one possible implementation, based on the content of each text block, the first audio data corresponding to each text block is obtained, including: inputting the content of each text block into the CosyVoice model to obtain the first audio data corresponding to each text block. The CosyVoice model has a sampling rate of 16kHz, supports Mandarin Chinese and emotional tone adjustment, and the generated first audio data completely matches the semantics of the text block content.
[0067] In step S160, based on the first audio data corresponding to each text block, lip-syncing is performed on the first digital human animation file corresponding to each text block to obtain the second digital human animation file corresponding to each text block, wherein the lip shape of the digital human in the second digital human animation file is aligned with the first audio data on the time axis.
[0068] Lip-tracking refers to the process of automatically generating lip movements for a digital human that precisely match the pronunciation, rhythm, and timing of an audio recording. The goal is to make the opening and closing of the virtual digital human's mouth and the changes in lip shape appear natural and realistic when "speaking," and to be completely synchronized with the spoken sound. Lip-tracking allows the digital human's lip movements to be highly consistent with the spoken language in terms of time, phonemes, and rhythm, thereby enhancing the realism, credibility, and user experience of virtual speech. Without lip-tracking, even with speech and body movements, a digital human will appear mechanical, artificial, and even uncomfortable due to "no moving mouth" or "misalignment between lip shape and speech."
[0069] As one possible implementation, based on the first audio data corresponding to each text block, lip-syncing is applied to the first digital human animation file corresponding to each text block to obtain a second digital human animation file corresponding to each text block. This includes: simultaneously inputting the first audio data and the first digital human animation file corresponding to the text block into a lip-syncing temporal generation model built based on the MuseTalk model, and using the output of the lip-syncing temporal generation model as the second digital human animation file corresponding to the text block. The MuseTalk model uses the speech audio signal corresponding to the first audio data and the action sequence frames corresponding to the first digital human animation file as joint inputs, and generates a digital human video sequence (i.e., the second digital human animation file) whose lip movement trajectory is strictly synchronized with the speech content through a temporal alignment mechanism. In this digital human video sequence, the lip shape changes of the digital human precisely correspond to the pronunciation time and phoneme content of the speech. In other implementations, other models can be used, such as models built based on temporal convolutional networks or generative adversarial networks.
[0070] Regarding step S160, in some embodiments of this disclosure, it may include, for example... Figure 2 Steps S1611 to S1613 are shown.
[0071] In step S1611, it is determined whether the first audio data corresponding to each text block is consistent with the time length of the first digital human animation file.
[0072] The duration of the first audio data is determined by the text content, speaking speed, pauses, etc. The first digital human animation file is a pre-configured file (e.g., uploaded by the user), and its duration is determined by the length of a preset animation segment. For example, the duration of the first digital human animation file corresponding to "waving" is fixed at 1.5 seconds. By comparing the durations of the first audio data and the first digital human animation file, it can be determined whether they are consistent.
[0073] In step S1612, if the duration of the first audio data is longer than the duration of the first digital human animation file, a general transition digital human animation file is spliced at the end of the first digital human animation file, and the splicing result is used as the third digital human animation file corresponding to each text block.
[0074] A universal transitional digital human animation file refers to a pre-made, loopable, and splicable basic animation segment for a digital human. It is used to seamlessly connect to and continue the natural posture of the digital human after the main motion animation (i.e., the first digital human animation file) ends, filling the "window period" caused by the speech duration exceeding the motion duration. Universal transitional digital human animation files can be manually created (based on motion combination analysis and design), automatically generated by algorithms (through motion trajectory or posture clustering), or a combination of both.
[0075] In one possible implementation, a generic transitional digital human animation file is spliced at the end of the first digital human animation file. First, the target duration of the generic transitional digital human animation file to be spliced is determined based on the difference in duration between the first audio data and the first digital human animation file. Then, the generic transitional digital human animation file is truncated and / or spliced based on this target duration to obtain the generic transitional digital human animation file to be spliced. Finally, this generic transitional digital human animation file to be spliced is spliced to the end of the first digital human animation file. In one example, the duration of the first audio data is 8.5 seconds, and the duration of the first digital human animation file is 8 seconds, so the target duration is 8.5 - 8 = 0.5 seconds. The generic transitional digital human animation file is 3 seconds long, so a 0.5-second segment is randomly extracted from the generic transitional digital human animation file as the generic transitional digital human animation file to be spliced. In another example, the duration of the first audio data is 8.5 seconds, and the duration of the first digital human animation file is 5 seconds, so the target duration is 8.5 - 5 = 3.5 seconds. If the generic transition digital human animation file is 3 seconds long, then the two generic transition digital human animation files are spliced together, and then a 3.5-second segment is randomly extracted from the spliced result as the generic transition digital human animation file to be spliced.
[0076] As a further implementation, before splicing the universal transition digital human animation file at the end of the first digital human animation file, the method further includes: creating a universal transition digital human animation file. In one example, the universal transition digital human animation file satisfies the following conditions: the motion information corresponding to the universal transition digital human animation file is standing (ensuring the neutral and stable posture of the digital human); any two frames in the universal transition digital human animation file satisfy the natural transition condition; and any frame in the universal transition digital human animation file satisfies the natural transition condition with the first and last frames of any other first digital human animation file (ensuring high compatibility of the universal transition digital human animation file). Considering that "standing" is the most basic and stable default posture of the human body and does not carry specific semantics, setting the motion information corresponding to the universal transition digital human animation file to standing can be applied to most scenarios and will not interfere with the understanding of the motion information corresponding to the first digital human animation files that are connected before and after. Considering that the universal transition digital human animation file needs to be dynamically cropped to any duration (such as padding 0.7 seconds, 1.3 seconds, 2.1 seconds, etc.) when used, if there are jumps, stutters, or discontinuous movements within the animation, unnatural segments will be exposed after cropping. Therefore, by setting a natural transition condition between any two frames within a universal transitional digital human animation file, arbitrary duration cropping can be supported (i.e., regardless of how much time is added, the cropped segments will look natural). Furthermore, by setting a natural transition condition between any frame within a universal transitional digital human animation file and both the first and last frames of the first digital human animation file, any cropped portion of the universal transitional digital human animation file can be seamlessly stitched with the first digital human animation file.
[0077] In step S1613, based on the first audio data corresponding to each text block, lip-syncing is performed on the third digital human animation file corresponding to each text block to obtain the second digital human animation file corresponding to each text block.
[0078] Through steps S1611 to S1613, it is ensured that the final generated digital human video fully covers the entire speech content in the time dimension, thereby achieving full lip-sync from the main action stage to the natural transition stage during lip-syncing. This implementation not only avoids the problem of lip-syncing interruption or body stiffness caused by the premature end of the first digital human animation file, but also ensures that the digital human's posture is consistent, lip movements are accurate, and performance is natural throughout the entire broadcast, significantly improving the realism, smoothness, and professionalism of the generated video.
[0079] In step S170, the first audio data and the second digital human animation file corresponding to each text block are combined to generate a digital human video.
[0080] The goal of the synthesis is to seamlessly stitch, render, and encapsulate aligned speech (first audio data) with animations featuring synchronized lip movements and body gestures (second digital human animation file) in chronological order into a complete, playable digital human video.
[0081] As one possible implementation, a digital human video is generated by synthesizing the first audio data and the second digital human animation file corresponding to each text block. This includes: concatenating the first audio data corresponding to each text block based on their order in the text to be played, to obtain the second audio data; concatenating the second digital human animation file corresponding to each text block based on their order in the text to be played, to obtain the fourth digital human animation file; and synthesizing the second audio data and the fourth digital human animation file to obtain the digital human video. This implementation ensures that the speech content and the digital human's actions are strictly consistent and seamlessly connected in overall timing, effectively avoiding problems such as audio-visual asynchrony, action jumps, or semantic breaks, thereby significantly improving the fluency and naturalness of the digital human's broadcast.
[0082] In one example, the second audio data is combined with the fourth digital human animation file to obtain a digital human video. This includes determining whether the splicing between two adjacent second digital human animation files constituting the fourth digital human animation file satisfies the natural transition condition. If the natural transition condition is not met, transition effects are added to the frames before and after the splicing position to obtain a fifth digital human animation file. The second audio data and the fifth digital human animation file are then combined to obtain the digital human video. This example effectively eliminates visual jumps caused by abrupt changes in movement or inconsistent postures, significantly improving the smoothness and naturalness of animation transitions.
[0083] For example, when determining whether the splicing between two adjacent second digital human animation files constituting the fourth digital human animation file satisfies the natural transition condition, it can be determined whether the last frame of the preceding second digital human animation file and the first frame of the following second digital human animation file satisfy the natural transition condition. Transition effects can use fade-in / fade-out, camera panning, etc. When using fade-in / fade-out effects, the last n frames of the preceding second digital human animation file are selected to add the fade-out effect, and the first n frames of the following second digital human animation file are selected to add the fade-in effect. The value of n is determined according to the transition duration.
[0084] As a further implementation, after obtaining the synthesized video, the process further includes fusing the synthesized video with background material to obtain a merged video. Visual enhancement processing is then applied to the merged video to generate a digital human video. For example, in fusing the synthesized video with background material, the foreground region of the digital human (preserving alpha channel transparency information) is obtained from the synthesized video, and image fusion algorithms such as alpha blending are used to seamlessly embed the obtained foreground region of the digital human into the target background material. The background material can be a static image (such as a conference room background, classroom scene, virtual stage, etc.) or a dynamic video matching the digital human's movements, a virtual 3D scene, etc. The visual enhancement processing can employ one or more processing methods such as color correction (such as white balance adjustment) and sharpening (enhancing edge clarity). Through this implementation, the digital human video not only appears realistic but also naturally blends into the target scene corresponding to the background material, achieving a visually indistinguishable effect.
[0085] Through steps S110 to S170, end-to-end automated digital human video generation from text input to digital human video output is achieved. Only the text to be read and background materials need to be input, eliminating the need for manual intervention in motion editing or video editing processes, significantly reducing the technical threshold and labor costs of digital human video production. This disclosure ensures that the digital human's movements are fluid throughout the entire reading process, with precise synchronization between lip movements and speech, effectively avoiding issues such as body stiffness or lip interruption caused by premature animation termination, and significantly improving the naturalness, smoothness, and visual realism of the generated digital human video.
[0086] Regarding step S120, in some embodiments of this disclosure, it may include, for example... Figure 3 Steps S1201 to S1204 are shown.
[0087] In step S1201, the text to be broadcast is segmented into semantic units to obtain multiple semantic units.
[0088] Semantic unit segmentation refers to dividing continuous text (the text to be broadcast) into several logical segments (semantic units) according to semantic integrity and rhythm. Segmentation criteria can include punctuation marks (periods, commas, question marks, exclamation marks, etc.), logical connectors ("firstly," "then," "but," "in addition," etc.), and colloquial pauses (such as "um…," "ah…," etc.). Semantic unit segmentation can be performed using NLP models (such as BERT-based semantic boundary detection).
[0089] In step S1202, action information matching is performed based on semantic units to obtain the action information corresponding to each semantic unit.
[0090] Action information matching refers to identifying the implicit or explicit action intention of each semantic unit, such as "welcome" corresponding to waving, and "introduce" corresponding to explanatory gestures.
[0091] In step S1203, adjacent semantic units with the same action information are merged.
[0092] By merging adjacent semantic units with the same action, we can avoid the same action being frequently started and stopped (such as playing the animation three times for three consecutive "pointing", which seems mechanical) and optimize resource scheduling.
[0093] In step S1204, the merged and retained action information is used as the action information contained in the text to be broadcast.
[0094] The merged action information is a sequence of action information after removing redundancy. For example, it can be represented by an action information list, where each row in the action information list corresponds to one action information, and the row number corresponds to the order in which the action information appears in the text to be broadcast.
[0095] Steps S1201 to S1204 not only improve the semantic accuracy of action information recognition, but also effectively avoid the mechanical feel and incoherence caused by frequent switching or fragmentation of actions in digital humans. This ensures that the actions of the digital human videos generated subsequently are natural and smooth, and that the semantics and behavior are highly consistent, significantly enhancing the realism and professional performance of the broadcast.
[0096] Regarding step S160, in some embodiments of this disclosure, it may include, for example... Figure 4 Steps S1621 to S1623 are shown. Step S1621 corresponds to... Figure 2 Steps S1611 and S1623 of the embodiment correspond to Figure 2 Step S1613 of the embodiment is detailed below. Figure 2 The relevant descriptions of the embodiments are not repeated here.
[0097] In step S1622, if the duration of the first audio data is less than the duration of the first digital human animation file, the first digital human animation file is truncated, and the truncated result is used as the third digital human animation file corresponding to each text block.
[0098] As one possible implementation, when truncating the first digital human animation file, the starting frame of the first digital human animation file is used as the starting point for truncating to ensure the integrity of the action semantics.
[0099] Steps S1621 to S1623 can effectively achieve audio-visual synchronization, avoiding the problem of lip movements or actions not matching the voice caused by the animation playback time exceeding the voice content.
[0100] Regarding step S160, in some embodiments of this disclosure, it may include, for example... Figure 5 Steps S1631 to S1633 are shown. Step S1631 corresponds to... Figure 2 Step S1611 and step S1633 of the embodiment correspond to Figure 2 Step S1613 of the embodiment is detailed below. Figure 2 The relevant descriptions of the embodiments are not repeated here.
[0101] In step S1632, if the duration of the first audio data is equal to the duration of the first digital human animation file, the first digital human animation file is used as the third digital human animation file.
[0102] From steps S1631 to S1633, since the first audio data and the first digital human animation file are naturally aligned in time, there is no need for truncation or stretching, which can ensure strict synchronization of lip movements, facial expressions, body movements and speech content.
[0103] According to any of the above embodiments, this disclosure also provides a digital human video generation apparatus 200. Figure 6 This is a schematic block diagram of a digital human video generation apparatus 200 according to one embodiment of this disclosure. Figure 6 As shown, the digital human video generation device 200 includes a text receiving module 210, an action information recognition module 220, a block processing module 230, an animation file acquisition module 240, an audio data generation module 250, a lip-syncing module 260, and a video synthesis module 270. The text receiving module 210 receives text to be played. The action information recognition module 220 performs action information recognition on the text to be played to obtain the action information contained in the text. The block processing module 230 performs block processing on the text to be played based on the action information to obtain multiple text blocks, where each text block corresponds to one action information, and adjacent text blocks correspond to different action information. The animation file acquisition module 240 acquires a first digital human animation file corresponding to each text block based on the action information corresponding to each text block, wherein the digital human in the first digital human animation file has the action corresponding to the action information. The audio data generation module 250 obtains first audio data corresponding to each text block based on the content of each text block. The lip-syncing module 260 is used to perform lip-syncing on the first digital human animation file corresponding to each text block based on the first audio data corresponding to each text block, to obtain the second digital human animation file corresponding to each text block. In the second digital human animation file, the lip shape of the digital human is aligned with the first audio data on the timeline. The video compositing module 270 is used to composite the first audio data and the second digital human animation file corresponding to each text block to generate a digital human video.
[0104] According to further embodiments of this disclosure, an electronic device is also provided. Figure 7 This diagram illustrates a schematic block diagram of an electronic device employing a processor-based hardware implementation according to an embodiment of the present disclosure. The hardware structure of the electronic device of the present disclosure can be implemented using a bus architecture. The bus architecture can include any number of interconnect buses and bridges, depending on the specific application and overall design constraints of the hardware. Bus 1100 connects various circuits including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc. Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one connecting line is used in this figure, but this does not imply that there is only one bus or one type of bus. Memory 1300 stores a computer program, and when processor 1200 executes the computer program, processor 1200 is able to perform the following processes. Receive the text to be played. Recognize the action information within the text to obtain the action information contained within it. Divide the text into blocks based on the action information, resulting in multiple text blocks, each corresponding to one piece of action information, with adjacent text blocks corresponding to different action information. Based on the action information corresponding to each text block, obtain the first digital human animation file corresponding to each text block, where the digital human in the first digital human animation file performs the actions corresponding to the action information. Based on the content of each text block, obtain the first audio data corresponding to each text block. Based on the first audio data corresponding to each text block, perform lip-syncing on the first digital human animation file corresponding to each text block to obtain the second digital human animation file corresponding to each text block, where the lip movements of the digital human in the second digital human animation file are aligned with the first audio data on the timeline. Combine the first audio data and the second digital human animation file corresponding to each text block to generate a digital human video.
[0105] This disclosure also provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of a readable storage medium include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.
[0106] This disclosure also provides a computer program product, the methods of which can be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, all or part of the processes or functions of this disclosure are performed.
[0107] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any available medium capable of access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or it can include both volatile and non-volatile types of storage media.
[0108] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0109] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0110] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0111] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0112] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment / mode or example, which are included in at least one embodiment / mode or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.
[0113] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.
Claims
1. A digital human video generation method, characterized by, The method comprises: receiving a text to be broadcast; performing action information identification on the text to be broadcast to obtain action information contained in the text to be broadcast; performing block processing on the text to be broadcast based on the action information to obtain a plurality of text blocks, wherein each text block corresponds to an action information, and adjacent text blocks correspond to different action information; obtaining a first digital human animation file corresponding to each text block based on the action information corresponding to each text block, wherein a digital human in the first digital human animation file has an action corresponding to the action information; obtaining first audio data corresponding to each text block based on the content of each text block; performing lip driving on the first digital human animation file corresponding to each text block based on the first audio data corresponding to each text block to obtain a second digital human animation file corresponding to each text block, wherein the lip shape of the digital human in the second digital human animation file is aligned with the first audio data on a time axis; and synthesizing the first audio data and the second digital human animation file corresponding to each text block to generate a digital human video. In the lip driving on the first digital human animation file corresponding to each text block based on the first audio data corresponding to each text block to obtain the second digital human animation file corresponding to each text block, it is determined whether the time length of the first audio data corresponding to each text block is consistent with the time length of the first digital human animation file; in the case where the time length of the first audio data is greater than the time length of the first digital human animation file, a general transition digital human animation file is spliced at the end position of the first digital human animation file, and the spliced result is taken as a third digital human animation file corresponding to each text block; the third digital human animation file corresponding to each text block is driven based on the first audio data corresponding to each text block to obtain the second digital human animation file corresponding to each text block.
2. The method of claim 1, wherein, Before obtaining the first digital human animation file corresponding to each text block, the method further comprises: creating and storing the first digital human animation file corresponding to different action information, wherein the first digital human animation file satisfies that the first frame image in the first digital human animation file and the last frame image of any other first digital human animation file satisfy a natural transition condition.
3. The method of claim 2, wherein, The natural transition condition comprises that the change amount of joint angle is less than a first threshold value, and the displacement amount of joint horizontal displacement is less than a second threshold value.
4. The method of claim 1, wherein, Before splicing the general transition digital human animation file at the end position of the first digital human animation file, the method further comprises: creating the general transition digital human animation file, wherein the general transition digital human animation file satisfies that the action information corresponding to the general transition digital human animation file is standing, any two frame images in the general transition digital human animation file satisfy the natural transition condition, and any frame image in the general transition digital human animation file and the first frame image and the last frame image in any other first digital human animation file satisfy the natural transition condition.
5. The method of claim 1, wherein, The first audio data corresponding to each text block is used to drive the first digital human animation file corresponding to each text block based on the first audio data, to obtain a second digital human animation file corresponding to each text block, including: It is judged whether the time length of the first audio data corresponding to each text block is consistent with the time length of the first digital human animation file; In the case where the time length of the first audio data is less than the time length of the first digital human animation file, the first digital human animation file is intercepted, and the interception result is taken as a third digital human animation file corresponding to each text block; and The first audio data corresponding to each text block is used to drive the third digital human animation file corresponding to each text block based on the first audio data, to obtain a second digital human animation file corresponding to each text block.
6. The method of claim 1, wherein, The first audio data corresponding to each text block and the second digital human animation file are synthesized to generate the digital human video, including: The first audio data corresponding to each text block is spliced based on the order of each text block in the to-be-broadcast text to obtain second audio data; The second digital human animation file corresponding to each text block is spliced based on the order of each text block in the to-be-broadcast text to obtain a fourth digital human animation file; and The second audio data and the fourth digital human animation file are synthesized to obtain the digital human video.
7. The method of claim 6, wherein, The second audio data and the fourth digital human animation file are synthesized to obtain a digital human video, including: It is judged whether the splicing between the adjacent two second digital human animation files constituting the fourth digital human animation file meets a natural transition condition; In the case where the natural transition condition is not met, transition special effects are added to the frame images located before and after the splicing position to obtain a fifth digital human animation file; and The second audio data and the fifth digital human animation file are synthesized to obtain a digital human video.
8. The method of claim 1, wherein, The first audio data corresponding to each text block and the second digital human animation file are synthesized to generate the digital human video, including: The first audio data corresponding to each text block and the second digital human animation file are synthesized to obtain a synthesized video; The synthesized video is fused with background material to obtain a fused video; and The fused video is subjected to visual enhancement processing to generate the digital human video.
9. An electronic device, comprising: including: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the digital human video generation method in any one of claims 1 to 8.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the digital human video generation method in any one of claims 1 to 8. The computer program is executed by the processor to implement the digital human video generation method in any one of claims 1 to 8.
Citation Information
Patent Citations
Realistic virtual human generation method and device based on text driving
CN115984429A
Method and system for driving digital human to speak through audio, and storage medium
CN117152319A
Video generation and live broadcast method, medium, device and computing equipment
CN118075548A
Digital human material generation and sharing method and device based on intelligent hardware
CN120640079A
Method and system for text-based avatar generation
WO2023096275A1