Digital human audio-video generation method and device, program product and electronic device
Patent Information
- Application Number
- CN202611014078.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0010]本申请提供了一种数字人音视频的生成方法、装置、程序产品及电子设备,以至少解决基于现有技术生成的数字人音视频的效果差的技术问题
[0104](1)建立精确的时间映射关系:通过检测每个音频片段在完整音频中的起始时间和结束时间,生成系统能够精确掌握每个文案片段对应的语音在时间轴上的具体位置。这种精确的时间映射解决了长文案场景下,如何准确对应“文案-语音-动作”三者时间关系的技术难题,为后续的视频片段生成和时长对齐提供了可靠的时间基准。
Smart Images

Figure CN122845893A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital human video generation technology, and more specifically, to a method, apparatus, program product, and electronic device for generating digital human audio and video. Background Technology
[0002] With the development of speech synthesis, image generation, and video generation technologies, digital humans have been widely applied in scenarios such as AI (Artificial Intelligence) instructors, corporate training, government publicity, virtual anchors, and intelligent customer service. Existing digital human generation technologies mainly include solutions such as image-based digital humans, video lip-syncing, motion-driven technologies, head motion transfer, and large-model image-to-video generation.
[0003] Image-based digital human solutions typically take a picture of a person and an audio clip as input, generating a speaking video from the static image. This approach generates talking human-like videos based on a reference image and audio, focusing on improving lip-sync, facial expressions, head movements, and stability in long videos. However, this type of method primarily works with the head, mouth, eyes, and facial expressions, typically not supporting upper body, arm, and hand gestures. This makes it difficult to meet the needs of AI instructors in scenarios requiring gestures such as welcoming, introducing, emphasizing, and guiding, and its inference speed remains limited.
[0004] Video lip-syncing solutions typically take existing video and target audio as input, redraw the mouth area of the person to match the lip movements with the speech. This solution provides high-quality lip-syncing for existing videos. While this type of method can preserve body movements from the original video, it requires suitable video footage; it cannot generate body movements from a single user image, nor can it solve the problem of adaptive matching of different text clips with action durations.
[0005] Action-driven solutions can take a person image and a reference action video as input, and generate a video from the target image according to the reference action. This solution drives the person's body movements based on a reference pose. This type of method can generate upper body or full-body movements, but the inference process is close to video generation, resulting in high computational costs. If action transfer is performed on each segment in a long text AI lecturer scenario, it will lead to long generation time and high computational consumption.
[0006] The head motion transfer scheme can generate head motion and partial lip-sync effects from a single photo driven by a reference video, and has high inference efficiency. However, its scope of application is mainly concentrated on the head and does not support complete body motion and gestures, which cannot meet the needs of narration-style digital humans for gesture expression.
[0007] Large-scale model-based video generation solutions can generate animated speaking videos based on images, prompts, and audio, demonstrating strong open generation capabilities. However, these methods suffer from high inference costs and slow speeds. Furthermore, since the actions are freely generated by the model, they lack stability and controllability, making it difficult to guarantee that a particular text segment corresponds to a specific action, or to ensure smooth splicing between multiple action segments.
[0008] In summary, existing technologies still have the following shortcomings: image-based digital humans lack body and gesture control; video lip-syncing relies on existing materials; motion-driven processes are costly; head migration does not support full-body movements; and large-model image-to-video generation is slow and the movements are unstable.
[0009] To address the technical problem of poor audio and video quality of digital humans generated using existing technologies, no effective solution has yet been proposed. There is an urgent need for a digital human motion control and video generation method that can combine motion tags, semantic motion matching, built-in motion template invocation, user image motion adaptation, low-cost transition template combination, and smooth splicing of multiple segments. Summary of the Invention
[0010] This application provides a method, apparatus, program product, and electronic device for generating digital human audio and video, to at least solve the technical problem of poor quality of digital human audio and video generated based on existing technology.
[0011] According to one aspect of this application, a method for generating audio and video of a digital human is provided, comprising: generating a sequence of action segments based on target text and action patterns input by a user, wherein each action segment in the sequence of action segments includes at least a segment number, text content, action code, and action name; generating an audio segment corresponding to each action segment based on each action segment in the sequence of action segments, and detecting the audio duration of each audio segment; generating an initial video segment corresponding to each action segment based on a digital human type set by a user, and detecting the video duration of each initial video segment; performing duration alignment processing on the initial video segment corresponding to each action segment based on the audio duration and video duration corresponding to each action segment to obtain an intermediate video segment with the same video duration as the corresponding audio duration; performing lip-sync processing on the intermediate video segment corresponding to each action segment based on the audio segment corresponding to each action segment to obtain a target video segment with lip movements synchronized with the corresponding audio content; and generating a target audio and video file based on the target video segment, audio segment, preset subtitles, and preset background corresponding to each action segment, wherein the target audio and video file is used to express the target text through a digital human speech.
[0012] Optionally, generating an action segment sequence based on the target text and action pattern input by the user includes: detecting the action pattern input by the user, wherein the action pattern is one of the following types: natural action pattern, manually inserted action pattern, and intelligent matching pattern; when the action pattern is natural action pattern / intelligent matching pattern, matching the semantic features of each text segment obtained by dividing the target text in a preset action library to obtain the index information of the action template corresponding to each text segment, wherein the index information of each action template includes at least an action code and an action name; when the action pattern is manually inserted action pattern, extracting pre-inserted action tags from the beginning of each text segment in the target text, and querying the preset action library based on each action tag to obtain the index information of the action template corresponding to each text segment; and generating an action segment sequence based on the text content corresponding to each text segment and the index information of the action template.
[0013] Optionally, an audio segment corresponding to each action segment is generated based on each action segment in the action segment sequence, and the audio duration of each audio segment is detected, including: performing speech synthesis operation based on the text content in each action segment according to the order corresponding to the action segment sequence to obtain the audio segment corresponding to each action segment; detecting the start time and end time of each audio segment in the complete audio corresponding to all action segments; and determining the audio duration of each audio segment based on the start time and end time of each audio segment.
[0014] Optionally, generating an initial video clip corresponding to each action segment based on the user-defined digital human type includes: if the user-defined digital human type is a built-in type, querying the preset action library to obtain the action template set corresponding to the user-selected built-in digital human image; querying the action template set corresponding to the built-in digital human image based on the action code in each action segment to obtain the action template corresponding to each action segment; and generating an initial video clip corresponding to each action segment based on the action template corresponding to each action segment.
[0015] Optionally, generating an initial video clip corresponding to each action segment based on the user-defined digital human type includes: preprocessing the user-uploaded digital human image to obtain a target image that meets the motion transfer conditions, where the motion transfer conditions are used to detect the completeness of the digital human image in the image; extracting auxiliary data corresponding to the target image, where the auxiliary data includes face bounding box position, human mask, human body key points, and facial key points; matching the target image in the built-in digital human image library, and using the matched built-in digital human image as the source image corresponding to the target image; and performing motion transfer based on the auxiliary data corresponding to the target image and the motion template set corresponding to the source image to obtain the initial video clip corresponding to each action segment.
[0016] Optionally, the user-uploaded digital human image is preprocessed to obtain a target image that meets the action transfer conditions. This includes: standardizing the user-uploaded digital human image to obtain a standardized image, wherein the standardization process is used to adjust the size of the digital human image based on coordinate mapping information; if the standardized image includes a single face, the single face is used as the target face; if the standardized image includes two or more faces, the face located in the center region of the standardized image and having the largest area is used as the target face; in the standardized image, the person to which the target face belongs is selected for image cutout and background replacement to obtain an initial image; if the integrity of the digital human image in the initial image does not meet the action transfer conditions, the digital human image in the initial image is updated based on an image completion algorithm to obtain the target image.
[0017] Optionally, matching is performed on the target image in a built-in digital human image library, and the matched built-in digital human image is used as the source image corresponding to the target image. This includes: detecting the gender of the target digital human image in the target image and the gender of each built-in digital human image in the built-in digital human image library; using the set of built-in digital humans with the same gender as the target digital human image as a candidate set; detecting the facial similarity and body shape similarity between each built-in digital human image in the candidate set and the target image; performing a weighted sum of the facial similarity and body shape similarity to obtain the target similarity between each built-in digital human image and the target image; and using the built-in digital human image with the highest target similarity in the candidate set as the source image corresponding to the target image.
[0018] Optionally, motion transfer is performed based on the auxiliary data corresponding to the target image and the motion template set corresponding to the source image to obtain the initial video segment corresponding to each motion segment. This includes: when the auxiliary data of the target image enters the motion transfer module for the first time, feature extraction is performed on the target image through the motion transfer module to obtain feature data, and the feature data and auxiliary data are cached as target features corresponding to the target image in a preset cache layer. The feature data includes target image features, target image encoding features, and latent variable features; based on the motion encoding in each motion segment, a query is performed in the motion template set corresponding to the source image to obtain the motion template corresponding to each motion segment; and the motion template is then retrieved in the preset cache layer. The template features of the action template corresponding to each pre-cached offline action segment are obtained. The template features include action posture sequence, key point data, person mask, action trajectory features, and posture guidance features. Based on the action template corresponding to each action segment, the action transfer strategy corresponding to each action segment is determined. The action template is one of the following types: speech action template, natural action template, or transition action template. The action transfer strategy is one of the following types: face action transfer strategy or human body action transfer strategy. Based on the action transfer strategy, template features, and target features corresponding to the target image, action transfer is performed to obtain the initial video segment corresponding to each action segment.
[0019] Optionally, the action transfer strategy for each action segment is determined based on the action template corresponding to each action segment, including: if the action template corresponding to the i-th action segment is a speech action template, the action transfer strategy for the i-th action segment is determined to be a human action transfer strategy; if the action template corresponding to the i-th action segment is a natural action template / transitional action template, the action transfer strategy for the i-th action segment is determined to be a face action transfer strategy.
[0020] Optionally, motion transfer is performed based on the motion transfer strategy, template features, and target features corresponding to each motion segment to obtain an initial video segment for each motion segment. This includes: when the motion transfer strategy for the i-th motion segment is a human motion transfer strategy, aligning the motion pose sequence in the template features and the human key points in the target features to obtain target pose data, wherein the target pose data is used to characterize the body pose and gesture position of the target digital human image in each frame of the motion pose sequence; generating a video frame sequence corresponding to the i-th motion segment using a preset motion transfer algorithm based on the target pose data and the target features corresponding to the target image; and performing enhancement operations on the video frame sequence corresponding to the i-th motion segment to obtain the initial video segment corresponding to the i-th motion segment.
[0021] Optionally, the initial video segment corresponding to each action segment is time-aligned based on the audio and video durations corresponding to each action segment to obtain an intermediate video segment with the same video duration as the corresponding audio duration. This includes: if the video duration corresponding to the i-th action segment is greater than or equal to the audio duration corresponding to the i-th action segment, the initial video segment corresponding to the i-th action segment is truncated based on the audio duration corresponding to the i-th action segment to obtain a retained video segment and a discarded video segment corresponding to the i-th action segment; if the video duration of the retained video segment is greater than or equal to the video duration of the discarded video segment, the discarded video segment played in the forward direction is used as the stabilized video segment corresponding to the i-th action segment; if the video duration of the retained video segment is less than the video duration of the discarded video segment, the discarded video segment played in the reverse direction is used as the stabilized video segment corresponding to the i-th action segment; the truncated video segment and the stabilized segment corresponding to the i-th action segment are concatenated to obtain an intermediate video segment corresponding to the i-th action segment, and a silent segment is generated based on the video duration of the stabilized video segment, and the silent segment is concatenated to the end of the audio segment corresponding to the i-th action segment.
[0022] Optionally, based on the audio and video durations corresponding to each action segment, the initial video segment corresponding to each action segment is time-aligned to obtain an intermediate video segment with the same video duration as the corresponding audio duration. This includes: if the video duration corresponding to the i-th action segment is less than the audio duration corresponding to the i-th action segment, the difference between the audio duration and video duration corresponding to the i-th action segment is taken as the remaining duration; the remaining duration corresponding to the i-th action segment is rounded up to obtain the target remaining duration corresponding to the i-th action segment; if the user-set digital human type is a built-in type, based on the audio and video durations corresponding to the i-th action segment... The target remaining time is used to retrieve L transitional action templates from the action template set corresponding to the user-selected built-in digital human avatar, where L is a positive integer. If the user-set digital human type is upload type, L transitional action templates are retrieved from the action template set corresponding to the source avatar based on the target remaining time corresponding to the i-th action segment. A transitional video segment is generated based on the L transitional action templates corresponding to the i-th action segment, where the video duration of the transitional video segment is equal to the target remaining time. The initial video segment and the transitional video segment corresponding to the i-th action segment are spliced together to obtain the intermediate video segment corresponding to the i-th action segment.
[0023] According to another aspect of this application, a digital human audio-visual generation apparatus is also provided, comprising: a first generation unit, configured to generate a sequence of action segments based on target text and action patterns input by a user, wherein each action segment in the action segment sequence includes at least a segment number, text content, action code, and action name; a second generation unit, configured to generate an audio segment corresponding to each action segment based on each action segment in the action segment sequence, and detect the audio duration of each audio segment; a third generation unit, configured to generate an initial video segment corresponding to each action segment based on a digital human type set by a user, and detect the video duration of each initial video segment; and a first processing unit. The system comprises four parts: a first processing unit, a second processing unit, and a third generation unit. The first processing unit performs duration alignment processing on the initial video segment corresponding to each action segment based on the audio and video durations corresponding to each action segment, resulting in an intermediate video segment with the same video duration as the corresponding audio duration. The second processing unit performs lip-sync processing on the intermediate video segment corresponding to each action segment based on the audio segment corresponding to each action segment, resulting in a target video segment with the digital human's lip movements synchronized with the corresponding audio content. The third generation unit generates a target audio-visual file based on the target video segment, audio segment, preset subtitles, and preset background corresponding to each action segment. The target audio-visual file is used to express the target text through a digital human's speech.
[0024] According to another aspect of this application, a computer program product is also provided, which stores a computer program, wherein, when the computer program is running, it controls the computer program product to execute any of the above-mentioned methods for generating digital human audio and video.
[0025] According to another aspect of this application, an electronic device is also provided, wherein the electronic device includes one or more processors and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the digital human audio and video generation method of any of the above.
[0026] In this application, firstly, an action segment sequence is generated based on the target text and action pattern input by the user. Each action segment in the action segment sequence includes at least a segment number, text content, action code, and action name. Then, an audio segment corresponding to each action segment is generated based on each action segment in the action segment sequence, and the audio duration of each audio segment is detected. Based on the digital human type set by the user, an initial video segment corresponding to each action segment is generated, and the video duration of each initial video segment is detected. Then, based on the audio duration and video duration of each action segment, the initial video segment corresponding to each action segment is time-aligned to obtain an intermediate video segment with the same video duration as the corresponding audio duration. Subsequently, based on the audio segment corresponding to each action segment, the intermediate video segment corresponding to each action segment is lip-synced to obtain a target video segment with the digital human's lip movements synchronized with the corresponding audio content. Based on the target video segment, audio segment, preset subtitles, and preset background corresponding to each action segment, a target audio-visual file is generated. The target audio-visual file is used to express the target text through digital human speech.
[0027] As described above, this application employs an action segment sequence-driven approach. By parsing the target text and action patterns input by the user, a structured sequence containing action codes is generated, and corresponding audio segments and initial video segments are generated independently. Subsequently, the initial video segment is time-aligned based on the detected audio and video durations to ensure that the duration of the intermediate video segments is strictly consistent with the audio segments. Then, lip-sync processing is performed on the intermediate video segments based on the audio segments. Finally, the target video segments, audio, subtitles, and background are integrated to generate the target audio-visual file. This achieves precise synchronization and unification of text content, voice audio, digital lip movements, and body movements on the timeline, thereby realizing the technical effect that the voice, lip movements, actions, and images in the generated digital human video remain coordinated, and the video segments have no sudden changes in action and the images are continuous. This solves the technical problem of poor effects in digital human audio-visual videos generated based on existing technologies. Attached Figure Description
[0028] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0029] Figure 1 This is a flowchart of an optional method for generating digital human audio and video according to an embodiment of this application;
[0030] Figure 2 This is an architecture diagram of an optional digital human motion control and video generation system according to an embodiment of this application;
[0031] Figure 3 This is a flowchart of an optional action parsing and matching module according to an embodiment of this application;
[0032] Figure 4 This is a flowchart of an optional user image preprocessing module according to an embodiment of this application;
[0033] Figure 5 This is a flowchart of an optional action migration module according to an embodiment of this application;
[0034] Figure 6 This is a flowchart of an optional transition template combination module according to an embodiment of this application;
[0035] Figure 7 This is a schematic diagram of an optional digital human audio and video generation apparatus according to an embodiment of this application;
[0036] Figure 8 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0039] It should also be noted that all information and data (including but not limited to information used for display and analysis) involved in this application are authorized by the user or fully authorized by all parties. For example, if there is an interface between this system and the relevant user or organization, before obtaining the relevant information, it is necessary to send a request to the aforementioned user or organization through the interface, and obtain the relevant information only after receiving consent from the aforementioned user or organization.
[0040] Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of relevant information and data involved in this application all comply with the relevant laws, regulations, and standards of the relevant regions, and necessary security measures have been taken. They do not violate public order and good morals. In addition, this application provides corresponding operation entry points for users to choose to agree to authorization or refuse authorization. If the user chooses to refuse authorization, the corresponding expert decision-making process will be initiated.
[0041] In one related technical embodiment, an image-driven digital human solution is provided. This type of solution typically inputs a reference image of a person and a piece of audio. It uses an audio feature extraction model to obtain information such as rhythm, phonemes, and prosody in the speech, then drives the image to generate a speaking video. The basic process is: image of person → audio input → audio feature extraction → head / mouth / expression generation → output speaking digital human video. For example, Sonic (an audio-driven portrait animation method) improves lip-sync, facial expressions, head movements, and long video stability through global audio perception. This method enables a single portrait image to speak based on audio, generating certain head and facial movements. The advantage of this type of solution is its simple input, making it suitable for quickly generating avatar-based explanatory videos; the disadvantage is that it mainly works on the head, mouth, and eyes, and usually does not support upper body, arm, and gesture movements, making it difficult to meet the needs of AI lecturers in scenarios such as welcoming, introducing, emphasizing, and guiding.
[0042] In one related technical embodiment, a video digital lip-syncing scheme is provided. This type of scheme typically takes an existing digital human video and target audio as input, redraws or replaces the mouth area in the video to synchronize the lip movements with the target audio. The basic process is: existing video → target audio → face / mouth area localization → audio-driven lip-syncing generation → output of the new lip-syncing video. For example, MuseTalk (a real-time, high-quality lip-syncing model) can modify the face area in the input video to synchronize it with the input audio and supports multiple languages such as Chinese, English, and Japanese audio; its disclosure mentions that it can achieve inference at over 30fps. The advantage of this type of scheme is that it can retain the body movements, clothing, and background of the original video, replacing only the mouth area; the disadvantage is that it requires suitable video material to already exist, cannot generate body movements from a single user image, and cannot solve the problem of adaptive completion and splicing when the action template and the text audio duration are inconsistent.
[0043] In one related technical embodiment, a motion-driven or posture-driven scheme is provided. This type of scheme typically takes a target person image and a reference motion video as input. It first extracts human posture, skeleton, or other motion control signals from the reference video, and then transfers the motion to the target person image to generate a video of the target person performing the same motion. The basic process is: target person image + reference motion video → human posture / motion feature extraction → motion transfer or posture-driven generation → output video with body motion. For example, MusePose (a motion-driven or posture-driven scheme) is commonly used to transfer the motion of a person in a reference video to a target person image. This type of scheme can generate upper body or full-body motion, suitable for scenarios requiring gestures and body movements. However, its inference process is close to video generation, resulting in high computational costs. If a complete motion transfer is performed on each text segment in a long text scenario for AI lecturers, it leads to long generation time and high computational consumption, making it difficult to meet the needs of batch, low-cost production.
[0044] In one related technical embodiment, a head motion transfer scheme is provided. This type of scheme typically takes a person image and a reference motion video as input. It mainly extracts head posture, facial expression changes, and some lip-sync information from the reference video and drives the target person image to generate corresponding head movements. Its basic process is: person image + reference head motion video → head posture / facial expression / lip-sync feature extraction → head region driving → output head animation video. For example, LivePortrait (a head motion transfer scheme) can generate head movements and facial expression changes from a static portrait image based on a reference video. It is suitable for rapid headshot-level animation generation. Its advantages are relatively fast speed and low barrier to entry; its disadvantage is that its main operating area is concentrated on the face and head, and it does not support complete upper body, arm, and gesture movements, thus failing to meet the needs of AI instructors for expressing body movements.
[0045] In one related technical embodiment, a large-scale model-based image-to-video (AGM) solution is provided. This type of solution typically takes a person image, text prompts, and audio as input, and a large-scale video generation model directly generates a speaking video with animation. The basic process is: person image + text prompts + audio → large-scale video generation model → output digital human video with animation. For example, Wan2.1 (a video generation model) can be used for text-to-video, image-to-video, and other generation tasks, aiming to improve video generation quality and open-ended generation capabilities. This type of method can generate richer actions and visual variations, but it suffers from high inference costs, slow speed, and insufficient stability and controllability because the actions are freely generated by the model based on the prompts. In AI lecturer scenarios, it is difficult to guarantee that a certain text corresponds to a specific action, and it is also difficult to ensure consistency and smooth splicing between multiple action segments.
[0046] As can be seen from the above, the technical problems that this application aims to solve in view of the technical defects existing in the above-mentioned related technical embodiments are as follows:
[0047] (1) Current image-based digital human technology mainly generates mouth, head, and eye movements based on audio, and usually does not support upper body, arm, and gesture movements, making it difficult to meet the needs of AI lecturers for expressing actions such as welcoming, introducing, emphasizing, and guiding. Therefore, this application aims to provide an action tag parsing and semantic action matching mechanism, so that users can manually specify actions, or the system can automatically match actions based on the semantics of the text.
[0048] (2) Current video lip-sync technology relies on existing video materials and cannot generate body movements from a single image uploaded by the user; although motion-driven technology can achieve motion transfer, the cost of full-body or upper-body motion transfer is high and it is not suitable for use in all segments of long text. Therefore, this application aims to provide a hierarchical motion generation strategy: the built-in digital human prioritizes the use of preset motion templates, and the user-uploaded image only performs full-body or upper-body motion transfer in speech motion segments with clear motion changes. During the transition frame interpolation stage, only head motion transfer is performed to reduce inference costs.
[0049] (3) Current large-scale model-based video generation solutions can generate videos with actions based on images, prompts, and audio, but the reasoning cost is high and the speed is slow. Moreover, since the actions are generated freely by the model, it is difficult to guarantee that the actions correspond precisely to the specified text segments, and it is also difficult to guarantee that the beginning and end states of multiple actions are consistent. Therefore, this application aims to achieve controllable generation of action selection, action duration, and segment connection through an action library, action tags, and combinable transition templates.
[0050] (4) The current action template scheme has the problem of inconsistent template duration and audio duration. When the template is longer than the audio, the complete playback will cause the action to be slow; when the template is shorter than the audio, stopping or switching directly will easily cause the picture to be discontinuous. Therefore, this application aims to provide a duration adaptive processing mechanism: when the template is longer than the audio, it is truncated; when the template is shorter than the audio, the remaining duration is made up by a combination of 1-second, 2-second, and 5-second transition templates, and alignment is achieved by mute when necessary.
[0051] Furthermore, given the common issue of incomplete upper bodies or arms in user-uploaded images, direct motion transfer can easily result in body distortion or missing arms. This application aims to improve the stability of motion transfer for user images by using upper body integrity detection, image completion, and similar built-in image matching to select a more suitable motion source.
[0052] In summary, this application aims to address the technical problems of insufficient motion control of existing AI lecturer digital humans, high cost of generating large models and uncontrollable motion, high cost of motion transfer, mismatch between template and audio duration, poor adaptability of user images, and unstable splicing of multiple segments.
[0053] The present invention will now be described in detail with reference to various embodiments.
[0054] Example 1
[0055] According to an embodiment of this application, an embodiment of a method for generating digital human audio and video is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0056] This application provides a digital human motion control and video generation system (hereinafter referred to as the generation system) for executing the digital human audio and video generation method of this application. Figure 1 This is a flowchart of an optional digital human audio and video generation method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0057] Step S101: Generate an action segment sequence based on the target text and action pattern input by the user. Each action segment in the action segment sequence includes at least a segment number, text content, action code, and action name.
[0058] Optionally, the target text represents the complete text content that the user wants the digital human to read. The target text can be set to plain text or text with action tags. The action mode is the mode that determines the action generation strategy, and can be one of the following: natural action mode, manually inserted action mode, or intelligent matching mode. For example, when the generation system is designed for long text reading scenarios, the action mode is manually inserted action mode, in which case the user can manually insert action tags into the target text. In the case of intelligent matching mode, AI automatically matches actions based on the content of the target text, thereby constructing a sequence of action segments.
[0059] Optionally, the generation system performs structuring processing through the above steps, transforming unstructured natural language text into a structured sequence of action fragments, thereby achieving a precise mapping between text content and action instructions. The generation of action fragment sequences supports manual user control or automatic intelligent matching by the system, solving the problems of single action expression and inability to dynamically adjust actions according to text semantics in traditional technologies. This improves the flexibility and controllability of action generation and provides a standardized data foundation for subsequent audio and video synchronization generation.
[0060] Step S102: Generate an audio segment corresponding to each action segment based on each action segment in the action segment sequence, and detect the audio duration of each audio segment.
[0061] Optionally, the generation system traverses the sequence of action segments, inputs the text content of each segment into the text-to-speech (TTS) model in turn, synthesizes the corresponding speech audio, and forms audio segments. At the same time, the generation system records the start time and end time of each audio segment, thereby detecting the audio duration of each audio segment.
[0062] Optionally, the generation system achieves text-to-speech conversion through the above steps, accurately quantifies the temporal attributes of each semantic segment, establishes a temporal benchmark between semantic content and acoustic features, provides an accurate temporal reference for the duration alignment of subsequent video segments, and ensures that subsequent processing can perform cropping, frame interpolation, or synchronization operations based on accurate duration information, avoiding audio-visual desynchronization caused by duration estimation errors.
[0063] Step S103: Generate initial video segments corresponding to each action segment based on the digital human type set by the user, and detect the video duration of each initial video segment.
[0064] Optionally, the digital human type can be either a built-in type or an uploaded type.
[0065] Optionally, the generation system first detects the type of digital human set by the user. If it is a built-in type, the generation system directly calls the corresponding action template (e.g., speech action template, natural action template, transition action template) from the action library based on the action code in the action segment to generate the initial video segment. If it is an uploaded type, the generation system first preprocesses the user's image and performs source image similarity matching. Then, it performs action migration based on the action template of the matched source image to generate the initial video segment. Finally, the system detects the video duration of each initial video segment.
[0066] Optionally, the generation system, through the above steps, adopts differentiated video generation strategies based on the different sources of the digital humans. For built-in digital humans, it directly calls preset templates, avoiding high-cost real-time inference and improving generation efficiency. For user-uploaded images, through preprocessing and hierarchical motion transfer, it reduces computing power consumption while ensuring the expressiveness of the motion, which helps to achieve low-cost and high-efficiency motion generation and obtain the true duration of the video, providing input data for subsequent duration alignment.
[0067] Step S104: Based on the audio duration and video duration corresponding to each action segment, perform duration alignment processing on the initial video segment corresponding to each action segment to obtain an intermediate video segment with the same video duration as the corresponding audio duration.
[0068] Optionally, duration alignment processing is an operation process that makes the video duration consistent with the audio duration through methods such as truncation, stabilization, transition template combination, and silent padding.
[0069] Optionally, the generation system compares the audio duration and video duration corresponding to each action segment. If the video duration is longer than the audio duration, the generation system truncates the video segment, performs stabilization processing, generates a stabilized segment, and splices the truncated segment and the stabilized segment together. At the same time, a silent audio is added to the end of the audio to lengthen the video of the spliced new video segment. If the video duration is shorter than the audio duration, the generation system calculates the remaining duration (i.e., the difference between the audio duration and the video duration) and supplements the video segment by combining transition segments generated based on transition templates of different specifications of 1 second, 2 seconds, or 5 seconds.
[0070] Optionally, the generation system, through the above steps, solves the problems of audio-visual asynchrony, sluggish movements, or abrupt scene transitions caused by the inconsistency between the action template duration and the audio duration. Through stabilization processing and transition template combinations, it ensures that video clips are strictly synchronized with the audio in the temporal dimension, while maintaining visual continuity and stability. Its advantage lies in achieving smooth splicing between multiple clips, avoiding abrupt changes in movement, and improving the overall smoothness and professionalism of the final video.
[0071] Step S105: Based on the audio segment corresponding to each action segment, perform lip-sync processing on the intermediate video segment corresponding to each action segment to obtain the target video segment in which the digital lip movements are synchronized with the corresponding audio content.
[0072] Optionally, the system calls a lip-sync model to extract phonemes, prosody, and other features from the audio, driving the mouth area of the person in the intermediate video clip to generate lip movements that strictly match the audio content, thereby processing to obtain the target video clip. In the target video clip, the lip movements and head actions of the digital human are synchronized with the speech content.
[0073] Optionally, the generation system addresses the issue of lip-syncing and speech disconnect during motion generation through the aforementioned steps. Regardless of whether the initial video clip is generated via template calling or motion transfer, lip-sync processing ensures consistency between the character's lip movements and speech content in the final video, thereby enhancing the naturalness and realism of the digital human's delivery and meeting the high-precision audio-visual synchronization requirements of AI lecturer scenarios.
[0074] Step S106: Generate a target audio-visual file based on the target video clip, audio clip, preset subtitles, and preset background corresponding to each action segment. The target audio-visual file is used to express the target text through a digital human speech.
[0075] Optionally, preset subtitles refer to a text display layer that is prepared in advance or automatically generated and synchronized with the audio; preset backgrounds refer to the background layer of the video screen, which can be replaced with a solid color or a specific scene.
[0076] Optionally, the generation system splices the target video and audio segments along the timeline according to the sequence of action segments. At the same time, it aligns the preset subtitles with the audio content and overlays them onto the video screen, and uses the preset background as the video base layer or replaces the original background, thereby encapsulating and generating a multimedia file containing video, audio, subtitles, and background, i.e., the target audio and video file.
[0077] Optionally, the generation system completes the final encapsulation and output of multimodal data through the above steps, realizing the collaborative generation of video, audio, subtitles and background. Through automated splicing and overlay, the generation improves video production efficiency, reduces the cost of manual post-production, and meets the user needs of AI lecturers, training courseware and other scenarios for batch and efficient generation of standardized video content.
[0078] As described above, this application employs an action segment sequence-driven approach. By parsing the target text and action patterns input by the user, a structured sequence containing action codes is generated, and corresponding audio segments and initial video segments are generated independently. Subsequently, the initial video segment is time-aligned based on the detected audio and video durations to ensure that the duration of the intermediate video segments is strictly consistent with the audio segments. Then, lip-sync processing is performed on the intermediate video segments based on the audio segments. Finally, the target video segments, audio, subtitles, and background are integrated to generate the target audio-visual file. This achieves precise synchronization and unification of text content, voice audio, digital lip movements, and body movements on the timeline, thereby realizing the technical effect that the voice, lip movements, actions, and images in the generated digital human video remain coordinated, and the video segments have no sudden changes in action and the images are continuous. This solves the technical problem of poor effects in digital human audio-visual videos generated based on existing technologies.
[0079] Optionally, Figure 2 This is an architecture diagram of an optional digital human motion control and video generation system according to an embodiment of this application, such as... Figure 2 As shown, the generation system comprises four layers: input and management layer, core logic layer, execution generation layer, and synthesis output layer. Simultaneously, the generation system collaboratively completes digital human motion control and video generation by calling 13 core modules: text input module, action library management module, action parsing and matching module, TTS and time alignment module, digital human type judgment module, built-in template calling module, user image preprocessing module, similar image matching module, action transfer module, transition template combination module, lip-sync module, video stitching module, and output module.
[0080] Optionally, in the input and management layers, the generation system receives long text input by the user (i.e., target text) through the text input module and determines the action generation method (i.e., action mode). The text input by the user can be plain text or text with action tags. The action generation method can be one of three: natural action mode, manual action insertion mode, or intelligent matching mode. For example, when using the manual action insertion mode, the user can insert preset action tags into the text. For instance, the long text might be: "[act2] Hello everyone, welcome to China Tower AI Lecturer; [act8] Now, let me introduce the system..." "Core Capabilities" refers to the preset action tags [act2] and [act8], which correspond to different action templates in the action library. The generation system determines the target action of the corresponding text segment based on the action tags. When using AI matching mode (i.e. intelligent matching mode), users do not need to manually insert action tags. The system performs semantic analysis on the text through a large model. Based on the semantics of the text, such as opening greetings, function introductions, key points, risk reminders, phase summaries, and acknowledgments, it automatically matches appropriate actions from the action library. The text input module outputs the user's text and action generation mode and passes them to the action parsing and matching module.
[0081] Optionally, in the input and management layers, the generation system uses the action library management module to manage the digital human image as the basic management unit. This module uniformly manages the pre-built digital human static images, speech action templates, natural action templates, transition action templates, and structured annotation data for each action template. Each digital human image corresponds to an independent material folder. For example, male images can be named m1, m2, m3, etc., and female images can be named w1, w2, w3, etc. It is preferred that the materials for each digital human image are collected or created using a solid color background to reduce the interference of the background on subsequent keying, background replacement, motion transfer, and visual feature matching.
[0082] Furthermore, the generation system establishes structured data records for each digital human image in the preset action library, including: (1) digital human image ID; (2) image feature tags: gender, facial features, 3D reconstruction SMPL body shape parameters, etc.; (3) static image image; (4) action template video; (5) natural action template video; (6) transition action template video; (7) auxiliary annotation data: mask video or mask image, face frame, 2D face key points, 2D human body key points, 3D face key points, etc. Among them, image feature tags are used for similar image matching when users upload images; action template videos are used to generate speech action segments, such as waving, welcoming, introducing, emphasizing, guiding, etc.; natural action templates are used for ordinary explanation segments where users have not specified specific actions; transition action templates include specifications such as 1 second, 2 seconds and 5 seconds, used to supplement the remaining video duration when the action template duration is shorter than the audio. Auxiliary annotation data is used to support the subsequent video generation process. For example, masked videos or masked images can be used for foreground extraction and background replacement; face bounding boxes can be used for lip-syncing and face region localization; 2D facial landmarks and 2D human landmarks can be used for motion transfer and pose alignment; and 3D facial landmarks can be used for head pose estimation, expression-driven algorithms, and face stability constraints.
[0083] In addition, the generation system also establishes index information for each action material in the preset action library, including action code, material name, corresponding action, action category, action purpose, applicable image, template duration, and beginning and end status markers. For example, action code 1 corresponds to "emphasizing with both hands", action code 2 corresponds to "waving hello with the right hand", action codes 9 to 14 correspond to natural actions, and action codes 15 and above can be expanded to include transition actions such as emphasis, clicking, liking, bowing, refusing, bowing, thinking, inviting, and writing.
[0084] Furthermore, to ensure stable splicing of multiple segments, the generation system allows the action template, natural action template, and transition action template to be set to have the same first and last frames, i.e., all returning to a unified initial state. The unified initial state can be a standard posture of the digital human standing facing forward, with the head facing forward, the body stable, and the hands in a natural position. Through the design of the unified initial state, different speech action templates, natural action templates, and transition action templates can be connected in a unified state, reducing abrupt changes in action and screen jumps in the final generated digital human audio and video.
[0085] In practical applications, for the built-in digital human path, the motion library management module directly provides the built-in template calling module with auxiliary data such as the corresponding image's motion template, natural motion template, transition template, and mask; for the user-uploaded image path, the motion library management module provides the similar image matching module with image feature tags and reference images, and provides the motion transfer module, lip-sync module, and transition template combination module with the selected source image's motion template, key points, mask, and transition template.
[0086] Optionally, in the core logic layer, a system call action parsing and matching module is generated. First, it detects the user-input action pattern, which can be one of the following types: natural action pattern, manually inserted action pattern, or intelligent matching pattern. Then, if the action pattern is natural action pattern / intelligent matching pattern, it matches the semantic features of each text fragment obtained by dividing the target text into segments in a preset action library to obtain the index information of the action template corresponding to each text fragment. The index information of each action template includes at least the action code and the action name. If the action pattern is manually inserted action pattern, it extracts the pre-inserted action tags from the beginning of each text fragment in the target text and queries the preset action library based on each action tag to obtain the index information of the action template corresponding to each text fragment. Subsequently, it generates an action fragment sequence based on the text content corresponding to each text fragment and the index information of the action template.
[0087] In other words, the action parsing and matching module is used to parse, match, and generate structured fragments from the input text based on the user's selected action pattern. The generation system supports three modes: natural action mode, manually inserted action mode, and AI matching mode. Figure 3 This is a flowchart of an optional action parsing and matching module according to an embodiment of this application, such as... Figure 3 As shown, the workflow of the action parsing and matching module is as follows:
[0088] (1) Natural movement pattern:
[0089] The natural action mode is the system default mode. This mode is suitable for scenarios where the user has not specified specific gestures and only wants the digital human to maintain a natural speaking state during the explanation. In this mode, the generation system sends the long text input by the user or the text after preliminary segmentation to the big model. The big model selects the most suitable natural action type from several preset natural action templates based on the overall semantics of the text, the explanation scenario and expression style.
[0090] For example, the generation system has preset 6 types of natural action templates for the natural action mode, covering the style of action in news broadcasting, stage speech, advertising, or other natural explanations. Each type of natural action template is pre-configured with corresponding descriptive words or semantic tags. The large model determines the natural action corresponding to each text segment based on the degree of matching between the text content and the descriptive words / semantic tags corresponding to the above natural actions. In the natural action mode, obvious gestures are generally not generated. The natural action templates are mainly used to maintain the digital human's natural speech, slight body movements, and stable explanation state.
[0091] (2) Manual insertion mode:
[0092] The manual action insertion mode is designed to meet users' needs for precise control over actions. Users can pre-select target actions in the front-end action selection interface, and the generation system will insert action tags at the corresponding positions in the long text.
[0093] For example, the generated long text might be: "[act2] Hello everyone, welcome to China Tower AI Instructor. [act8] Now, let me introduce the core capabilities of the system." In this case, the action parsing and matching module parses the action code in the action tag and binds that action code to the corresponding text fragment. Assuming that when [act2] is recognized, the system establishes a correspondence between the action template corresponding to action code 2 and the subsequent text fragment. This mode does not require the large model to reselect specific actions; the system only completes action parsing based on the tags manually inserted by the user.
[0094] (3) AI matching mode:
[0095] AI matching mode is used to automatically match explicit actions to text without the user manually inserting them. In this mode, the system first automatically segments the long text input by the user into multiple short sentences or semantic fragments; then, it calls a large model or semantic matching model to match each short sentence with preset action semantic tags in the action library, and selects the most suitable action for that short sentence from the action library. Action semantic tags can include types such as opening greetings, welcomes, introductions, guidance, emphasis, reminders, reflection, presentations, summaries, and thanks. For example, short sentences with action semantic tags containing the semantics of "Hello everyone" or "Welcome to use" can be matched with waving or welcoming actions; short sentences with action semantic tags containing the semantics of "Let me introduce" or "Please look here" can be matched with introduction or guidance actions; short sentences with action semantic tags containing the semantics of "Please pay attention" or "Important points" can be matched with emphasis actions; and short sentences with action semantic tags containing the semantics of "Thank you for watching" or "Thank you everyone" can be matched with thanks or bowing actions.
[0096] It's important to note that regardless of whether the natural action mode, manual action insertion mode, or AI matching mode is used, the action parsing and matching module ultimately outputs a structured sequence of action segments in a unified format. Each action segment includes at least a segment number, text content, action code, and action name. For example: {"segment_id":1,"text":"Hello everyone, welcome to China Tower AI Instructor.","action_id":2,"action_name":"Wave hello with your right hand"}, where "segment_id" is the segment number, "text" is the text content, "action_id" is the action code, and "action_name" is the action name.
[0097] Furthermore, the structured action segments in the structured action segment sequence may also include fields such as action type, whether it is a speech action segment, whether action transfer is required, action template path, and candidate action confidence, which are used to realize subsequent TTS time alignment, template calling, action transfer, and video splicing processing.
[0098] As can be seen from the above, through the action parsing and matching module, the generation system can achieve three types of action control according to different user needs: ensuring the stability and naturalness of ordinary explanations through the natural action mode, ensuring precise user control of actions through the manual insertion action mode, and achieving automatic action matching for each sentence of long text through the AI matching mode, thereby improving the flexibility and controllability of AI lecturer video production. Specifically, the technical effects achieved by the action parsing and matching module are as follows:
[0099] (1) Improve the flexibility and accuracy of action control: In the manual action insertion mode, users can directly query the preset action library through action tags, realizing precise control of specific text fragment actions, and meeting users' high-precision needs for action expression of key nodes (such as welcome, emphasis, and thanks); In the intelligent matching mode, the system automatically matches the action template index based on the semantic features of the text fragment, realizing the automatic generation of sentence-by-sentence actions in long text scenarios, reducing the user's operation threshold and improving the generation efficiency; In the natural action mode, the system selects a natural broadcasting style according to the overall semantics, ensuring the coherence and naturalness of ordinary explanation fragments, thereby solving the problem of single action expression or inflexible control in the existing technology, and can provide action control capabilities in multiple dimensions from fully automatic to semi-automatic to fully manual according to different application scenarios and user needs.
[0100] (2) Achieving accurate mapping between actions and text: By matching based on semantic features or querying based on action tags, the system can accurately establish a correspondence between each text fragment and a specific action template in the preset action library, generating index information containing action codes and action names. Subsequently, based on the text content and index information, an action fragment sequence is generated, transforming unstructured text data into structured action instruction data. Thus, this application solves the problem of disconnect between actions and speech / content in traditional digital human technology, ensuring that the actions of digital humans can accurately reflect the semantic focus and emotional color of the text, and enhancing the expressiveness and appeal of the video.
[0101] Optionally, in the core logic layer, the system call TTS and time alignment module first performs speech synthesis based on the text content in each action segment according to the order corresponding to the action segment sequence to obtain the audio segment corresponding to each action segment. Then, it detects the start time and end time of each audio segment in the complete audio corresponding to all action segments. Finally, it determines the audio duration of each audio segment based on the start time and end time of each audio segment.
[0102] In other words, the TTS and time alignment module is used to convert each text segment into speech audio based on the structured action segment sequence output by the action parsing and matching module, and to establish a time mapping relationship between text segments, audio segments, and action templates. The system first receives the structured results output by the action parsing and matching module. Regardless of whether the result comes from natural action mode, manually inserted action mode, or AI matching mode, each segment contains at least a segment number, text content, action code, and action name. Subsequently, the system performs TTS speech synthesis according to the segment order, generates corresponding audio, and records the start time, end time, and audio duration of each segment in the complete audio. For example, inputting {"segment_id":1,"text":"Hello everyone, welcome to China Tower AI Instructor.","action_id":2,"action_name":"Hello with a right hand wave"} will result in the following output after processing by this module: {"segment_id":1,"text":"Hello everyone, welcome to China Tower AI Instructor.","audio_start":0.00,"audio_end":3.60,"audio_duration":3.60,"action_id":2,"action_name":"Hello with a right hand wave"}. This time mapping result is used to subsequently determine the relationship between the action template duration and the corresponding audio duration. When the action template duration is longer than the audio duration, the system can truncate the action template; when the action template duration is shorter than the audio duration, the system can calculate the remaining duration and supplement it through transition template combinations. This time mapping also provides a basis for lip-syncing, silence completion, multi-segment splicing, and the final video timeline generation.
[0103] As can be seen from the above, the TTS and time alignment module achieves the following technical effects by serially generating audio and accurately recording timestamps:
[0104] (1) Establishing a precise time mapping relationship: By detecting the start and end times of each audio segment in the complete audio, the generation system can accurately grasp the specific position of the speech corresponding to each text segment on the time axis. This precise time mapping solves the technical problem of how to accurately correspond the time relationship between "text-speech-action" in long text scenarios, and provides a reliable time benchmark for subsequent video segment generation and duration alignment.
[0105] (2) Support for differentiated duration processing strategies: By calculating the audio duration, the system can compare it with the video duration of the subsequently generated initial video segment. If the video duration is greater than or less than the audio duration, the system can flexibly choose duration alignment processing methods such as truncation, stabilization, or transition template combination based on the specific duration difference (such as remaining duration or truncation position). Thus, the generation system ensures that the duration of the final output video segment is strictly consistent with the duration of the audio segment, avoiding the problem of audio-visual asynchrony.
[0106] (3) Improve the automation and accuracy of the generation process: This step does not require manual intervention and automatically completes the feature extraction from text to audio and the acquisition of time dimension data, which improves the automation level of AI lecturer video generation. At the same time, positioning based on the timeline of the complete audio ensures the temporal continuity when splicing multiple segments and enhances the logic and coherence of the final generated video.
[0107] Optionally, in the core logic layer, a system call to the digital human type determination module is generated to detect the digital human type set by the user, wherein the digital human type is one of the following:
[0108] (1) The system has built-in digital human type (i.e., built-in type), which is a digital human with a static image, action template, natural action template and transition template pre-configured by the system; if it is determined to be a built-in digital human, the system will enter the built-in template call path.
[0109] (2) User-uploaded image digital human (i.e. upload type), that is, the user provides a custom image of a person, which needs to be further processed by the system to generate a corresponding action video; if it is determined to be a user-uploaded image digital human, the system will enter the user image preprocessing, similar image matching and action migration path.
[0110] Optionally, in the execution and generation layer, if the user-set digital human type is a built-in type, the generation system calls the built-in template calling module. First, it queries the preset action library to obtain the action template set corresponding to the user-selected built-in digital human image. Then, based on the action code in each action segment, it queries the action template set corresponding to the built-in digital human image to obtain the action template corresponding to each action segment. Finally, based on the action template corresponding to each action segment, it generates the initial video segment corresponding to each action segment.
[0111] In other words, the built-in template calling module is used to directly call the video template corresponding to the built-in image from the action library when the target digital human is a built-in digital human in the system, based on the structured action segment results output by the action parsing and matching module. Since the action template, natural action template and transition template of the built-in digital human have been pre-made and stored in the action library, the built-in template calling module does not need to perform full-body action transfer and can directly generate the corresponding video segment through template calling.
[0112] Specifically, for speech action segments obtained through manual insertion or AI matching, the built-in template invocation module retrieves the target action template from the action template set corresponding to the current built-in digital human avatar based on the action code in the action segment, thereby generating the corresponding initial video segment. For example, when the action_id in the structured segment is 2, the system calls the action template corresponding to "waving hello with the right hand" under that built-in avatar; when the action_id is 8, the system calls the action template corresponding to "introducing to the upper right with the right hand".
[0113] In addition, for natural action mode, or ordinary explanation segments that are determined to be natural actions by the action parsing and matching module, the built-in template calling module calls the natural action template corresponding to the current built-in digital human image. The natural action template is usually used in ordinary explanation scenarios such as news broadcasting, stage speeches, and advertising, and can maintain the digital human's natural broadcasting state when the user does not specify a specific action.
[0114] As can be seen from the above, by calling the built-in template calling module, the generation system can quickly generate action segments using preset video templates in the built-in digital human scenario, avoiding the high-cost action generation or motion transfer for each segment, thereby improving the efficiency and stability of AI lecturer video generation. Specifically, the built-in template calling module, through the template direct calling mechanism under the built-in digital human path, can achieve the following technical effects:
[0115] (1) Reduce reasoning cost and generation time: For built-in digital humans, the system does not need to execute computationally intensive full-body or half-body motion transfer algorithms, nor does it need to call large-scale image-generated video models. The system only needs to perform data query and index matching in the preset motion library and directly read the pre-made video template file. This retrieval-based processing method avoids the huge computational power consumption of real-time video generation, thereby improving the generation efficiency of AI lecturer long text videos and realizing low-cost video production.
[0116] (2) Ensure consistency and stability of action performance: Since the action templates are carefully made or captured in advance, the action quality of the built-in digital human (such as gesture standardization, action smoothness, and background purity) is controllable and fixed. By accurately retrieving the corresponding action template based on action coding, it is ensured that each text segment can obtain predefined high-quality action feedback, avoiding action deformation, artifacts or instability that may be caused by real-time generation, and improving the quality of the final video.
[0117] (3) Simplify the processing flow and improve system reliability: The direct template calling mechanism omits complex steps such as image preprocessing, feature extraction, pose alignment, and motion transfer, and only retains the template query and playback logic. This simplified pipeline processing reduces potential technical failure points (such as key point detection failure, feature matching error, etc.) and improves the stability and robustness of the system in batch generation scenarios.
[0118] Optionally, in the execution and generation layer, when the user sets the digital human type to the upload type, the generation system first preprocesses the user-uploaded digital human image through the user image preprocessing module to obtain a target image that meets the motion transfer conditions. The motion transfer conditions are used to detect the completeness of the digital human image in the image and extract the auxiliary data corresponding to the target image. The auxiliary data includes the face bounding box position, the person mask, the human body key points, and the face key points. Then, the similar image matching module matches the target image in the built-in digital human image library and uses the matched built-in digital human image as the source image corresponding to the target image. Subsequently, the motion transfer module performs motion transfer based on the auxiliary data corresponding to the target image and the motion template set corresponding to the source image to obtain the initial video segment corresponding to each motion segment.
[0119] Optionally, the step of preprocessing the user-uploaded digital human image through the user image preprocessing module to obtain a target image that meets the action transfer conditions includes: First, standardizing the user-uploaded digital human image to obtain a standardized image, wherein the standardization process is used to adjust the size of the digital human image based on coordinate mapping information. Then, if the standardized image includes a single face, the single face is taken as the target face; if the standardized image includes two or more faces, the face located in the center region of the standardized image and having the largest area is taken as the target face. Subsequently, in the standardized image, the person to which the target face belongs is selected for image cutout and background replacement to obtain an initial image. Then, if the integrity of the digital human image in the initial image does not meet the action transfer conditions, the digital human image in the initial image is updated based on an image completion algorithm to obtain the target image.
[0120] Figure 4 This is a flowchart of an optional user image preprocessing module according to an embodiment of this application, such as... Figure 4 As shown, the workflow of the user image preprocessing module includes:
[0121] First, the system receives static images of people uploaded by users and determines whether the images contain coordinate mapping information of the preset avatar_list. If the images contain coordinate mapping information, the system adjusts the size and basic position of the images based on this information. If the images do not contain coordinate mapping information, the system directly proceeds to the face detection process.
[0122] The avatar_list coordinate mapping information represents the position of the face / person in the image. This information includes fx, fy, fw, and fh. fx and fy are the starting coordinates, while fw and fh refer to the width and height dimensions. The user image preprocessing module adjusts the image size based on the avatar_list coordinate mapping information. This means resizing the user-uploaded static image of a person to the specified width fw and height fh, and then aligning it according to the fx and fy coordinates (i.e., basic pasting).
[0123] In the face detection stage, the generation system determines whether there are valid faces in the image. When multiple faces are detected, the face located in the center of the image with the largest area is selected as the target face by default to avoid background figures or non-target figures interfering with subsequent processing.
[0124] Subsequently, the system performs portrait cutout on the target person and replaces the original background with a solid color background to reduce the impact of complex backgrounds on similar image matching, motion transfer, and video compositing. After the background is unified, the system performs human key point detection to obtain the positions of key parts such as the head, shoulders, arms, and upper body.
[0125] The generation system then determines whether the image meets the motion transfer criteria based on key human body points. The criteria include whether the upper body is complete, whether the shoulders are clear, and whether the arms have sufficient visible area. If the motion transfer criteria are met, the auxiliary data extraction process begins. If the criteria are not met, for example, if only the head is included, the shoulders are missing, or the arms are not fully visible, an image completion algorithm is called to complete the missing shoulder, upper body, and arm areas, generating a standardized human image (i.e., a standardized image) that meets the motion transfer criteria.
[0126] After completing the above processing, the generation system synchronously extracts and caches auxiliary data. The auxiliary data includes at least the face bounding box position, the person foreground mask, the human body key points, and the facial key points. Among them, the face bounding box is used for subsequent lip-syncing, the person foreground mask is used for background replacement, and the human body and facial key points are used for similar image matching, pose alignment, and motion transfer. Finally, the generation system outputs standardized human images and their auxiliary information through the user image preprocessing module for use in subsequent similar image matching and transition processing.
[0127] As can be seen from the above, by calling the user image preprocessing module, a strict image preprocessing and integrity verification mechanism is established, which can achieve the following technical effects:
[0128] (1) Improve the input quality and stability of action transfer: The image size was unified through standardization, and the accurate locking of the target person was ensured by the face detection strategy (prioritizing the face in the central area with the largest area), avoiding interference from background people or non-target people. The subsequent image matting and solid color background replacement eliminated the noise influence of complex background on similar image matching and feature extraction in action transfer, and improved the model's accuracy in recognizing the main features of the person.
[0129] (2) Solving the problem of motion distortion caused by incomplete user-uploaded images: To address the common issues of missing upper body parts, arms obscuring the image, or incomplete shoulders in user-uploaded images, this solution introduces a motion transfer condition judgment and image completion mechanism. When an incomplete image is detected, the completion algorithm is automatically invoked to generate the missing body parts. This mechanism ensures that the images input to the motion transfer module have complete human structural features, thereby avoiding motion transfer failure, limb proportion imbalance, or boundary artifacts caused by missing body parts, and improving the naturalness and rationality of digital human motion generation.
[0130] Optionally, the step of matching the target image in the built-in digital human image library using the similar image matching module includes: first, detecting the gender of the target digital human image in the target image and the gender of each built-in digital human image in the built-in digital human image library; then, taking the set of built-in digital humans with the same gender as the target digital human image as the candidate set; then, detecting the facial similarity and body shape similarity between each built-in digital human image in the candidate set and the target image; then, performing a weighted summation of the facial similarity and body shape similarity to obtain the target similarity between each built-in digital human image and the target image; finally, taking the built-in digital human image with the highest target similarity in the candidate set as the source image corresponding to the target image.
[0131] In other words, the similar image matching module is used to select the source image that is closest to the user's uploaded image from the built-in digital human image library in the path of the digital human image uploaded by the user. This provides a reference for subsequent action migration and transition template calls. The similar image matching module first performs gender recognition on the user's uploaded image and the built-in digital human image, and then performs candidate screening of the built-in images based on the gender information. For built-in images with inconsistent genders, they can be directly excluded or their matching weight can be reduced.
[0132] First, in the candidate image set, the generation system further calculates facial similarity and body shape similarity. Facial similarity is obtained by extracting facial feature vectors from user images and built-in images through a facial recognition model and calculating the cosine similarity between them. Body shape similarity is obtained by estimating the human body shape parameter β of user images and built-in images through SMPL (Skinned Multi-Person Linear model) or SMPL-X (extended version of skinned multi-person linear model) 3D human body reconstruction models and calculating it based on the distance between the β parameters. The β parameters are used to characterize structural features such as human body shape, shoulder width, torso proportion, and limb proportion.
[0133] Next, the similarity matching module calculates a comprehensive score S (i.e., target similarity) for each built-in digital human image using the formula S = w1 × Sface + w2 × Sbody. Then, based on the comprehensive score S, it selects a source image from the digital human image library that matches the user's image. Here, S represents the comprehensive similarity, Sface represents the face similarity, Sbody represents the body shape similarity, and w1 and w2 are weight parameters. The body shape similarity is calculated using the formula Sbody = 1 / (1 + ||βuser - βsource||2), where βuser represents the human body shape parameters corresponding to the user's uploaded image, and βsource represents the human body shape parameters corresponding to the candidate built-in image.
[0134] Subsequently, the generation system selects the built-in image with the highest overall similarity as the source image, and calls the corresponding action template, natural action template, transition template, mask and key point data of the source image for subsequent action transfer and video generation.
[0135] As can be seen from the above, the similar image matching module enables the generation system to simultaneously consider both facial appearance consistency and human body structure consistency, avoiding excessive differences in body shape between the source image and the target image due to relying solely on overall image semantic matching. This improves the stability of shoulder, arm, and upper body postures during motion transfer. Specifically, the similar image matching module can achieve the following technical effects:
[0136] (1) Improving the structural adaptability and naturalness of motion transfer: Traditional techniques often rely solely on appearance or overall semantics for image matching, which can easily lead to significant differences between the source image and the target image in terms of body shape (such as shoulder width, torso proportion, and limb length), resulting in limb distortion, disproportion, or posture misalignment during motion transfer. This solution introduces body shape similarity detection based on models such as SMPL to ensure that the selected source image is highly consistent with the target image in terms of human geometry. This structural matching is the foundation of high-quality motion transfer, which can significantly reduce deformation errors during the motion transfer process, making the generated digital human movements more natural and ergonomic.
[0137] (2) Enhance visual consistency and realism: By simultaneously considering facial similarity and body shape similarity, and using a weighted summation method to determine the final target similarity, the system can prioritize the built-in image that is closest to the target user in both appearance (face) and structure (body shape). This not only ensures the continuity of facial features, but also ensures the coordination of body structure, thereby enhancing the visual realism and user immersion of the digital human in the final generated video.
[0138] (3) Optimize matching efficiency and accuracy: By narrowing the search range through gender screening, a candidate set is formed, avoiding the waste of computing power caused by full calculation in the entire database. Then, fine-grained similarity calculation is performed in the candidate set, which not only ensures the accuracy of matching, but also improves the processing efficiency of the system.
[0139] Optionally, the step of performing motion transfer based on auxiliary data corresponding to the target image and a set of motion templates corresponding to the source image through the motion transfer module includes: when the auxiliary data of the target image enters the motion transfer module for the first time, the motion transfer module extracts features from the target image to obtain feature data, and caches the feature data and auxiliary data as target features corresponding to the target image in a preset cache layer. The feature data includes target image features, target image encoding features, and latent variable features. Then, based on the motion encoding in each motion segment, a query is performed in the set of motion templates corresponding to the source image to obtain the motion template corresponding to each motion segment. Finally, a query is performed in the preset cache layer to obtain the pre-defined motion template. First, the template features of the action template corresponding to each action segment are cached offline. The template features include action posture sequence, key point data, person mask, action trajectory features, and posture guidance features. Then, based on the action template corresponding to each action segment, the action transfer strategy corresponding to each action segment is determined. The action template is one of the following types: speech action template, natural action template, or transition action template. The action transfer strategy is one of the following types: face action transfer strategy or human body action transfer strategy. Then, based on the action transfer strategy, template features, and target features corresponding to the target image, action transfer is performed to obtain the initial video segment corresponding to each action segment.
[0140] Specifically, the step of determining the action transfer strategy corresponding to each action segment based on the action template corresponding to each action segment includes: if the action template corresponding to the i-th action segment is a speech action template, the action transfer strategy corresponding to the i-th action segment is determined to be a human action transfer strategy; if the action template corresponding to the i-th action segment is a natural action template / transitional action template, the action transfer strategy corresponding to the i-th action segment is determined to be a face action transfer strategy.
[0141] Specifically, the steps for motion transfer based on the motion transfer strategy, template features, and target features corresponding to each motion segment include: when the motion transfer strategy corresponding to the i-th motion segment is a human motion transfer strategy, aligning the motion pose sequence in the template features and the human key points in the target features to obtain target pose data, wherein the target pose data is used to characterize the body pose and gesture position of the target digital human image in each frame of the motion pose sequence; based on the target pose data and the target features corresponding to the target image, generating the video frame sequence corresponding to the i-th motion segment through a preset motion transfer algorithm; and performing enhancement operations on the video frame sequence corresponding to the i-th motion segment to obtain the initial video segment corresponding to the i-th motion segment.
[0142] Figure 5This is a flowchart of an optional motion transfer module according to an embodiment of this application. The motion transfer module is used to transfer the source motion to the target image uploaded by the user in the digital human path of the user-uploaded image, based on the source image and its motion template determined by the similar image matching module, and generate a corresponding digital human motion video clip, such as... Figure 5 As shown, the workflow of the action transfer module includes at least the stages of basic data reception, feature loading and caching, hierarchical action transfer, inference acceleration, and post-processing synthesis.
[0143] During the basic data receiving phase, the motion transfer module first receives the basic data required for motion transfer. The basic data includes two categories: target data and source data. The target data includes user-uploaded and preprocessed target images, face bounding boxes, human body key points, and facial key points. The source data includes source image motion templates, source motion pose sequences, human masks, motion trajectory features, etc. The above basic data is used for subsequent pose alignment, feature injection, motion generation, and region constraint processing.
[0144] During the feature loading and caching stage, considering that the same user image typically remains unchanged throughout the generation of an AI lecturer video, the generation system needs to cache user image-related features. Specifically, when processing the target image for the first time, the system extracts and caches intermediate results such as target image features, image encoding features, latent variable features, face bounding boxes, and key points. When subsequent action segments continue to use the same user image, the system directly reuses the cached user features, avoiding repeated image encoding and feature extraction. Simultaneously, for source action templates in the action library, since they are pre-set fixed materials, the system can extract and save source action features offline in advance, including the pose sequence, key point sequence, person mask, action trajectory, and pose guidance features of the source template. During online generation, the action transfer module directly reads the pre-calculated source action features, avoiding repeated extraction of source poses and motion trajectories, thereby reducing online inference time.
[0145] In the tiered motion transfer stage, the motion transfer module selects different levels of motion transfer methods based on the action type corresponding to the current text segment: For speech action segments, such as waving, introducing, emphasizing, guiding, welcoming, liking, bowing, etc., which have obvious body movements or gestures, the motion transfer module performs full-body or upper-body motion transfer. This process aligns the poses based on the pose sequence in the source motion template and the key point information in the target image, and generates a video segment in which the target image completes the corresponding action; For natural motion segments or transitional frame interpolation segments, since they do not contain obvious body movements and full-body motion transfer is costly, the system does not perform full-body motion transfer, but only performs head motion transfer, light head driving, or local natural micro-movements, so that the head, eyes, and subsequent lip shapes keep changing naturally in sync, while keeping the body area stable. Through this tiered strategy, the system only uses high-cost motion transfer in necessary action segments, and adopts low-cost generation methods in ordinary segments and transitional segments.
[0146] In the inference acceleration and post-processing synthesis stages, the generation system can further employ various acceleration strategies. For diffusion-based or iterative action generation models, the system can reduce the number of iteration steps through model distillation, fewer-step sampling, or consistency generation. For example, distillation methods such as LCM (Latent Consistency Models) or DMD (Distillation with Diffusion Models) can be used to compress the generation process. The system can also reduce inference overhead by employing engineering acceleration methods such as half-precision inference, ONNX (Open Neural Network Exchange) or TensorRT (a deep learning inference toolkit) compilation and deployment, fixed input size, and batch frame processing.
[0147] In addition, the motion transfer module can also use a combination of low-resolution generation and post-processing to initially generate low-resolution motion video clips, and then obtain motion video clips at the target resolution through super-resolution, scaling or local enhancement. This method can reduce the cost of generating a single frame while ensuring the continuity of motion. Finally, the motion transfer module outputs the motion video clips (i.e., the initial video clips) corresponding to the target digital human for subsequent processing by the transition template combination module, lip-sync module and video stitching module.
[0148] As can be seen from the above, the action transfer module, through feature caching mechanism, hierarchical action transfer strategy, and targeted pose-driven generation, can achieve the following technical effects:
[0149] (1) Reduce inference latency and computation cost: By extracting target image features and template features offline and caching them to a preset cache layer, the repeated feature extraction of the same user image and fixed template during the same video generation process is avoided. This cache strategy of extracting once and reusing multiple times greatly reduces the computational overhead during online inference and improves the overall efficiency of long text video generation.
[0150] (2) Achieving a balance between low cost and high quality in motion generation: The motion transfer module introduces a hierarchical motion transfer strategy. For speech motion templates containing obvious gestures, a high-cost human motion transfer strategy is adopted. Through posture alignment and preset motion transfer algorithms, videos containing complete body movements and gestures are generated, ensuring the expressiveness of key expressive segments. For natural motion templates and transitional motion templates, a low-cost face motion transfer strategy is adopted, focusing only on head and facial changes, avoiding unnecessary full-body motion transfer calculations. This differentiated processing ensures the overall richness of the video's motion while effectively reducing computational power consumption, making it particularly suitable for the presence of a large number of bland segments in long text scenarios.
[0151] (3) Ensuring the accuracy and smoothness of motion generation in the video: In the human motion transfer strategy, video frame sequences are generated based on target pose data (derived from the alignment of the source image's motion pose sequence with the target human body's key points) and target features, ensuring that the target image can accurately reproduce the body pose and gesture position of the source motion. Subsequent enhancement operations further improve the resolution and visual quality of the generated video, avoiding blurring or artifacts caused by direct generation, and ensuring high-quality output of the initial video clips.
[0152] (4) Improve system resource utilization: By using a preset cache layer to store static or semi-static feature data (such as template features and target image features), the storage and computation are decoupled, enabling the system to cope more flexibly with concurrent requests and long text generation tasks, thereby improving the system's scalability and resource utilization efficiency.
[0153] Optionally, in the execution and generation layer, the generation system calls the transition template combination module to compare the video duration and audio duration corresponding to each action segment.
[0154] Specifically, when the transition template combination module detects that the video duration corresponding to the i-th action segment is greater than or equal to the audio duration corresponding to the i-th action segment, firstly, the initial video segment corresponding to the i-th action segment is truncated based on the audio duration corresponding to the i-th action segment to obtain a retained video segment and a discarded video segment corresponding to the i-th action segment. Then, if the video duration of the retained video segment is greater than or equal to the video duration of the discarded video segment, the discarded video segment played in the forward direction is used as the stabilized video segment corresponding to the i-th action segment. If the video duration of the retained video segment is less than the video duration of the discarded video segment, the discarded video segment played in the reverse direction is used as the stabilized video segment corresponding to the i-th action segment. Then, the truncated video segment and the stabilized segment corresponding to the i-th action segment are spliced together to obtain the intermediate video segment corresponding to the i-th action segment. A silent segment is generated based on the video duration of the stabilized video segment and spliced to the end of the audio segment corresponding to the i-th action segment.
[0155] Specifically, when the transition template combination module detects that the video duration corresponding to the i-th action segment is less than the audio duration corresponding to the i-th action segment, the difference between the audio duration and the video duration corresponding to the i-th action segment is taken as the remaining duration. Then, the remaining duration corresponding to the i-th action segment is rounded up to obtain the target remaining duration corresponding to the i-th action segment. Then, if the user-set digital human type is a built-in type, L transition action templates are retrieved from the action template set corresponding to the user-selected built-in digital human image based on the target remaining duration corresponding to the i-th action segment, where L is a positive integer. If the user-set digital human type is an uploaded type, L transition action templates are retrieved from the action template set corresponding to the source image based on the target remaining duration corresponding to the i-th action segment. Subsequently, a transition video segment is generated based on the L transition action templates corresponding to the i-th action segment, where the video duration of the transition video segment is equal to the target remaining duration. Finally, the initial video segment and the transition video segment corresponding to the i-th action segment are spliced together to obtain the intermediate video segment corresponding to the i-th action segment.
[0156] Figure 6 This is a flowchart of an optional transition template combination module according to an embodiment of this application. The transition template combination module is used to solve the problem of inconsistency between the duration of action video segments and the corresponding audio duration, and to ensure that adjacent action segments can be smoothly connected in a unified state, such as... Figure 6As shown, the transition template combination module first receives the audio duration information output by the TTS and time alignment module, as well as the video segment duration information output by the built-in template calling module or motion transfer module. Then, it compares the duration of the current video segment with the duration of the corresponding audio segment, and performs truncation-stabilization processing or transition supplementation processing based on the comparison result.
[0157] Specifically, when the video segment duration is greater than or equal to the audio duration, the system truncates the video segment according to the audio duration. Since the first and last frames of the action template are set to a unified initial state, if the current truncated frame is not in the unified initial state after truncation, the system further determines the distance between the truncated frame and the start and end frames of the action template. If the truncated frame is closer to the start frame, a reverse playback method is used to quickly revert the action from the truncated frame to the start frame; if the truncated frame is closer to the end frame, forward playback continues to the end frame. The above stabilization process can be configured with a silent segment so that the stabilization action does not affect the original audio content. In this way, the system can still return the digital human to a unified initial state after the action is truncated, avoiding posture jumps when splicing with the next action segment.
[0158] Specifically, when the video clip duration is shorter than the audio duration, the generation system calculates the remaining duration and uses transition templates to fill it in. Each built-in digital human character is pre-configured with transition action templates of different specifications, such as 1 second, 2 seconds, and 5 seconds. The first and last frames of each transition template are in a unified initial state. The system selects and combines different specifications of transition templates according to the remaining duration. For example, when there are 7 seconds left, a 5-second transition template and a 2-second transition template can be combined. When there are 8 seconds left, a 5-second, 2-second, and 1-second transition template can be combined. When the remaining audio duration is not an integer number of seconds, the generation system can add silence at the end of the audio clip or perform slight duration alignment processing on the transition clip to convert the remaining duration into an integer number of seconds that can be obtained by combining 1-second, 2-second, and 5-second templates. In this way, the system can combine a small number of transition templates to create supplementary clips of different lengths without regenerating long action videos.
[0159] For the built-in digital human path, the transition template combination module directly calls the transition template corresponding to the current built-in image and splices it after the action template to make up for the remaining audio duration. Since both the transition template and the action template return to the same initial state, the next action segment can be stably connected after the segment is completed.
[0160] For the digital human path of the user-uploaded image, the transition template combination module prioritizes calling the transition template corresponding to the source image determined by the similar image matching module. Considering that the cost of full body motion transfer is high, during the transition frame interpolation stage, the system only performs head motion transfer or head lightweight drive on the combined transition template, and does not perform full body or upper body motion transfer, so that the head, eyes and subsequent lip shape keep changing naturally in sync, while keeping the body area stable, thereby reducing the inference cost.
[0161] As described above, the transition template combination module ultimately outputs a video segment sequence that matches the duration of the current audio segment. This includes the original action segment, necessary stabilization segments, and supplementary transition template segments. Through this mechanism, the transition template combination module can simultaneously achieve the alignment of action video and audio durations, the stabilization of the state after action truncation, and the stable splicing between multiple segments. The functions of the transition template combination module are as follows:
[0162] (1) Solving the problem of splicing abrupt transitions when the video duration is longer than the audio duration: When the initial video segment is longer than the audio, the solution is not to simply truncate it, but to introduce a stabilization mechanism. By comparing the size of the retained part and the discarded part, the discarded part is selected for forward or reverse playback as the stabilization video segment, so that the digital human's posture smoothly transitions to a unified initial state. This design avoids abrupt changes in action or hard cuts in the picture caused by truncation, and ensures the stability of the posture at the end of the segment, laying the foundation for the smooth connection of subsequent segments. At the same time, by generating a silent segment and splicing it to the end of the audio, strict synchronization of audio and video duration is achieved, solving the problem of audio and video length mismatch.
[0163] (2) Solving the problem of missing actions when the video duration is shorter than the audio duration: When the initial video segment is shorter than the audio, the solution uses a combination of transitional action templates to fill in the gaps. The remaining duration of the target is determined by rounding up, and L transitional action templates are retrieved from the action template set and spliced together. Since the transitional templates are also set to the same initial state at the beginning and end, this frame interpolation method not only fills in the time gaps, but also maintains the continuity of the digital human's posture, avoiding the static or abrupt switching of the scene caused by the premature end of the action.
[0164] (3) Compatible with both built-in and uploaded digital human types, reducing generation costs: The transition template combination module has designed template query paths for both built-in and uploaded types. For built-in types, it directly calls the preset transition template; for uploaded types, it calls the transition template of the source image. This design ensures that transition materials matching the image characteristics can be obtained regardless of the source of the digital human. Especially in the uploaded type, by reusing the transition template of the source image, the high-cost reasoning of generating long-term transition actions for each user image is avoided. Only a small number of standard template combinations are needed to achieve duration adaptation, thereby reducing the consumption of computing resources.
[0165] (4) Achieve smooth connection and consistency between multiple segments: Whether through stabilization or transition template combination, the final generated intermediate video segments ensure that they point to a unified initial state or a smooth transition state, thereby ensuring that the switching between adjacent action segments is natural and coherent in the long text broadcast, eliminating the phenomenon of action jumps and discontinuous images, and improving the overall viewing experience and professionalism of the final output video.
[0166] Optionally, in the compositing and output layer, a system call to the lip-sync module is generated. This module synchronizes the lip movements of the characters in the video after the video clip and corresponding audio clip have been aligned in duration. The lip-sync module first receives the video clip processed by built-in template calls, motion transfer, or transition template combinations, and the corresponding audio clip output by the TTS and time alignment module. It then confirms whether the video clip duration matches the audio clip duration. If they are aligned, it calls the video digital human lip-sync model to drive the lip movements of the face or mouth area in the video, ensuring the lip movements are synchronized with the audio content. Built-in digital human video clips, user-uploaded image motion transfer clips, and transition frame interpolation clips can all undergo unified lip-sync processing through the lip-sync module. The lip-sync module outputs a video clip with lip movements matching the audio and sends it to the subsequent video splicing module.
[0167] Optionally, in the compositing and output layer, a system call video stitching module is generated. This module stitches multiple lip-synced action segments, natural action segments, and transition segments into a complete video in chronological order. To avoid abrupt changes in action, posture jumps, or discontinuities during segment transitions, the system can perform smoothing at segment boundaries, such as boundary frame fusion, short fade-in / fade-out, head posture smoothing, and silent segment alignment. In this invention, both the action template and the transition template can be set to a unified initial state, meaning that the start and end states of each segment return to the same standard posture. Through this design, the video stitching module can complete segment connection based on a unified state, thereby improving the continuity and stability of long, multi-action videos.
[0168] Optionally, in the synthesis and output layer, a system call output module is generated. This module encapsulates the spliced video, audio, subtitles, and background into the final AI lecturer digital human video. This module can output video files in different formats according to application requirements, and can also push the generated results to AI lecturer platforms, training platforms, publicity systems, or other business systems. In the final output video, the digital human can perform corresponding actions based on different text fragments, maintaining coordination and consistency between speech, lip movements, body movements, and fragment splicing.
[0169] In the above embodiments, users can manually specify actions via action tags, or the large model can automatically match actions. Optionally, action specification can also be achieved through front-end button selection, timeline dragging, drop-down menus, voice commands, or text input of action names. AI action matching is not limited to large models and can be replaced by keyword rule matching, text classification models, semantic vector retrieval models, or action semantic tag matching models.
[0170] In the above embodiments, the preset action library is organized using digital human figures as the basic unit. Optionally, the action library can also be organized according to action type, explanation scenario, industry template, character style, or action intensity. Action templates are not limited to video files, but can also be stored in the form of human keypoint sequences, pose sequences, SMPL / SMPL-X parameters, skeletal animation files, alpha channel videos, or motion feature files.
[0171] In the above embodiments, user image preprocessing includes face detection, portrait matting, solid color background replacement, human keypoint detection, and image completion. Optionally, portrait matting can employ semantic segmentation, instance segmentation, matting (foreground extraction), or SAM (Segment Anything Model) segmentation models; human structure detection can employ 2D keypoints, 3D keypoints, DensePose (dense human pose estimation), SMPL, or SMPL-X reconstruction, etc. When an image does not meet the action transfer conditions, inpainting can be substituted for image completion through image expansion, standard body template stitching, human body generation, or manual selection of half-body templates.
[0172] In the above embodiments, similar image matching can be completed based on gender, facial similarity, and SMPL body shape parameters. Optionally, similar image matching can also be performed by combining features such as clothing color, clothing type, hairstyle, face orientation, shoulder width, composition ratio, skin color, and age group. Similarity calculation methods can employ CLIP (Contrastive Language–Image Pre-training), ArcFace (Additive Angular Margin Loss for Deep Face Recognition), visual Transformer, multimodal retrieval model, pose matching model, or learned ranking model.
[0173] In the above embodiments, the speech action segments perform full-body or upper-body motion transfer, while natural movements and transition segments employ lightweight head transfer. Optionally, motion transfer can employ pose-driven video generation, keypoint-driven generation, skeleton-driven generation, SMPL / SMPL-X-driven generation, video-to-video motion transfer, image-to-video generation, or 3D digital human rendering. The hierarchical strategy can also be extended to multiple levels such as full-body transfer, half-body transfer, arm transfer, head transfer, lip-sync only, or static retention.
[0174] In the above embodiments, the remaining duration is supplemented by a combination of 1-second, 2-second, and 5-second transition templates, and the splicing stability is ensured by truncation stabilization. Optionally, the transition template can be set to any specification such as 0.5 seconds, 1 second, 3 seconds, 5 seconds, and 10 seconds. Duration adaptation and state stabilization can also be achieved by using methods such as looping natural motion, variable speed templates, frame interpolation, time resampling, posture interpolation, motion compression, or preset stabilization segments.
[0175] In summary, this application can achieve the following technical effects:
[0176] (1) This application divides long text into structured segments with action information through action tag parsing and AI action matching mechanism, so that each text segment can correspond to a clear action, natural action or transitional action. Therefore, compared with the image digital human solution that only generates head movement and lip movement based on audio, this application can realize the corresponding control between text content, voice segments and digital human actions, and improve the controllability of actions and the richness of expression in the AI lecturer's explanation process.
[0177] (2) This application maintains the built-in digital human's motion templates, natural motion templates, and transition templates in advance through the motion library management module. When the digital human is the built-in image of the system, the system can directly call the preset video templates without having to regenerate full-body motion or perform large model image-to-video reasoning. Therefore, this application can reduce the computational load of video generation and improve the generation efficiency and image stability of the built-in digital human's long text video.
[0178] (3) This application sets up a motion transfer image preprocessing module in the user-uploaded image path. Through face detection, portrait cutout, solid color background replacement, human body key point detection and necessary image completion, the non-standard images uploaded by users are converted into standardized human images suitable for motion transfer. Therefore, this application can reduce motion transfer failure, body deformation and boundary abnormality caused by incomplete head, missing upper body, invisible arms or complex background.
[0179] (4) This application selects a source image that is closer to the user-uploaded image in terms of facial features, gender and body shape parameters from the built-in image library through the similar image matching module, and calls the corresponding action template of the source image for migration. Therefore, compared with the method of randomly selecting the source action template, this application can improve the matching degree between the source image and the target image in terms of appearance and human body structure, and reduce the problems of posture misalignment, arm deformation and inconsistency in proportion during the action migration process.
[0180] (5) This application adopts a hierarchical motion transfer strategy, performing full-body or upper-body motion transfer only in speech motion segments, and using head motion transfer or low-cost processing methods in natural motion segments and transition frame interpolation segments. Therefore, this invention can reduce unnecessary full motion transfer calculations while retaining the necessary body movements and gesture expression capabilities, thereby reducing the inference cost of generating long text videos.
[0181] (6) This application uses a combination mechanism of 1-second, 2-second, and 5-second transition templates to compensate for situations where the action template is shorter than the audio. Simultaneously, when the action template is longer than the audio, the distance between the truncated frame and the start and end frames is used to determine whether to play back in reverse to the start frame or play forward to the end frame, thus returning the action to a unified initial state. Therefore, this application can solve the problems of motion sluggishness, screen pauses, and splicing jumps caused by the inconsistency between the action template duration and the audio duration.
[0182] (7) This application sets the first and last frames of the action template, natural action template and transition template to a unified initial state, so that different action segments are spliced on the basis of the same posture. Therefore, this application can reduce the sudden changes in posture, body jumps and discontinuity of the picture when multiple action segments of long text are generated continuously, and improve the continuity and stability of the final AI lecturer video.
[0183] (8) This application calls the lip-sync model after the video clip and audio clip are aligned in duration, so that the built-in template clip, motion transfer clip and transition clip can all be synthesized in a unified manner. Therefore, this application can ensure the consistency of the digital human’s speech, lip shape and motion clips on the timeline, and improve the naturalness and usability of the final video.
[0184] Example 2
[0185] This application embodiment can also provide a digital human audio and video generation device. It should be noted that the digital human audio and video generation device of this application embodiment can be used to execute the digital human audio and video generation method provided in this application embodiment. The following is a description of the digital human audio and video generation device provided in this application embodiment.
[0186] According to an embodiment of this application, an apparatus for implementing the above-described digital human audio and video generation method is also provided. Figure 7 This is a schematic diagram of an optional digital human audio / video generation apparatus according to an embodiment of this application, such as... Figure 7 As shown, the device includes: a first generation unit 701, a second generation unit 702, a third generation unit 703, a first processing unit 704, a second processing unit 705, and a fourth generation unit 706.
[0187] Optionally, the first generation unit 701 is used to generate a sequence of action segments based on the target text and action pattern input by the user, wherein each action segment in the action segment sequence includes at least a segment number, text content, action code, and action name; the second generation unit 702 is used to generate an audio segment corresponding to each action segment based on each action segment in the action segment sequence, and to detect the audio duration of each audio segment; the third generation unit 703 is used to generate an initial video segment corresponding to each action segment based on the digital human type set by the user, and to detect the video duration of each initial video segment; the first processing unit 704 is used to process each action segment... The first processing unit 705 performs duration alignment processing on the initial video segment corresponding to each action segment based on the audio and video durations of the segment, resulting in an intermediate video segment with the same video duration as the corresponding audio duration. The second processing unit 705 performs lip-sync processing on the intermediate video segment corresponding to each action segment based on the audio segment corresponding to each action segment, resulting in a target video segment with the digital human's lip movements synchronized with the corresponding audio content. The fourth generation unit 706 generates a target audio-visual file based on the target video segment, audio segment, preset subtitles, and preset background corresponding to each action segment, wherein the target audio-visual file is used to express the target text through digital human speech.
[0188] Optionally, the first generation unit 701 includes: a first detection subunit, a first matching subunit, a first query subunit, and a first generation subunit.
[0189] Specifically, the first detection subunit is used to detect the user-input action pattern, wherein the action pattern is one of the following types: natural action pattern, manually inserted action pattern, and intelligent matching pattern; the first matching subunit is used to match the semantic features of each text fragment obtained by dividing the target text into natural action pattern / intelligent matching pattern in a preset action library to obtain the index information of the action template corresponding to each text fragment, wherein the index information of each action template includes at least the action code and the action name; the first query subunit is used to extract the pre-inserted action tag from the beginning of each text fragment in the target text when the action pattern is manually inserted action pattern, and query the preset action library based on each action tag to obtain the index information of the action template corresponding to each text fragment; the first generation subunit is used to generate an action fragment sequence based on the text content corresponding to each text fragment and the index information of the action template.
[0190] Optionally, the second generation unit 702 includes: a synthesis subunit, a second detection subunit, and a first determination subunit.
[0191] Specifically, the synthesis subunit is used to perform speech synthesis operations based on the text content in each action segment according to the order corresponding to the action segment sequence, so as to obtain the audio segment corresponding to each action segment; the second detection subunit is used to detect the start time and end time of each audio segment in the complete audio corresponding to all action segments; and the first determination subunit is used to determine the audio duration of each audio segment based on the start time and end time of each audio segment.
[0192] Optionally, the third generation unit 703 includes: a second query subunit, a third query subunit, and a second generation subunit.
[0193] Specifically, the second query subunit is used to query the preset action library to obtain the set of action templates corresponding to the user-selected built-in digital human image when the digital human type set by the user is the built-in type; the third query subunit is used to query the set of action templates corresponding to the built-in digital human image based on the action code in each action segment to obtain the action template corresponding to each action segment; the second generation subunit is used to generate the initial video segment corresponding to each action segment based on the action template corresponding to each action segment.
[0194] Optionally, the third generation unit 703 further includes: a preprocessing subunit, an extraction subunit, a second matching subunit, and a migration subunit.
[0195] Specifically, the preprocessing subunit is used to preprocess the user-uploaded digital human image when the user-set digital human type is upload type, to obtain a target image that meets the motion transfer conditions, wherein the motion transfer conditions are used to detect the completeness of the digital human image in the image; the extraction subunit is used to extract the auxiliary data corresponding to the target image, wherein the auxiliary data includes face bounding box position, human mask, human body key points and face key points; the second matching subunit is used to perform matching based on the target image in the built-in digital human image library, and use the matched built-in digital human image as the source image corresponding to the target image; the transfer subunit is used to perform motion transfer based on the auxiliary data corresponding to the target image and the motion template set corresponding to the source image, to obtain the initial video segment corresponding to each motion segment.
[0196] Optionally, the preprocessing subunit includes: a standardization module, a first determination module, a second determination module, a replacement module, and an update module.
[0197] Specifically, the standardization module is used to standardize the digital human image uploaded by the user to obtain a standardized image, wherein the standardization process is used to adjust the size of the digital human image based on coordinate mapping information; the first determination module is used to identify the single face as the target face when the standardized image includes a single face; the second determination module is used to identify the face located in the center region of the standardized image with the largest area as the target face when the standardized image includes two or more faces; the replacement module is used to perform image cutout and background replacement operations on the person to whom the target face belongs in the standardized image to obtain an initial image; and the update module is used to update the digital human image in the initial image based on an image completion algorithm to obtain the target image when the integrity of the digital human image in the initial image does not meet the action transfer conditions.
[0198] Optionally, the second matching subunit includes: a first detection module, a third determination module, a second detection module, a weighted summation module, and a fourth determination module.
[0199] Specifically, the first detection module is used to detect the gender of the target digital human image in the target image and the gender of each built-in digital human image in the built-in digital human image library; the third determination module is used to select the set of built-in digital humans with the same gender as the target digital human image as the candidate set; the second detection module is used to detect the facial similarity and body shape similarity between each built-in digital human image in the candidate set and the target image; the weighted summation module is used to perform weighted summation of facial similarity and body shape similarity to obtain the target similarity between each built-in digital human image and the target image; the fourth determination module is used to select the built-in digital human image with the highest target similarity in the candidate set as the source image corresponding to the target image.
[0200] Optionally, the migration subunit includes: an extraction module, a first query module, a second query module, a fifth determination module, and a migration module.
[0201] Specifically, the extraction module is used to extract features from the target image through the motion transfer module when the auxiliary data of the target image first enters the motion transfer module, obtaining feature data, and caching the feature data and auxiliary data as target features corresponding to the target image in a preset cache layer. The feature data includes target image features, target image encoding features, and latent variable features. The first query module is used to query the set of motion templates corresponding to the source image based on the motion encoding in each motion segment, obtaining the motion template corresponding to each motion segment. The second query module is used to query the preset cache layer to obtain the motion template corresponding to each motion segment that has been pre-cached offline. The template features of the board include action posture sequences, key point data, person masks, action trajectory features, and posture guidance features; the fifth determination module is used to determine the action transfer strategy corresponding to each action segment based on the action template corresponding to each action segment, wherein the action template is one of the following types: speech action template, natural action template, transition action template, and the action transfer strategy is one of the following types: face action transfer strategy, human body action transfer strategy; the transfer module is used to perform action transfer based on the action transfer strategy corresponding to each action segment, the template features, and the target features corresponding to the target image, to obtain the initial video segment corresponding to each action segment.
[0202] Optionally, the fifth determining module includes: a first determining submodule and a second determining submodule.
[0203] Specifically, the first determining submodule is used to determine the action transfer strategy corresponding to the i-th action segment as a human action transfer strategy when the action template corresponding to the i-th action segment is a speech action template; the second determining submodule is used to determine the action transfer strategy corresponding to the i-th action segment as a face action transfer strategy when the action template corresponding to the i-th action segment is a natural action template / transitional action template.
[0204] Optionally, the migration module includes: an alignment submodule, a generation submodule, and an enhancement submodule.
[0205] Specifically, the alignment submodule is used to perform pose alignment based on the action pose sequence in the template features of the i-th action segment and the human body key points in the target features when the action transfer strategy corresponding to the i-th action segment is a human action transfer strategy, to obtain target pose data. The target pose data is used to characterize the body pose and gesture position of the target digital human image in each frame of the action pose sequence. The generation submodule is used to generate the video frame sequence corresponding to the i-th action segment based on the target pose data and the target features corresponding to the target image through a preset action transfer algorithm. The enhancement submodule is used to perform enhancement operations on the video frame sequence corresponding to the i-th action segment to obtain the initial video segment corresponding to the i-th action segment.
[0206] Optionally, the first processing unit 704 includes: a truncation subunit, a second determination subunit, a third determination subunit, and a first splicing subunit.
[0207] Specifically, the first subunit is used to truncate the initial video segment corresponding to the i-th action segment based on the audio duration of the i-th action segment when the video duration of the i-th action segment is greater than or equal to the audio duration of the i-th action segment, thus obtaining a retained video segment and a discarded video segment corresponding to the i-th action segment; the second determining subunit is used to use the discarded video segment played in the forward direction as the stabilized video segment corresponding to the i-th action segment when the video duration of the retained video segment is greater than or equal to the video duration of the discarded video segment; the third determining subunit is used to use the discarded video segment played in reverse direction as the stabilized video segment corresponding to the i-th action segment when the video duration of the retained video segment is less than the video duration of the discarded video segment; the first splicing subunit is used to splice the truncated video segment and the stabilized segment corresponding to the i-th action segment to obtain an intermediate video segment corresponding to the i-th action segment, and generate a silent segment based on the video duration of the stabilized video segment, splicing the silent segment to the end of the audio segment corresponding to the i-th action segment.
[0208] Optionally, the first processing unit 704 further includes: a fourth determining subunit, an integer subunit, a fourth query subunit, a fifth query subunit, a third generating subunit, and a second splicing subunit.
[0209] Specifically, the fourth determining subunit is used to determine the remaining duration as the difference between the audio duration and video duration of the i-th action segment when the video duration of the i-th action segment is less than the audio duration of the i-th action segment; the rounding subunit is used to round up the remaining duration of the i-th action segment to obtain the target remaining duration of the i-th action segment; and the fourth query subunit is used to query L transitions from the set of action templates corresponding to the user-selected built-in digital human image, based on the target remaining duration of the i-th action segment, when the user-set digital human type is a built-in type. The system comprises: an action template, where L is a positive integer; a fifth query subunit, used to query L transition action templates from the action template set corresponding to the source image based on the target remaining duration corresponding to the i-th action segment, when the user-set digital human type is the upload type; a third generation subunit, used to generate a transition video segment based on the L transition action templates corresponding to the i-th action segment, where the video duration of the transition video segment is equal to the target remaining duration; and a second splicing subunit, used to splice the initial video segment and the transition video segment corresponding to the i-th action segment to obtain the intermediate video segment corresponding to the i-th action segment.
[0210] It should be noted that the above-mentioned units 701 to 706 correspond to steps S101 to S106 in the method embodiment. The above-mentioned units and corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments.
[0211] Example 3
[0212] Embodiments of this application can also provide an electronic device. Figure 8 This is a structural block diagram of an electronic device according to an embodiment of this application, such as... Figure 8 As shown, the electronic device includes: one or more ( Figure 8 (Only one is shown) processor 802, memory 804, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0213] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned method for generating digital human audio and video.
[0214] The memory may include high-speed random access memory (RAM), and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, which can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks (LANs), mobile communication networks, and combinations thereof.
[0215] The processor can access information and applications stored in memory via a transmission device to perform the following steps: generating a sequence of action segments based on the target text and action pattern input by the user, wherein each action segment in the sequence includes at least a segment number, text content, action code, and action name; generating an audio segment corresponding to each action segment based on each action segment in the sequence, and detecting the audio duration of each audio segment; generating an initial video segment corresponding to each action segment based on the user-defined digital human type, and detecting the video duration of each initial video segment; performing duration alignment processing on the initial video segments corresponding to each action segment based on the audio duration and video duration of each action segment to obtain an intermediate video segment with the same video duration as the corresponding audio duration; performing lip-sync processing on the intermediate video segments corresponding to each action segment based on the audio segment of each action segment to obtain a target video segment with lip movements synchronized with the corresponding audio content; generating a target audio-visual file based on the target video segment, audio segment, preset subtitles, and preset background of each action segment, wherein the target audio-visual file is used to express the target text through digital human speech.
[0216] The processor can access information and applications stored in memory via a transmission device to perform the following steps: Detecting user-inputted action patterns, where the action pattern is one of the following types: natural action pattern, manually inserted action pattern, and intelligent matching pattern; In the case of natural action pattern / intelligent matching pattern, matching the semantic features of each text fragment obtained by dividing the target text in a preset action library to obtain the index information of the action template corresponding to each text fragment, wherein the index information of each action template includes at least the action code and action name; In the case of manually inserted action pattern, extracting pre-inserted action tags from the beginning of each text fragment in the target text, and querying the preset action library based on each action tag to obtain the index information of the action template corresponding to each text fragment; Generating an action fragment sequence based on the text content corresponding to each text fragment and the index information of the action template.
[0217] The processor can access the information and application programs stored in the memory via the transmission device to perform the following steps: according to the order corresponding to the action segment sequence, perform speech synthesis based on the text content in each action segment to obtain the audio segment corresponding to each action segment; detect the start time and end time of each audio segment in the complete audio corresponding to all action segments; determine the audio duration of each audio segment based on the start time and end time of each audio segment.
[0218] The processor can access the information and application stored in the memory via the transmission device to perform the following steps: if the user-set digital human type is a built-in type, query the preset action library to obtain the action template set corresponding to the user-selected built-in digital human image; query the action template set corresponding to the built-in digital human image based on the action code in each action segment to obtain the action template corresponding to each action segment; generate the initial video segment corresponding to each action segment based on the action template corresponding to each action segment.
[0219] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: If the user-set digital human type is upload type, preprocess the user-uploaded digital human image to obtain a target image that meets motion transfer conditions, where the motion transfer conditions are used to detect the completeness of the digital human image in the image; extract auxiliary data corresponding to the target image, including face bounding box position, person mask, human body key points, and facial key points; perform matching based on the target image in the built-in digital human image library, and use the matched built-in digital human image as the source image corresponding to the target image; perform motion transfer based on the auxiliary data corresponding to the target image and the motion template set corresponding to the source image to obtain the initial video segment corresponding to each motion segment.
[0220] The processor can access information and applications stored in memory via a transmission device to perform the following steps: standardizing the user-uploaded digital human image to obtain a standardized image, wherein the standardization process is used to adjust the size of the digital human image based on coordinate mapping information; if the standardized image includes a single face, designating that single face as the target face; if the standardized image includes two or more faces, designating the face located in the center of the standardized image and having the largest area as the target face; performing image cutout and background replacement operations on the person to whom the target face belongs in the standardized image to obtain an initial image; if the completeness of the digital human image in the initial image does not meet the action transfer conditions, updating the digital human image in the initial image based on an image completion algorithm to obtain the target image.
[0221] The processor can access the information and application programs stored in the memory via the transmission device to perform the following steps: detect the gender of the target digital human image in the target image and the gender of each built-in digital human image in the built-in digital human image library; select the set of built-in digital humans with the same gender as the target digital human image as the candidate set; detect the facial similarity and body shape similarity between each built-in digital human image in the candidate set and the target image; perform a weighted sum of the facial similarity and body shape similarity to obtain the target similarity between each built-in digital human image and the target image; and select the built-in digital human image with the highest target similarity in the candidate set as the source image corresponding to the target image.
[0222] The processor can access the information and application programs stored in the memory via the transmission device to execute the following steps: When the auxiliary data of the target image enters the motion transfer module for the first time, the motion transfer module extracts features from the target image to obtain feature data, and caches the feature data and auxiliary data as target features corresponding to the target image in a preset cache layer. The feature data includes target image features, target image encoding features, and latent variable features. Based on the motion encoding in each motion segment, the processor queries the motion template set corresponding to the source image to obtain the motion template corresponding to each motion segment. The processor queries the preset cache layer to obtain the template features of the motion template corresponding to each motion segment that is cached offline beforehand. The template features include motion posture sequence, key point data, person mask, motion trajectory features, and posture guidance features. Based on the motion template corresponding to each motion segment, the processor determines the motion transfer strategy corresponding to each motion segment. The motion template is one of the following types: speech motion template, natural motion template, transitional motion template, and the motion transfer strategy is one of the following types: face motion transfer strategy, human motion transfer strategy. Based on the motion transfer strategy, template features, and target features corresponding to the target image, the processor performs motion transfer to obtain the initial video segment corresponding to each motion segment.
[0223] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: if the action template corresponding to the i-th action segment is a speech action template, determine that the action transfer strategy corresponding to the i-th action segment is a human action transfer strategy; if the action template corresponding to the i-th action segment is a natural action template / transition action template, determine that the action transfer strategy corresponding to the i-th action segment is a face action transfer strategy.
[0224] The processor can access the information and application programs stored in the memory via the transmission device to execute the following steps: When the motion transfer strategy corresponding to the i-th motion segment is a human motion transfer strategy, perform pose alignment based on the motion posture sequence in the template features corresponding to the i-th motion segment and the human body key points in the target features to obtain target pose data. The target pose data is used to characterize the body posture and gesture position of the target digital human image in each frame of the motion posture sequence. Based on the target pose data and the target features corresponding to the target image, generate a video frame sequence corresponding to the i-th motion segment using a preset motion transfer algorithm. Perform enhancement operations on the video frame sequence corresponding to the i-th motion segment to obtain the initial video segment corresponding to the i-th motion segment.
[0225] The processor can access the information and application program stored in the memory via the transmission device to perform the following steps: If the video duration corresponding to the i-th action segment is greater than or equal to the audio duration corresponding to the i-th action segment, truncate the initial video segment corresponding to the i-th action segment based on the audio duration corresponding to the i-th action segment to obtain a retained video segment and a discarded video segment corresponding to the i-th action segment; if the video duration of the retained video segment is greater than or equal to the video duration of the discarded video segment, use the forward-played discarded video segment as the stabilized video segment corresponding to the i-th action segment; if the video duration of the retained video segment is less than the video duration of the discarded video segment, use the reverse-played discarded video segment as the stabilized video segment corresponding to the i-th action segment; concatenate the truncated video segment and the stabilized segment corresponding to the i-th action segment to obtain an intermediate video segment corresponding to the i-th action segment, and generate a silent segment based on the video duration of the stabilized video segment, concatenating the silent segment to the end of the audio segment corresponding to the i-th action segment.
[0226] The processor can access the information and application stored in the memory via the transmission device to execute the following steps: If the video duration corresponding to the i-th action segment is less than the audio duration corresponding to the i-th action segment, the difference between the audio duration and video duration corresponding to the i-th action segment is taken as the remaining duration; the remaining duration corresponding to the i-th action segment is rounded up to obtain the target remaining duration corresponding to the i-th action segment; if the user-set digital human type is a built-in type, L transitional action templates are retrieved from the action template set corresponding to the user-selected built-in digital human image based on the target remaining duration corresponding to the i-th action segment, where L is a positive integer; if the user-set digital human type is an uploaded type, L transitional action templates are retrieved from the action template set corresponding to the source image based on the target remaining duration corresponding to the i-th action segment; a transitional video segment is generated based on the L transitional action templates corresponding to the i-th action segment, where the video duration of the transitional video segment is equal to the target remaining duration; the initial video segment and the transitional video segment corresponding to the i-th action segment are spliced together to obtain the intermediate video segment corresponding to the i-th action segment.
[0227] This application provides a scheme for generating digital human audio and video. It employs an action segment sequence-driven approach, generating a structured sequence containing action codes by parsing the user-input target text and action patterns. Corresponding audio and initial video segments are then generated independently. Subsequently, the initial video segment is time-aligned based on the detected audio and video durations to ensure strict consistency between the intermediate video segment duration and the audio segment duration. Lip-sync processing is then applied to the intermediate video segment based on the audio segment. Finally, the target video segment, audio, subtitles, and background are integrated to generate the target audio and video file. This achieves precise synchronization and unification of text content, voice audio, digital lip movements, and body movements on the timeline, resulting in a digital human video where speech, lip movements, actions, and visuals remain coordinated, video segments exhibit no abrupt changes in action, and the visuals are continuous. This solves the technical problem of poor audio and video quality in digital humans generated using existing technologies.
[0228] Those skilled in the art will understand that Figure 8 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, PDAs, mobile internet devices, PADs, and other terminal devices. Figure 8 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 8 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 8 The different configurations shown.
[0229] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0230] Example 4
[0231] Embodiments of this application may also provide a storage medium.
[0232] Optionally, in this embodiment of the application, the storage medium can be used to store the program code executed by the digital human audio and video generation method provided in the above method embodiment.
[0233] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0234] This application also provides a computer program product, which, when executed on a data processing device, is adapted to perform the steps of a method for generating digital human audio and video.
[0235] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0236] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0237] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0238] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0239] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0240] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0241] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for generating digital human audio and video, characterized in that, include: An action segment sequence is generated based on the target text and action pattern input by the user, wherein each action segment in the action segment sequence includes at least a segment number, text content, action code, and action name; Based on each action segment in the action segment sequence, an audio segment corresponding to each action segment is generated, and the audio duration of each audio segment is detected; Based on the user-defined digital human type, an initial video segment corresponding to each action segment is generated, and the video duration of each initial video segment is detected; Based on the audio duration and video duration corresponding to each action segment, the initial video segment corresponding to each action segment is time-aligned to obtain an intermediate video segment with the same video duration as the corresponding audio duration. Based on the audio segment corresponding to each action segment, lip-sync processing is performed on the intermediate video segment corresponding to each action segment to obtain the target video segment in which the digital lip movements are synchronized with the corresponding audio content. A target audio-visual file is generated based on the target video clip, audio clip, preset subtitles, and preset background corresponding to each action segment, wherein the target audio-visual file is used to express the target text through a digital human speech.
2. The method for generating digital human audio and video according to claim 1, characterized in that, Generate a sequence of action segments based on the target text and action patterns input by the user, including: The action pattern of the user input is detected, wherein the action pattern is one of the following types: natural action pattern, manually inserted action pattern, and intelligent matching pattern; When the action mode is the natural action mode / intelligent matching mode, the semantic features of each text fragment obtained by dividing the target text are matched in a preset action library to obtain the index information of the action template corresponding to each text fragment. The index information of each action template includes at least the action code and the action name. When the action mode is the manual insertion action mode, the pre-inserted action tags are extracted from the beginning of each text segment in the target text, and the index information of the action template corresponding to each text segment is obtained by querying the preset action library based on each action tag. The sequence of action segments is generated based on the index information of the text content and action template corresponding to each text segment.
3. The method for generating digital human audio and video according to claim 1, characterized in that, Based on each action segment in the action segment sequence, an audio segment corresponding to each action segment is generated, and the audio duration of each audio segment is detected, including: According to the order corresponding to the action segment sequence, a speech synthesis operation is performed based on the text content in each action segment to obtain the audio segment corresponding to each action segment; Detect the start and end times of each audio segment in the complete audio corresponding to all action segments; The duration of each audio segment is determined based on its start and end times.
4. The method for generating digital human audio and video according to claim 1, characterized in that, Based on the user-defined digital human type, an initial video clip is generated for each action segment, including: If the user-defined digital human type is a built-in type, the set of action templates corresponding to the user-selected built-in digital human image is retrieved from the preset action library. Based on the action code in each action segment, the action template corresponding to the built-in digital human image is queried in the action template set to obtain the action template corresponding to each action segment; An initial video segment corresponding to each action segment is generated based on the action template corresponding to each action segment.
5. The method for generating digital human audio and video according to claim 1, characterized in that, Based on the user-defined digital human type, an initial video clip is generated for each action segment, including: When the user sets the digital human type to the upload type, the user-uploaded digital human image is preprocessed to obtain a target image that meets the motion transfer conditions, wherein the motion transfer conditions are used to detect the completeness of the digital human image in the image; Extract the auxiliary data corresponding to the target image, wherein the auxiliary data includes the face bounding box position, the person mask, the human body key points, and the face key points; The target image is matched against the built-in digital human image library, and the matched built-in digital human image is used as the source image corresponding to the target image. Based on the auxiliary data corresponding to the target image and the action template set corresponding to the source image, motion transfer is performed to obtain the initial video segment corresponding to each action segment.
6. The method for generating digital human audio and video according to claim 5, characterized in that, The digital human image uploaded by the user is preprocessed to obtain a target image that meets the motion transfer conditions, including: The digital human image uploaded by the user is standardized to obtain a standardized image, wherein the standardization process is used to adjust the size of the digital human image based on coordinate mapping information; In the case where the standardized image includes a single face, that single face is taken as the target face; When the standardized image includes two or more faces, the face located in the central region of the standardized image and having the largest area is taken as the target face. In the standardized image, the target face is cut out and the background is replaced to obtain the initial image; If the completeness of the digital human image in the initial image does not meet the motion transfer condition, the digital human image in the initial image is updated based on the image completion algorithm to obtain the target image.
7. The method for generating digital human audio and video according to claim 5, characterized in that, Matching is performed on the target image within a built-in digital human image library, and the matched built-in digital human image is used as the source image corresponding to the target image, including: Detect the gender of the target digital human image in the target image and the gender of each built-in digital human image in the built-in digital human image library; The set of built-in digital humans with the same gender as the target digital human image is used as the candidate set; Detect the facial similarity and body shape similarity between each built-in digital human image in the candidate set and the target image; The face similarity and body shape similarity are weighted and summed to obtain the target similarity between each built-in digital human image and the target image; The built-in digital human image with the highest target similarity in the candidate set is taken as the source image corresponding to the target image.
8. The method for generating digital human audio and video according to claim 5, characterized in that, Based on the auxiliary data corresponding to the target image and the motion template set corresponding to the source image, motion transfer is performed to obtain the initial video segment corresponding to each motion segment, including: When the auxiliary data of the target image enters the action transfer module for the first time, the action transfer module performs feature extraction on the target image to obtain feature data, and caches the feature data and the auxiliary data as target features corresponding to the target image in a preset cache layer. The feature data includes target image features, target image encoding features, and latent variable features. Based on the action code in each action segment, a query is performed in the action template set corresponding to the source image to obtain the action template corresponding to each action segment; The template features of the action template corresponding to each action segment, which is pre-cached offline, are obtained by querying the preset cache layer. The template features include action posture sequence, key point data, character mask, action trajectory features, and posture guidance features. Based on the action template corresponding to each action segment, an action transfer strategy corresponding to each action segment is determined, wherein the action template is one of the following types: speech action template, natural action template, transition action template, and the action transfer strategy is one of the following types: face action transfer strategy, human body action transfer strategy. Based on the motion transfer strategy, template features, and target features corresponding to each motion segment, motion transfer is performed to obtain the initial video segment corresponding to each motion segment.
9. The method for generating digital human audio and video according to claim 8, characterized in that, Based on the action template corresponding to each action segment, determine the action transfer strategy corresponding to each action segment, including: If the action template corresponding to the i-th action segment is the speech action template, then the action transfer strategy corresponding to the i-th action segment is determined to be the human action transfer strategy. If the action template corresponding to the i-th action segment is the natural action template / the transitional action template, then the action transfer strategy corresponding to the i-th action segment is determined to be the face action transfer strategy.
10. The method for generating digital human audio and video according to claim 8, characterized in that, Based on the motion transfer strategy, template features, and target features corresponding to each motion segment, motion transfer is performed to obtain an initial video segment corresponding to each motion segment, including: When the motion transfer strategy corresponding to the i-th motion segment is the human motion transfer strategy, the pose alignment is performed based on the motion pose sequence in the template features corresponding to the i-th motion segment and the human key points in the target features to obtain target pose data. The target pose data is used to characterize the body pose and gesture position of the target digital human image in each frame of the motion pose sequence. Based on the target pose data and the target features corresponding to the target image, a video frame sequence corresponding to the i-th action segment is generated by a preset motion transfer algorithm; An enhancement operation is performed on the video frame sequence corresponding to the i-th action segment to obtain the initial video segment corresponding to the i-th action segment.
11. The method for generating digital human audio and video according to claim 1, characterized in that, Based on the audio and video durations corresponding to each action segment, the initial video segment corresponding to each action segment is time-aligned to obtain an intermediate video segment with the same video duration as the corresponding audio duration, including: If the video duration corresponding to the i-th action segment is greater than or equal to the audio duration corresponding to the i-th action segment, the initial video segment corresponding to the i-th action segment is truncated based on the audio duration corresponding to the i-th action segment to obtain the retained video segment and the discarded video segment corresponding to the i-th action segment. If the video duration of the retained video segment is greater than or equal to the video duration of the discarded video segment, the discarded video segment played in the forward direction will be used as the stabilization video segment corresponding to the i-th action segment. If the video duration of the retained video segment is less than the video duration of the discarded video segment, the discarded video segment played in reverse will be used as the stabilization video segment corresponding to the i-th action segment. The truncated video segment and the stabilized segment corresponding to the i-th action segment are spliced together to obtain the intermediate video segment corresponding to the i-th action segment. A silent segment is generated based on the video duration of the stabilized video segment, and the silent segment is spliced to the end of the audio segment corresponding to the i-th action segment.
12. The method for generating digital human audio and video according to claim 1, characterized in that, Based on the audio and video durations corresponding to each action segment, the initial video segment corresponding to each action segment is time-aligned to obtain an intermediate video segment with the same video duration as the corresponding audio duration, including: If the video duration corresponding to the i-th action segment is less than the audio duration corresponding to the i-th action segment, the difference between the audio duration and the video duration corresponding to the i-th action segment shall be taken as the remaining duration. The remaining duration corresponding to the i-th action segment is rounded up to obtain the target remaining duration corresponding to the i-th action segment; When the user-set digital human type is a built-in type, L transition action templates are obtained from the action template set corresponding to the user-selected built-in digital human image based on the target remaining time corresponding to the i-th action segment, where L is a positive integer; When the user-set digital human type is the upload type, L transitional action templates are obtained by querying the action template set corresponding to the source image based on the target remaining duration corresponding to the i-th action segment. A transition video segment is generated based on L transition action templates corresponding to the i-th action segment, wherein the video duration of the transition video segment is equal to the target remaining duration; The initial video segment and the transition video segment corresponding to the i-th action segment are spliced together to obtain the intermediate video segment corresponding to the i-th action segment.
13. A digital human audio-visual generation device, characterized in that, include: The first generation unit is used to generate a sequence of action segments based on the target text and action pattern input by the user, wherein each action segment in the sequence of action segments includes at least a segment number, text content, action code and action name; The second generation unit is used to generate an audio segment corresponding to each action segment based on each action segment in the action segment sequence, and to detect the audio duration of each audio segment; The third generation unit is used to generate an initial video segment corresponding to each action segment based on the digital human type set by the user, and to detect the video duration of each initial video segment; The first processing unit is used to perform duration alignment processing on the initial video segment corresponding to each action segment based on the audio duration and video duration corresponding to each action segment, so as to obtain an intermediate video segment with the same video duration and corresponding audio duration. The second processing unit is used to perform lip-sync processing on the intermediate video segment corresponding to each action segment based on the audio segment corresponding to each action segment, so as to obtain a target video segment in which the digital lip movements are synchronized with the corresponding audio content. The fourth generation unit is used to generate a target audio-visual file based on the target video segment, audio segment, preset subtitles and preset background corresponding to each action segment, wherein the target audio-visual file is used to express the target text through a digital human speech.
14. A computer program product, characterized in that, The computer program product includes a computer program, wherein, when the computer program is executed, it controls the computer program product to perform the digital human audio and video generation method according to any one of claims 1 to 12.
15. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the digital human audio and video generation method according to any one of claims 1 to 12.