Digital human commentary video generation method and device

By structuring PPT files and using dual-stream parallel technology, the problems of semantic and action disconnect and synchronization rigidity in digital human presentation videos were solved, generating high-quality videos with logical coherence and natural visuals, thus enhancing the realism and immersion of the videos.

CN121691848BActive Publication Date: 2026-06-02山东鲁商科技集团有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
山东鲁商科技集团有限公司
Filing Date
2026-02-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to synchronize semantics and actions in digital human narration videos, leading to a disconnect between semantics and actions and a rigid synchronization, which affects the naturalness and immersion of the video.

Method used

By extracting key information from PPT files, a structured script file is generated. This script is then processed in parallel with two streams, including background image sequences and audio and motion data from a personalized digital human model. Through bidirectional dynamic alignment technology, precise synchronization of the audio and motion data is achieved, ultimately generating a high-quality digital human presentation video.

Benefits of technology

It achieves precise synchronization of semantics and actions in digital human narration videos, generating high-quality videos that are logically coherent and visually natural, thus enhancing the realism and immersion of the videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121691848B_ABST
    Figure CN121691848B_ABST
Patent Text Reader

Abstract

The application discloses a digital person explanation video generation method and device, and relates to the technical field of computers. The method comprises the following steps: extracting key information from a PPT file to obtain a structured script file corresponding to the PPT file; performing double-flow parallel processing on the PPT file and a personalized digital person model based on the structured script file to obtain a background image sequence and audio data and action data corresponding to the personalized digital person model; performing segment division on the background image sequence to obtain a plurality of background image sequence segments; performing bidirectional dynamic alignment on the audio data and the action data corresponding to each background image sequence segment to obtain a digital person video segment corresponding to each background image sequence segment; and fusing the digital person video segments corresponding to each background image sequence segment to obtain a digital person explanation video corresponding to the PPT file. In this way, a more reasonable and natural PPT digital person explanation video can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method and device for generating digital human narration videos. Background Technology

[0002] With the popularization of online education and digital training, a large number of PPT courseware need to be converted into video format for dissemination and learning. This is typically achieved through manual screen recording and professional post-production, but these methods suffer from long production cycles and high costs. With the rapid development of AI technology, digital human technology is increasingly being applied to education, training, and other fields. However, in the video generation process, this technology struggles to establish a deep mapping relationship between the narration (audio stream) and animation nodes (video stream). Rendering engines cannot perceive pauses, changes in speech rate, and the semantic endings of sentences, leading to problems such as semantic and action disconnect (digital human body movements cannot automatically match appropriate body language based on the tone or specific semantics of the narration) and synchronization stiffness (for example, when a complex body movement is not yet completed, the narration may have already ended and the next page has been switched, resulting in abruptly cut-off actions and extremely unnatural visuals). Summary of the Invention

[0003] This application provides a method and device for generating digital human presentation videos to solve the following technical problem: how to effectively solve the problems of semantic and action separation and synchronization rigidity in digital human presentation videos, thereby generating more reasonable and natural PPT digital human presentation videos.

[0004] In a first aspect, embodiments of this application provide a method for generating digital human narration videos, the method comprising:

[0005] Key information is extracted from the PPT file to obtain the corresponding structured script file;

[0006] Based on the structured script file, the PPT file and the personalized digital human model are processed in parallel with two streams to obtain a background image sequence and audio and motion data corresponding to the personalized digital human model.

[0007] The background image sequence is segmented to obtain multiple background image sequence segments;

[0008] The audio data and motion data corresponding to each background image sequence segment are dynamically aligned bidirectionally to obtain a digital human video segment corresponding to each background image sequence segment.

[0009] The digital human video segments corresponding to each of the background image sequence segments are merged to obtain the digital human explanation video corresponding to the PPT file.

[0010] Secondly, embodiments of this application also provide a digital human narration video generation device, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0011] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions, which, when executed, implement the method described in any of the above claims.

[0012] The digital human narration video generation method, device, and medium provided in this application have the following beneficial effects:

[0013] First, by extracting key information from the PPT file, the unstructured PPT pages are transformed into structured data that machines can understand and process, resulting in a structured script file for the corresponding PPT file. This establishes a single data source and control center, laying the foundation for subsequent precise synchronization. Next, based on the structured script file, the PPT file and the personalized digital human model undergo dual-stream parallel processing, improving processing efficiency and significantly shortening the overall generation time. This yields a background image sequence and corresponding audio and motion data for the personalized digital human model. The background rendering of the PPT and the audio and motion of the digital human are both driven by the structured script file, ensuring that the data output from both channels is logically interconnected and consistent, thus guaranteeing the accuracy and consistency of content generation. Finally, the background image sequence is segmented, breaking down a large video generation task into multiple independent and manageable subtasks, achieving task modularization and decoupling. Multiple background image sequence segments are obtained, and each segment can be processed individually, greatly reducing subsequent modification costs and time. Then, the audio and motion data corresponding to each background image sequence segment are dynamically aligned bidirectionally. Based on the audio duration, speech rate, and stress, as well as the duration and semantics of the motion, an algorithm finds an optimal time matching scheme to achieve frame-level precise synchronization, resulting in digital human video segments corresponding to each background image sequence segment. This significantly enhances the realism and immersion of the video, thereby improving the naturalness of the digital human's performance. Finally, the digital human video segments corresponding to each background image sequence segment are merged. By splicing together independent video segments with synchronized content, a complete video file is generated to achieve seamless integration of the final product. This ensures that all segments are logically coherent and visually transition naturally, ultimately generating a high-quality digital human explanation video that meets user expectations. Attached Figure Description

[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0015] Figure 1 A flowchart of a digital human narration video generation method provided in this application embodiment;

[0016] Figure 2 This is a schematic diagram of the internal structure of a device provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] It is understood that in the embodiments of this application, data related to user information (such as user accounts) is involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0019] In the following description, the terms “first, second, ...” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first, second, ...” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0022] 1) Text to Speech (TTS): TTS technology is an artificial intelligence technology that can automatically convert text information on a computer into audio files that sound like human speech. Common TTS technologies include splicing synthesis (e.g., MBROLA), parametric synthesis (e.g., HTS), and deep learning synthesis (e.g., Tacotron). It can be widely used in fields such as intelligent assistants, public broadcasting, entertainment industry, content creation, and education.

[0023] 2) Digital Human Model: A digital human model is a virtual character with a realistic human appearance and behavior, realized through computer technology. It can simulate the interaction of real humans and is an important bridge connecting the virtual and real worlds. A complete digital human model usually includes key parts such as 3D modeling and rendering, motion capture and driving, artificial intelligence and interaction.

[0024] With the popularization of online education and digital training, a large number of PPT courseware need to be converted into video format for easy dissemination and learning. Traditional video production methods typically fall into two main categories:

[0025] 1) Manual screen recording: The instructor plays the PPT while explaining and recording the screen. This method requires a high level of expression from the instructor, and if a mistake is made, it needs to be re-recorded, which is difficult to modify later and has a long production cycle.

[0026] 2) Professional post-production: Using professional software such as Adobe Premiere and Final Cut Pro to edit and combine PPT images with audio. This method has a long production cycle, high cost, is difficult to mass-produce, and requires professional video editors.

[0027] In recent years, with the rapid development of AI technology, digital human technology has been gradually applied to education, training and other fields. In the process of video generation, it is difficult to establish a deep mapping relationship between the narration (audio stream) and the PPT animation nodes (video stream). The rendering engine cannot perceive the pauses, changes in speech rate and semantic endings of sentences, resulting in problems such as semantic and action separation (the digital human's body movements cannot automatically match appropriate body language according to the tone or specific semantics of the narration) and synchronization stiffness (for example, when a complex body movement has not been completed, the narration may have ended and the next page has been switched, resulting in the action being abruptly cut off and the picture being extremely unnatural).

[0028] Based on this, this application provides a method for generating digital human presentation videos, which can effectively solve the problems of semantic and action separation and synchronization rigidity in digital human presentation videos, thereby generating more reasonable and natural PPT digital human presentation videos.

[0029] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 This application provides a flowchart of a method for generating digital human presentation videos. This method can be applied to various scenarios involving the generation of PPT presentation videos, such as education and training (online course production, corporate internal training, etc.), corporate marketing and communication (product introductions and launches, automated generation of marketing content, etc.), and information services and public utilities (government information and policy interpretation, digital guides, etc.). Certain input parameters or intermediate results in the process can be manually adjusted to help improve accuracy.

[0031] This application provides a method for generating digital human narration videos. It should be noted that the execution entity in these embodiments can be a server or any terminal device with data processing capabilities. For example, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal device can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, in-vehicle terminal, etc., but is not limited to these.

[0032] like Figure 1 As shown in the figure, the digital human narration video generation method provided in this application embodiment specifically includes the following steps:

[0033] Step 101: Extract key information from the PPT file to obtain the structured script file corresponding to the PPT file.

[0034] It should be noted that a structured script file refers to a text file that organizes unstructured or semi-structured raw content (such as a PPT document) according to specific logical rules and format standards, forming a text file with a clear hierarchical structure and semantic markup.

[0035] In some embodiments, step 101 described above can be implemented as follows: the PPT file is decomposed into multimodal elements, and each decomposed element is digitized to obtain multiple structured data; multidimensional semantic analysis is performed on each structured data to obtain semantic text corresponding to each structured data; and each semantic text is fused and recombined to obtain a structured script file corresponding to the PPT file.

[0036] Thus, by breaking down PPT documents into multiple modal elements and digitizing them, the problem of traditional technologies being able to process only a single text modality while ignoring unstructured visual information can be solved. This ensures the completeness and accuracy of content parsing. Furthermore, by performing deep semantic analysis on each structured data point, uneditable pixel information is transformed into processable structured semantic text, filling the "visual blind spot" of traditional OCR technology. This ensures that the generated presentation script completely corresponds to the content of the presentation screen, avoiding information loss and content disconnect. In addition, by fusing and recombining semantic text, semantic information obtained from multiple dimensions is combined to obtain a structured script file corresponding to the PPT file, laying the foundation for accurate synchronization in the future.

[0037] As an example, suppose there is a PowerPoint presentation about cars. Breaking the presentation down into multiple modal elements yields four elements: title text, 3D rendering, performance comparison bar chart, and notes. Digitizing each element, we get the following structured data: Title text: { "type": "text", "subtype": "title", "content": "Ultimate Performance, Drive at Your Fingertips", "position": {...}, "style": {...}}; 3D rendering: { "type": "image", "id": "img_001", "data": "[cropped image data]", "position": {...}}; and performance comparison bar chart: { "type":"chart", "id": "img_002", "data": "[chart image data]", "extracted_semantic_data": { "chart_type": "bar", "entities": ["Car M", "Competitor A", "Competitor B"], "metrics": ["0-100 km / h acceleration", ... "Top Speed", "Range"], "conclusion": "Car M leads in many key performance indicators"}, "position": {...}}, The structured data corresponding to the note text is { "type": "text", "subtype": "note", "content": "Here we can emphasize our 800V supercharging technology, 5 minutes of charging provides 200 kilometers of range." Subsequently, multi-dimensional semantic analysis was performed on each structured data point. The semantic text corresponding to the title text was: "'Ultimate performance, at your fingertips' is a powerful and liberating slogan; the tone should be enthusiastic and confident." The semantic text corresponding to the 3D rendering image was: "This is the heart of the Automotive M—our newly developed 'Falcon' electric drive system, representing the industry's top engineering aesthetics and power efficiency." The semantic text corresponding to the performance comparison bar chart was: "Please look at this comparison chart; it's clear at a glance. Whether it's 0-100 km / h acceleration or overall range, the Automotive M overwhelmingly surpasses mainstream competitors in the market. The data doesn't lie; this is our most direct answer to 'ultimate performance.' Especially in terms of range, we have nearly 30% more than competitor A." The implied meaning text corresponding to the remarks text was: "Powerful performance cannot be separated from efficient energy replenishment."We were the first to adopt the industry-leading 800V supercharging platform, which truly enables 200 kilometers of charging in just 5 minutes, allowing you to completely say goodbye to range anxiety. Finally, the various semantic texts were merged and recombined to obtain the structured script file corresponding to the PPT file.

[0038] In some embodiments, the above-mentioned fusion and recombination of each semantic text to obtain a structured script file corresponding to the PPT file can be achieved in the following way: performing contextual understanding on each semantic text, and concatenating each semantic text based on the understood context content to obtain concatenated semantic text; performing sentiment annotation and action instruction generation on the concatenated semantic text to obtain a structured script file corresponding to the PPT file.

[0039] In this way, by giving AI a global perspective through contextual understanding, it can perceive the logical progression between pages, the emotional shifts between chapters, and the core argument of the entire PPT. This makes the subsequent splicing no longer a mechanical addition of sentences, but an organic reorganization that follows the laws of human cognition, ensuring that the entire explanation video generated later is logically seamless and progressively develops in thought, thus avoiding problems such as content jumps or blurred focus.

[0040] As an example, suppose we need to transform a PowerPoint presentation into a video narrated by a professional and composed digital human anchor. Through multi-dimensional element decomposition, digitization, and multi-dimensional semantic analysis, we obtain multiple semantic texts: "Title: Challenges and Opportunities Coexist," "Challenge 1: Global Inflationary Pressures," "Challenge 2: Supply Chain Restructuring," "Opportunity 1: Green Energy Transition," "Opportunity 2: Artificial Intelligence Revolution," and "Note: Cautiously Optimistic." Then, we perform contextual understanding on these multiple semantic texts and intelligently splice them together based on the contextual understanding, reorganizing them into a coherent and natural narrative: "Looking back on the past year, we undoubtedly faced..." Facing severe challenges, such as persistent global inflationary pressures and the accelerating restructuring of global supply chains, we also see tremendous development opportunities. The sustainable development transformation, exemplified by green energy, is timely, while the global artificial intelligence revolution opens up unprecedented possibilities. Based on the concatenated semantic text, emotional annotation and action instructions are generated. For example, when discussing "challenges," the emotional label is a slightly serious tone, and the action instructions are: slightly furrowed brows, hands crossed in front of the body, and a slight forward lean, conveying the severity of the problem. When the topic shifts to "opportunities," the emotional label changes to an upbeat, confident tone, and the action instructions become: relaxed brows, arms outstretched in a gesture of embracing the future, and a smile. Ultimately, a structured script for the corresponding PowerPoint presentation is generated.

[0041] Step 102: Based on the structured script file, perform dual-stream parallel processing on the PPT file and the personalized digital human model to obtain a background image sequence and the audio data and motion data corresponding to the personalized digital human model.

[0042] It should be noted that dual-stream parallel processing refers to splitting a large task into two or more subtasks that can run independently and simultaneously, thereby making full use of computing resources and significantly shortening the total processing time. In this application, the dual streams are mainly visual stream (generating a sequence of background images) and digital human stream (generating audio data and motion data). The background image sequence refers to a set of images composed of a series of continuous static images, which can form a dynamic video frame when played in a specific order and at a specific rate. Here, the background image sequence refers to a set of images composed of generating a high-resolution image for each page of the PPT and arranging all the generated images according to their correct time order of appearance in the PPT.

[0043] In some embodiments, the structured script includes digital information, voice text, emotion tags, and action instructions. Step 102 above can be implemented in the following ways: based on the digital information, the PPT file is digitized to obtain the background image sequence; based on the emotion tags, the voice text is converted into audio to obtain audio data corresponding to the personalized digital human model; based on the action instructions, the personalized digital human model is generated to obtain action data corresponding to the personalized digital human model.

[0044] In this way, by using four dimensions—digital information, voice text, emotion tags, and action commands—within the structured script to independently yet collaboratively drive three parallel generation streams of visual, auditory, and physical behavior, the generation efficiency and content consistency can be greatly improved. Digital information directly guides the accurate rendering of the background image, ensuring a strict correspondence between the visual presentation and the PPT content. Emotion tags, as core parameters, are directly injected into the TTS engine, giving the audio data an inherently correct emotional tone, thus solving the problem of the disconnect between voice and content emotion. Action commands provide clear semantic basis for the generation of digital human behavior, driving it to make body language that highly matches the content being explained, thereby generating more natural behavioral actions.

[0045] It should be noted that digital information refers to all non-textual, quantifiable structural and layout information extracted from the PPT file, specifically including page structure and type, element coordinates and dimensions, element hierarchy, chart metadata, animations and transitions, etc.; speech text refers to a fluent and natural presentation script suitable for digital human narration, formed by polishing, reorganizing and conversationalizing the text information on the PPT file, specifically including opening and closing remarks, transitions, content explanations, and emphasis and summary, etc.; emotion tags are the emotional color and tone labeled for each part of the speech text, which is the key to driving the TTS engine to generate an emotional rather than flat voice, specifically including global emotion and local emotion; action instructions refer to the action data of the digital human corresponding to each sentence of the presentation, specifically including head posture, gaze direction, facial expressions, gestures / body movements, etc.

[0046] As an example, taking a car product presentation scenario, step 101 generates structured text containing digital information, voice text, emotion tags, and action instructions. The script is then distributed to two independent pipelines that work synchronously. In the visual stream, the rendering engine renders each slide of the PPT as a high-definition image based on the digital information. If the PPT contains animation, it also renders each frame of the animation, ultimately resulting in a sequence of background images containing all the images. Simultaneously, in the digital human stream, the TTS engine converts the voice text content into audio data corresponding to the emotion tags, based on the voice text and its corresponding emotion tags. Furthermore, the digital human motion generation engine, based on the action instructions, drives the various organs of the personalized digital human model, transforming discrete instructions into a series of smooth skeletal animation data, thus obtaining the motion data of the corresponding personalized digital human model. In this way, the visual stream and the digital human stream complete their respective tasks in parallel, completely independently yet with high coordination, providing reliable data support for subsequent processing.

[0047] In some embodiments, before performing step 102, the following processing may also be performed: extracting identity features from the material data to obtain static appearance features and dynamic posture features; initializing and configuring a pre-trained general digital human model based on the static appearance features to obtain an initialized digital human model corresponding to the material data; and adjusting the behavioral features of the initialized digital human model based on the dynamic posture features to obtain the personalized digital human model.

[0048] In this way, by decoupling the material data into static appearance features and dynamic posture features, and initializing a general model based on the static appearance features, we can ensure that the visual elements of the digital human are highly consistent with the target person, thus obtaining an initialized digital human model. At the same time, we can use dynamic posture features to adjust the behavioral features of the initialized digital human model, learn the unique facial rhythms, micro-expression habits and personalized body language patterns of the target person, eliminate the "mechanical feel" and "uncanny valley effect" that are common in digital humans, and realize the personalized customization of digital human models.

[0049] It should be noted that static appearance characteristics refer to the relatively static visual attributes of an individual whose identity remains constant. These characteristics do not change fundamentally with time, expression, or action. Specifically, they can include facial geometry, skin and texture, hair characteristics, and body shape characteristics. Dynamic posture characteristics refer to the unique and dynamic behavioral patterns and personal habits that an individual exhibits when expressing emotions, speaking, or making actions. Specific content can include facial micro-expression habits, head posture and gaze patterns, unique rhythm and cadence, and signature movements.

[0050] As an example, a company wants to create a digital avatar of its CEO, Mr. Wang, for his annual speech at the company's annual meeting. The system inputs a three-minute high-definition video of Mr. Wang speaking into the camera. The system deeply analyzes this video, extracting key features. Static appearance features include "unique facial contours, nose bridge height, eye spacing, smile lines at the corners of the eyes, and a signature hairstyle." Dynamic posture features include "unconsciously pushing up his glasses when emphasizing important points, and habitually glancing briefly to the upper left when thinking." Subsequently, based on these static appearance features, the system configures a generic digital human model to an initial model suitable for Mr. Wang's appearance. Then, based on the dynamic appearance features, the initial model learns Mr. Wang's behavioral patterns, such as the habit of pushing up his glasses and shifting his gaze. Through thousands of training sessions and adjustments, the digital human learns to naturally push up his glasses when speaking about key points and to subtly change his gaze when thinking, thus completing fine-tuning to obtain a personalized digital human model of Mr. Wang.

[0051] Step 103: Divide the background image sequence into segments to obtain multiple background image sequence segments.

[0052] It should be noted that segments can be divided according to the content structure of the PPT file, such as by PPT page, by chapter or theme, or by content module. Segments can also be divided according to the time information of events, such as by pauses and semantic units in TTS audio, or by preset action instruction units. They can also be divided according to the visual effects of dynamic media, such as by animation nodes within the PPT or by transition effects between pages. The specific division can be based on the actual situation and is not limited here.

[0053] In some embodiments, step 103 above can be implemented in the following way: parsing the content of the PPT file, and dividing the content of the PPT file into multiple sub-contents based on the parsed content; for each sub-content, determining the starting background image and ending background image corresponding to the sub-content, and taking all background images between the starting background image and the ending background image in the background image sequence as the background image sequence fragment corresponding to the sub-content.

[0054] In this way, by extracting and reorganizing the original information of the PPT into multiple sub-contents with clear themes and logical boundaries, and dividing the background image sequence into multiple segments according to the correspondence between each word content and each background image in the background image sequence, the integrity of the content of each segment can be guaranteed, thus ensuring a high degree of consistency in logic, audiovisual and content of the subsequently generated digital human explanation video.

[0055] As an example, assuming the background image sequence is {p1, p2, p3, p4, p5, p6, p7, p8}, parsing the PPT file reveals that p1, p2, and p3 represent the first chapter, p4 and p5 the second chapter, p6 and p7 the third chapter, and p8 the fourth chapter. Then, for the first chapter, p1 is determined to be the starting background image, and p3 the ending background image. p1, p2, and p3 are then considered as the corresponding background image sequence segments for the first chapter. Similarly, p4 and p5 are the background image sequence segments for the second chapter, p6 and p7 for the third chapter, and p8 for the fourth chapter.

[0056] Step 104: Perform bidirectional dynamic alignment of the audio data and motion data corresponding to each background image sequence segment to obtain the digital human video segment corresponding to each background image sequence segment.

[0057] It should be noted that bidirectional dynamic alignment is an intelligent, negotiation-based timing synchronization mechanism that enables audio data and motion data to adjust and adapt to each other based on the same common goal, ultimately finding a globally optimal processing solution.

[0058] In some embodiments, step 104 described above can be implemented as follows: for each background image sequence segment, determine a first duration of the audio data and a second duration of the motion data corresponding to the background image sequence segment; based on the speech distortion cost, motion distortion cost, and pause cost, determine a cost function corresponding to each background image sequence segment, wherein the cost function is used to quantify the unnaturalness generated after adjusting the audio data and motion data corresponding to the background image sequence segment; based on the cost function corresponding to each background image sequence segment, determine a target duration sequence that minimizes the cost value of the corresponding background image sequence, wherein the target duration sequence includes the target duration corresponding to each background image sequence segment; based on the target duration sequence, adjust the audio data and motion data corresponding to each background image sequence segment respectively to obtain a digital human video segment corresponding to each background image sequence segment.

[0059] Thus, the optimization model based on dynamic programming can achieve intelligent coordination and global optimal alignment of audio and action timing in digital human video generation, fundamentally solving the "either / or" dilemma of multimodal content synchronization, and realizing an intelligent leap from passive adaptation to active negotiation. This ensures the naturalness and professionalism of the final video. Furthermore, through a two-way dynamic adjustment mechanism, it can ensure that the lip movements, facial expressions, and speech of the digital human model are accurately synchronized at the frame level, while its body movements also maintain physical rationality and fluency. This generates a digital human video clip with highly coordinated audiovisual elements and natural and smooth expression, which can significantly improve the immersion and credibility of digital human videos.

[0060] As an example, suppose the first duration of the audio data corresponding to the i-th background image sequence segment is... The second duration of the motion data is The final target rendering duration is According to the preset alignment cost rule, the cost function corresponding to each background image sequence segment is determined as shown in formula (1). The local cost of the i-th background image sequence segment is calculated by the cost function formula (1). , , , These are weighting coefficients used to adjust the importance of different modal distortion costs. The cost of speech distortion represents the auditory loss caused by speech stretching, which can be calculated using formula (2). The cost of motion distortion, representing the visual disharmony caused by motion stretching, can be calculated using formula (3). As a cost for pause / idle time, if the target duration Much larger and This indicates that an Idle frame needs to be inserted, resulting in a time redundancy penalty. Subsequently, a dynamic programming algorithm is used to determine the target duration sequence corresponding to the minimum cost of the background image sequence. Based on each target duration in the target duration sequence, the audio data and motion data corresponding to each background image sequence segment are adjusted bidirectionally to obtain the digital human video segment corresponding to each background image sequence segment. Here, the target duration must ultimately meet certain physical constraints, namely, the target duration cannot make the audio too fast or the motion too slow, and allows for the insertion of a certain period of waiting time when necessary.

[0061] (1)

[0062] (2)

[0063] (3)

[0064] In some embodiments, the above-described determination of the target duration sequence with the minimum cost value corresponding to each background image sequence segment based on the cost function corresponding to each background image sequence segment can be implemented in the following manner: For each background image sequence segment, the following processes are performed respectively: obtaining the candidate duration corresponding to the background image sequence segment; adding the cumulative minimum cost value of all background image sequence segments preceding the background image sequence segment in the background image sequence to the local cost value of each candidate duration corresponding to the background image sequence segment to obtain the cumulative cost value set of the background image sequence segment; determining the cumulative minimum cost value of each candidate duration corresponding to the background image sequence segment based on the cumulative cost value set; in response to the completion of the calculation of the last background image sequence segment, performing reverse tracing based on the cumulative minimum cost value corresponding to the last background image sequence segment to obtain the target duration sequence with the minimum cost of the background image sequence.

[0065] In this way, the processing of each segment is based on the optimal decisions of all previous segments. By calculating the cumulative cost, the system maintains the best path with the lowest cost among all possible paths from the starting point to the current point. This ensures that when processing any segment, the overall naturalness of the whole is not sacrificed for the sake of local optimization (such as making the current segment perfectly aligned). This avoids the chain reaction of misalignment of subsequent segments caused by local optimization. At the same time, since the candidate duration and local cost of each segment are calculated independently and parallel processing is supported, the computational efficiency can be effectively improved. This provides a solid algorithmic foundation for the real-time or near real-time generation of complex long videos, thereby ensuring the smoothness and harmony of the final product on the timeline.

[0066] As an example, suppose the background image sequence includes three background image sequence segments, namely segment 1, segment 2, and segment 3. First, obtain the possible candidate durations corresponding to each background image sequence segment. The candidate durations for segment 1 are t1, t2, and t3; for segment 2, t4 and t5; and for segment 3, t6 and t7. Then, add the cumulative minimum cost value of all background image sequence segments preceding the current segment to the local cost value of each candidate duration for the current segment. This yields the cumulative cost value set for the background image sequence segments. For segment 1, calculate the total cost value corresponding to candidate duration t1. 10. The total cost value corresponding to candidate duration t2 is 2, and the total cost value corresponding to candidate duration t3 is 5, resulting in the cumulative cost value set {t1:10; t2:2; t3:5}. Then, for segment 2, the total cost value corresponding to candidate duration sequences {t1, t4} is calculated to be 12, the total cost value corresponding to candidate duration sequences {t1, t5} is 11, the total cost value corresponding to candidate duration sequences {t2, t4} is 4, the total cost value corresponding to candidate duration sequences {t2, t5} is 3, the total cost value corresponding to candidate duration sequences {t3, t4} is 7, and the total cost value corresponding to candidate duration sequences {t3, t5} is 6, resulting in the cumulative cost value set {t1, t4:12; Given the sequence {t1,t5:11; t2,t4:4; t2,t5:3; t3,t4:7; t3,t5:6}, based on the cumulative cost value set, determine the minimum cumulative cost value for each candidate duration corresponding to the current background image sequence segment, resulting in {t1,t5:11; t2,t5:3; t3,t5:6}. Similarly, for segment 3, the total cost value corresponding to the candidate duration sequence {t1, t5, t6} is calculated to be 15, the total cost value corresponding to the candidate duration sequence {t1, t5, t7} is 16, the total cost value corresponding to the candidate duration sequence {t2, t5, t6} is 7, the total cost value corresponding to the candidate duration sequence {t2, t5, t7} is 8, the total cost value corresponding to the candidate duration sequence {t3, t5, t6} is 10, and the total cost value corresponding to the candidate duration sequence {t3, t5, t7} is 11. The cumulative cost value set is obtained as {t1, t5, t6: 15; t1, t5, t7: 16; t2, t5, t6: 7; t2, t5, t7: 8; t3, t5, t6: 10; t3, t5, t6: 10; t4, t5, t6: 15; t5, t7: 16; t2, t5, t6: 7; t2, t5, t7: 8; t3, t5, t6: 10; t5, t6: 10; t5, t6: 10; t3, t5, t6: 10; t5, t6: 15; t5, t6: 15; t5, t7: 16; t2, t5, t6: 7; t2, t5, t7: 8; t3, t5, t6: 10; t5, t6: 10; t5, t6: 10; t5, t6: 10; t5, Based on the cumulative cost set, the minimum cumulative cost for each candidate duration corresponding to the current background image sequence segment is determined, resulting in {t1,t5,t6:15; t2,t5,t6:7; t3,t5,t6:10}. After the calculation of segment 3 is completed, the calculation of the last background image sequence segment is completed, and the minimum cumulative cost value corresponding to segment 3 is determined to be 7. By tracing back, the corresponding target duration sequence is obtained as {t2,t5,t6}.

[0067] In some embodiments, the above-described adjustment of the audio data and motion data corresponding to each background image sequence segment based on the target duration sequence to obtain a digital human video segment corresponding to each background image sequence segment can be implemented in the following manner: For each background image sequence segment, the following processing is performed: In response to the target duration corresponding to the background image sequence segment being greater than the first duration, the audio data corresponding to the background image sequence segment is subjected to silence filling or rhythmic extension processing; In response to the target duration corresponding to the background image sequence segment being greater than the second duration, the motion data corresponding to the background image sequence segment is subjected to micro-motion filling processing; In response to the target duration corresponding to the background image sequence segment being less than the second duration, the motion data corresponding to the background image sequence segment is subjected to linear acceleration or frame extraction processing; The adjusted audio data and adjusted motion data of the background image sequence segment are combined into a video composite to obtain the digital human video segment corresponding to the background image sequence segment.

[0068] Thus, by employing three different optimization strategies, when the target duration exceeds the audio duration, the system optimizes through "silence filling" or "rhythmic extension." The former preserves natural pauses, while the latter lengthens vowels without altering pitch, resulting in a more natural sound than simply slowing down speech, simulating the authentic state of a human thinking or emphasizing. When the target duration exceeds the action duration, "micro-motion filling" is introduced to eliminate the mechanical nature of the digital human's movements. By inserting micro-motions to fill the time gaps after actions, the transitions between digital human speech are made more natural and reasonable. When the target duration is shorter than the action duration, the system uses "linear acceleration" or "frame dropping," which shortens the physical duration of actions without severely affecting key action trajectories, ensuring the overall rhythm is compact. Finally, by combining these three highly contextualized adjustment strategies with video synthesis processing, the system ensures that different adjustment alignment strategies can be adapted to different target durations. The resulting synthesized video clips are not only time-accurate but also highly natural and harmonious in both sound and visual appeal, achieving a high degree of unity between technical optimization and artistic perception.

[0069] It should be noted that silence filling refers to inserting a blank audio segment of a specific duration with no sound at the beginning, end, or middle of an audio clip, at a preset pause point; prosodic extension is a more advanced and natural method of extending audio duration by intelligently stretching vowels and silent parts of speech to increase the total duration while maintaining the clarity of consonants and the overall intonation; micro-motion filling refers to intelligently inserting a series of context-appropriate, small, and non-critical actions after the core action is completed; linear acceleration refers to shortening the duration of an action by increasing the playback frame rate of the action sequence; and frame extraction refers to selectively removing unimportant or redundant frames from an action sequence to shorten the duration of the action.

[0070] As an example, suppose the target duration for segment 1 is determined to be 6 seconds, segment 2 to be 11 seconds, and segment 3 to be 9 seconds using a dynamic programming algorithm. For segment 1, the first duration of the audio data is 5 seconds, which is less than the target duration. Therefore, according to the audio adjustment strategy, a prosodic extension processing method is selected, inserting very subtle and natural pauses between words in the sentence and slightly lengthening the final sound of the last word to obtain the adjusted audio data. The second duration of the motion data is 6 seconds, which is equal to the target duration, requiring no processing. Then, the adjusted audio data and motion data for segment 1 are concatenated and combined, and combined with the background image sequence of segment 1 to form the first digital human video segment. Next, for segment 2, the audio data duration is exactly 11 seconds, which is equal to the target duration, requiring no adjustment. The second duration of the motion data is 9 seconds, which is longer than the target duration. According to the motion filling strategy, after the digital human introduces the last artist... Next, intelligent micro-motions are added (e.g., the digital human slowly raises its hand while slightly turning its body towards the camera with an encouraging smile) to make the movements more lifelike, resulting in adjusted motion data. Then, the audio data corresponding to segment 2 is spliced ​​and synthesized with the adjusted motion data, and combined with the background image sequence of segment 2 to form the second digital human video segment. After that, for segment 3, the audio data corresponds to a duration of 8 seconds, which is less than the target duration. The system also uses prosodic expansion, pausing slightly at the end of the sentence to give the audience time to think, resulting in adjusted audio data. The motion data corresponds to a second duration of 11 seconds, which is longer than the target duration. According to the motion compression strategy, the two non-critical transition frames of "turning around" and "raising hand" are linearly accelerated very slightly to reduce the motion duration, resulting in adjusted motion data. Then, the adjusted audio data corresponding to segment 3 is spliced ​​and synthesized with the adjusted motion data, and combined with the background image sequence of segment 3 to form the third digital human video segment.

[0071] Step 105: Merge the digital human video segments corresponding to each of the background image sequence segments to obtain the digital human explanation video corresponding to the PPT file.

[0072] As an example, firstly, all video clips are arranged end-to-end on the timeline according to the original logical order of the PPT, forming a continuous video stream. Then, smooth visual transition effects can be added at the junctions between clips, based on the original settings of the PPT or intelligent recommendations from AI. Next, the audio streams of each clip can be connected, and crossfading can be applied at the junctions to ensure consistency in volume, timbre, and background noise. Finally, the spliced ​​video stream (including background images, digital human figures, and transition effects) and audio stream are finally synthesized and compressed into a universal, playable video file format (such as MP4) using a professional encoder (e.g., H.264) to complete the video clip fusion operation, thereby obtaining the digital human presentation video corresponding to the PPT file.

[0073] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.

[0074] In related technologies, PPT-to-video tools (such as PowerPoint's built-in "Export as Video" function and some online conversion tools) typically only convert PPT animations and pages into video streams, or simply convert text into speech and mechanically play PPT images. In recent years, with the rapid development of AI technology, digital human technology has been gradually applied to education, training, and other fields, and generating PPT presentation videos using digital human technology has gradually become a reality. However, the following technical problems exist in the implementation of related technologies:

[0075] (1) The digital human image is monotonous and unrealistic, resulting in a lack of spatiotemporal consistency of audiovisual modalities and a significant “uncanny valley” effect. That is, traditional digital human generation technology usually adopts simple “audio-lip-sync” rule mapping or low-precision 3D model driving. This method results in stiff facial expressions and lack of micro-expressions in the generated digital human. Moreover, the lip-sync and pronunciation often show non-linear offset on the time axis (i.e., the lip-sync does not match). For example, under high-frequency speech or complex emotional expression, the facial feature points of the digital human are prone to jitter or drift, resulting in a visually unnatural feeling (i.e., the “uncanny valley” effect), which seriously affects the immersiveness and credibility of information transmission.

[0076] (2) Difficulty in synchronizing audio and video leads to a lack of semantic-level timeline mapping. Furthermore, existing technologies cannot establish a deep mapping relationship between the narration (audio stream) and PPT animation nodes (video stream). The rendering engine cannot perceive pauses, changes in speech rate, and semantic endings of sentences, resulting in an inability to automatically stretch or compress the PPT screen's dwell time based on the actual length of the audio. Secondly, the computational power and logic of dynamic rendering are lacking. Existing solutions are mostly simple "screen recording" or "page-by-page splicing," lacking real-time "temporal resampling" algorithms. When the length of the narration exceeds the preset screen duration, the system cannot automatically perform "frame insertion" or "screen freeze" operations; conversely, it cannot perform "speed playback" or "editing" operations, ultimately leading to the narration being truncated or the screen and audio being severely misaligned. Thirdly, the separation of semantics and action means that existing technologies cannot deeply understand the logical focus and emotional tone of the PPT page content. That is, the digital human's body movements are usually randomly generated or in fixed cycles, making it impossible to adjust the narration based on the content of the narration. The automatic matching of appropriate body language to the tone of voice (e.g., passionate, soothing) or specific semantics (referring to the screen, emphasizing numbers) results in a superficial connection between the "performance" and the "speaking." In addition, the unidirectional drive leads to rigid synchronization. Existing technologies are usually unidirectional synchronization—either "fixed-length audio, forced video" or "fixed-length video, truncated audio." There is a lack of a two-way dynamic alignment mechanism that can simultaneously consider "voice duration" and "action performance duration." For example, when a complex body action (such as waving to show the whole picture) has not been completed, the voice may have already ended and the next page has been switched, resulting in the action being abruptly cut off and the picture looking extremely unnatural.

[0077] (3) High modification and maintenance costs. Existing PPT to video technology usually adopts the "full linear rendering" mode, where all frames of the video are regarded as a whole stream. The high modification cost caused by strong coupling and the lack of modular indexing mechanism inside the video stream means that when the user only modifies a chart data on a certain page of PPT or adjusts the wording of a certain explanatory sentence, the system cannot locate the specific impact scope of the modification on the video stream, and causes an ineffective waste of computing resources. Since the unchanged parts cannot be identified, the system is forced to re-execute the entire process (TTS generation, motion calculation, image rendering, encoding and synthesis), resulting in a large amount of computing power being wasted on repeatedly calculating unchanged video segments, which is not only time-consuming, but also seriously hinders real-time preview and rapid iteration.

[0078] (4) Inability to analyze complete PPT information. Due to insufficient multimodal document parsing capabilities, it is impossible to understand unstructured information. Existing technologies usually only support the extraction and analysis of a single text modality (uni-modal), and heavily rely on traditional OCR or text parsing interfaces. When the input document (such as PPT) contains a large amount of unstructured visual information such as charts, formulas, experimental data graphs or flowcharts, existing solutions cannot extract the semantic logic, resulting in information loss. This "visual blind spot" directly causes the generated explanation script to be disconnected from the content of the presentation screen, and it is impossible to achieve accurate and in-depth interpretation of the text and images.

[0079] In some embodiments, this application proposes a generative digital human-driven architecture based on Latent Space Feature Disentanglement, which supports the extraction of personalized features from a small amount of video data (e.g., a 3-minute narrated video) or a single image provided by the user, and achieves high-fidelity driving through semantic instructions. First, the system receives source materials from the user (such as a single RGB image or a short video clip). Then, it uses a pre-trained 3D deformable model (3DMM) or identity encoder to analyze the source materials frame by frame. Next, in the latent space, the image features are decomposed into two orthogonal (non-interfering) vectors to obtain static appearance features (encoding the user's facial geometry, skin texture, facial proportions, and other identity-constant features) and dynamic pose features (encoding head pose, gaze direction, and expression coefficients). Finally, based on the short video provided by the user (such as 3 minutes of data), the general rendering network is subjected to low-rank adaptation (LoRA) or lightweight fine-tuning to enable it to accurately fit the user's specific micro-expression habits (such as habitual eyebrow raising and lip pursing), thereby eliminating the "mechanical feel".

[0080] In some embodiments, this application addresses the user's need to "change appearance / background" by introducing in-painting technology based on the diffusion model. By automatically segmenting the character's torso or background area to generate a mask, and combining it with text prompts (such as "wearing a business suit" or "tech-savvy background"), the torso texture and background pixels are regenerated while strictly maintaining the facial identity features. This achieves asset decoupling between "facial identity" and "environment / clothing," allowing multiple looks to be obtained without reshooting.

[0081] In some embodiments, this application utilizes a Large Language Model (LLM) to analyze the deep semantics and emotional connotations of the input text, outputs action description instructions (e.g., "spread your hands to welcome"), and constructs a text-to-Gesture Network. This generation network is based on a probabilistic graphical model, maps discrete text instructions to a continuous sequence of 3D skeletal keypoints, and introduces an action interpolation algorithm to ensure the temporal smoothness of the generated limb movements, and automatically generates transition actions for the first and last frames to ensure continuity during loop playback.

[0082] In some embodiments, this application fuses the generated "limb skeleton flow" with the extracted "static appearance features," employing an improved Wav2Lip or FaceFormer architecture. Using the Mel-spectrogram of speech as conditional input, it predicts the lip top and bottom point displacements for each frame, generating lip shape parameters strictly synchronized with the speech. Then, using a Neural Radiation Field (NeRF) or Generative Adversarial Network (GAN) generator, all the aforementioned parameters (appearance, pose, movement, and lip shape) are rendered into the final 2D video frames. This model possesses generalization capabilities, reusing the same driving logic for real-life portraits, 3D models, and cartoon characters.

[0083] In some embodiments, this application utilizes a Large Language Model (LLM) to perform in-depth analysis of the text and image content of each PPT slide and the user's presentation requirements, extracting key information points and generating a presentation script that conforms to spoken language. Simultaneously, the LLM automatically labels tone tags (e.g., excitement, seriousness, questioning) and body language instructions (e.g., pointing to the upper left corner, spreading hands for emphasis, nodding) based on the script's semantics. Subsequently, combining the tone tags, a high-fidelity TTS engine generates an emotionally charged speech stream, accurately calculating the audio duration and phoneme-level timestamps of each sentence. Simultaneously, the digital human engine, based on the action instructions output by the LLM, pre-... The system calculates the physical duration required to complete the physical movements. Then, it uses a dynamic programming algorithm to perform bidirectional temporal matching and optimization of the speech stream and the action stream to solve the problem of inconsistency between audio duration and action duration. For specific implementation, please refer to the implementation method of step 104 above. Finally, based on the aligned speech stream, the system drives the digital human's lip-sync and facial micro-expressions (such as blinking and eyebrow raising) in real time. Based on the aligned action timeline, the system drives the skeletal system to execute preset actions. Finally, the PPT screen, digital human actions, facial expressions and speech are synthesized frame by frame to output a highly natural and perfectly synchronized audio-visual presentation video.

[0084] In some embodiments, to address the problem that existing technologies cannot understand unstructured visual elements (such as abstract charts and data graphs) and missing semantics in mixed-format documents (PPT / PDF), this application proposes a cascaded multimodal parsing and semantic reconstruction architecture. Through a three-stage process of "layout deconstruction - visual translation - logical induction", it achieves accurate conversion from unstructured documents to high semantic density video scripts. First, the input PPT document stream is converted into a high-resolution image sequence. Then, a deep learning-based object detection network (such as Faster R-CNN or YOLO-X) is used to perform a panoramic scan of the page, identifying and locating individual elements within the page. Next, the identified elements are categorized into different types (text blocks, raster images, vector charts, formulas). Simultaneously, the bounding box coordinates and Z-order of each element are extracted to construct an initial Document Object Model (DOM) tree, preserving layout information for subsequent pixel-level reconstruction. Subsequently, individual image patches are cropped from the extracted image and chart regions and input into a multimodal vision-language model (VLM) (such as CLIP, BLIP, or GPT-4V visual encoders). The model outputs a dense caption for the image, such as: "A flowchart showing the hierarchical structure of a neural network." Finally, combined with the page context, prompts are generated through... The Engineering model guides the extraction of deep information, such as "analyzing data trends in charts" or "explaining the core logic of the architecture diagram," thereby transforming uneditable pixel information into structured semantic text that can be processed by LLM, filling the "visual blind spot" of traditional OCR. Finally, the extracted original text is merged with the generated visual description text, and the long text understanding capability of the Large Language Model (LLM) is used to analyze the semantic relationships between pages, automatically identify "chapter titles," "transition pages," and "body pages," and cluster discrete pages into logical "chapter" and "section." Finally, a standardized JSON / XML intermediate representation file is generated. This file not only contains the polished explanation script but also strictly binds the corresponding visual element references and their spatiotemporal trigger logic in the video frame, ensuring that the generated video content is logically rigorous and hierarchically clear.

[0085] The above are embodiments of the method proposed in this application. Based on the same inventive concept, embodiments of this application also provide a device, the structure of which is as follows: Figure 2 As shown.

[0086] Figure 2This is a schematic diagram of the internal structure of a device provided in an embodiment of this application. Figure 2 As shown, the device includes:

[0087] At least one processor 201;

[0088] And a memory 202 that is communicatively connected to at least one processor;

[0089] The memory 202 stores instructions that can be executed by at least one processor. The instructions are executed by at least one processor 201 to enable at least one processor 201 to perform the steps of the method corresponding to any of the above embodiments.

[0090] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium stores computer-executable instructions configured to perform the steps of the method corresponding to any of the above embodiments.

[0091] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0092] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.

[0093] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0094] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0095] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0096] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0097] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0098] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0099] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0100] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0101] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for generating digital human narration videos, characterized in that, The method includes: Key information is extracted from the PPT file to obtain the corresponding structured script file; Based on the structured script file, the PPT file and the personalized digital human model are processed in parallel with two streams to obtain a background image sequence and audio and motion data corresponding to the personalized digital human model. The background image sequence is segmented to obtain multiple background image sequence segments; For each background image sequence segment, a first duration of audio data and a second duration of motion data corresponding to the background image sequence segment are determined, and a speech distortion cost for a target duration of the background image sequence segment is constructed based on the first duration, and a motion distortion cost for a target duration of the background image sequence segment is constructed based on the second duration. Based on speech distortion cost, motion distortion cost, and pause cost, a cost function is determined for each background image sequence segment. The cost function is used to quantify the unnaturalness produced by adjusting the audio data and motion data corresponding to the background image sequence segment. For each of the background image sequence segments, the following processes are performed: The candidate duration corresponding to the background image sequence segment is obtained, and the local cost value corresponding to each candidate duration is calculated based on the cost function corresponding to the background image sequence segment; when the background image sequence segment is the first background image sequence segment, the local cost value corresponding to each candidate duration is used as the cumulative minimum cost value corresponding to each candidate duration; when the background image sequence segment is not the first background image sequence segment, the cumulative minimum cost value corresponding to each candidate duration of the previous background image sequence segment is added to the local cost value corresponding to each candidate duration of the background image sequence segment to obtain the cumulative cost value set of the background image sequence segment; for each candidate duration in the background image sequence segment, the minimum cumulative cost value corresponding to the candidate duration in the cumulative cost value set is used as the cumulative minimum cost value of the candidate duration; After the cumulative minimum cost value of each candidate duration corresponding to the last background image sequence segment is calculated, the minimum cumulative minimum cost value of multiple candidate durations corresponding to the last background image sequence segment is used for reverse tracing to obtain the target duration sequence with the minimum cost of the background image sequence; wherein, the target duration sequence includes the target duration corresponding to each background image sequence segment; Based on the target duration sequence, the audio data and motion data corresponding to each background image sequence segment are adjusted to obtain a digital human video segment corresponding to each background image sequence segment; The digital human video segments corresponding to each of the background image sequence segments are merged to obtain the digital human explanation video corresponding to the PPT file.

2. The method according to claim 1, characterized in that, The step of extracting key information from the PPT file to obtain a structured script file corresponding to the PPT file includes: The PPT file is decomposed into multimodal elements, and each decomposed element is digitized to obtain multiple structured data. Perform multi-dimensional semantic analysis on each of the structured data to obtain the semantic text corresponding to each of the structured data; Each of the semantic texts is merged and recombined to obtain a structured script file corresponding to the PPT file.

3. The method according to claim 2, characterized in that, The step of fusing and recombining each semantic text to obtain a structured script file corresponding to the PPT file includes: Each semantic text is subjected to contextual understanding, and based on the understood contextual content, each semantic text is concatenated to obtain the concatenated semantic text; Sentiment annotation and action command generation are performed on the concatenated semantic text to obtain a structured script file corresponding to the PPT file.

4. The method according to claim 1, characterized in that, The structured script includes digital information, voice text, emotion tags, and action instructions; The process involves performing dual-stream parallel processing on the PPT file and the personalized digital human model based on the structured script file to obtain a background image sequence and corresponding audio and motion data of the personalized digital human model, including: Based on the digital information, the PPT file is digitally processed to obtain the background image sequence; Based on the emotion tags, the voice text is converted into audio to obtain the audio data corresponding to the personalized digital human model; Based on the action instructions, actions are generated on the personalized digital human model to obtain action data corresponding to the personalized digital human model.

5. The method according to claim 1, characterized in that, Before performing dual-stream parallel processing on the PPT file and the personalized digital human model based on the structured script file to obtain the background image sequence and the audio and motion data corresponding to the personalized digital human model, the method further includes: Identity features are extracted from the material data to obtain static appearance features and dynamic posture features; Based on the static appearance features, the pre-trained general digital human model is initialized and configured to obtain the initial digital human model corresponding to the material data; Based on the dynamic posture features, the behavioral features of the initial digital human model are adjusted to obtain the personalized digital human model.

6. The method according to claim 1, characterized in that, The background image sequence is segmented to obtain multiple background image sequence segments, including: The PPT file is parsed, and based on the parsed content, the content of the PPT file is divided into multiple sub-contents; For each sub-content, a starting background image and an ending background image corresponding to the sub-content are determined, and all background images between the starting background image and the ending background image in the background image sequence are taken as the background image sequence segment corresponding to the sub-content.

7. The method according to claim 1, characterized in that, The step of adjusting the audio and motion data corresponding to each background image sequence segment based on the target duration sequence to obtain a digital human video segment corresponding to each background image sequence segment includes: For each of the aforementioned background image sequence segments, the following processing is performed: In response to the target duration corresponding to the background image sequence segment being greater than the first duration, the audio data corresponding to the background image sequence segment is subjected to mute filling or rhythmic extension processing. In response to the target duration corresponding to the background image sequence segment being greater than the second duration, micro-motion filling processing is performed on the motion data corresponding to the background image sequence segment; In response to the fact that the target duration corresponding to the background image sequence segment is less than the second duration, linear acceleration or frame extraction is performed on the motion data corresponding to the background image sequence segment. The adjusted audio data and adjusted motion data of the background image sequence fragment are combined to obtain a digital human video fragment corresponding to the background image sequence fragment.

8. A digital human narration video generation device, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-7.