Method and system for making virtual digital human to explain PPT
By preprocessing PPT documents and editing video scenes, a digital human narration video with a transparent background is generated, solving the problem of dynamic element rendering and improving the PPT presentation effect and generation efficiency.
Patent Information
- Application Number
- CN202511273084.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies cannot effectively render and display dynamic content such as audio, video, and animated images contained in PPT pages, which affects the presentation effect of virtual digital humans explaining PPTs, and the generation efficiency is low.
By generating PPT presentation documents, preprocessing, parsing, video scene editing, audio and video separation, and layer compositing, a digital human narration video with a transparent background is generated, and multi-task parallel processing is combined to improve generation efficiency.
It achieves effective rendering of dynamic elements in PPT, enhancing the presentation effect, and improves generation speed and performance through hardware and software optimization.
Smart Images

Figure CN121397259A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of virtual digital humans and audio-visual technology, and in particular to a method and system for creating virtual digital human presentation slides. Background Technology
[0002] With the rapid development of artificial intelligence technology, virtual digital human technology has gradually matured and is widely used in entertainment, education, news broadcasting, and other fields. Virtual digital humans achieve highly realistic effects through customized digital human avatars and AI-driven speech synthesis and facial expression simulation. Using virtual digital humans to explain PowerPoint presentations is a common application scenario in the education field. Existing technical solutions typically convert PowerPoint pages into static background images and then use audio to drive the digital human to synthesize PowerPoint presentation videos. However, these solutions often suffer from compatibility issues with the PowerPoint document content, meaning they cannot render and display dynamic page content such as audio, video, and animated images contained in the PowerPoint slides, thus affecting the overall presentation effect of using virtual digital humans to explain PowerPoint documents. Therefore, there is an urgent need for a method and system for creating virtual digital human presentations to solve the existing technical problems. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art, and proposes a method and system for creating virtual digital human presentation PPTs.
[0004] In a first aspect, embodiments of the present invention provide a method for creating a virtual digital human to present a PowerPoint presentation, comprising:
[0005] Generate PPT presentation documents based on user needs;
[0006] The PPT presentation document is preprocessed to generate explanatory text information for each page of the PPT presentation document;
[0007] The pre-processed PPT presentation document is parsed to obtain the set of graphic elements and explanatory text information contained in each slide of the PPT presentation document;
[0008] Edit the video scenes of the parsed PPT presentation document, including at least creating new video scenes, setting preset scene templates, setting slide display settings, setting digital human appearance, setting digital human voice, setting narration text, and setting video output.
[0009] Using the explanatory text information, digital human image, and digital human voice of a single slide in a PPT presentation document as input data, a first target video is generated. The first target video is a virtual digital human explanation video file generated by combining the virtual digital human image with sound samples and single-page explanatory text.
[0010] Extract the digital human foreground target from the first target video, remove the video background, and generate a second target video, which is a digital human explanation video file with a transparent background.
[0011] The second target video is subjected to audio-video separation processing, the audio file is extracted and input into the subtitle generation unit to generate a subtitle file in standard subtitle format;
[0012] The set of graphic elements contained in a single slide and the video of the digitized human explaining the slide are taken as input, and a video layer compositing operation is performed to generate a third target video, which is a video file of the virtual digitized human explaining a single slide.
[0013] The third target video set and video transition animation set generated from all slides of the PPT presentation document are used as input. The video segment merging operation is performed to generate the fourth target video, which is a video file in which a virtual digital human explains the entire PPT presentation document.
[0014] Furthermore, based on user needs, a PPT presentation document is generated. Specific steps include: executing a first preset workflow according to the user-inputted PPT theme information and generation conditions to generate a PPT presentation document that meets user requirements; the first preset workflow, based on a workflow automation tool, defines multiple workflow nodes and executes them in an automated manner to achieve automatic generation of the PPT presentation document; the first preset workflow includes a first input node, which receives the user-inputted PPT theme information and generation conditions; the first preset workflow also includes an AI agent node, which is bound to a specific large language model, using the information from the first input node as prompts for the large language model, calling the large language model via an API interface, and generating PPT outline and content data in a specified format; the first preset workflow also includes a data conversion node, which parses the PPT content generated by the large model and converts its data format into a PPT document structure; the first preset workflow also includes a code execution node, which inputs the PPT document data into the PPT document generation tool to complete the PPT document generation operation.
[0015] Further, the PPT presentation document is preprocessed to generate explanatory text information for each page of the PPT presentation document. Specific steps include: using the PPT presentation document as input, executing a second preset workflow to generate explanatory text information for each page of the PPT, and saving the generated explanatory text information to the notes section of the PPT page; the second preset workflow is based on a workflow automation tool, defining multiple workflow nodes, and using automated workflow execution to automatically generate the explanatory text for the PPT pages; the second preset workflow includes a file reading node, which is used to read the content of the PPT file; the second preset workflow also includes a code execution node, which executes code... The second preset workflow involves parsing the document structure of the read PPT file to obtain a set of PPT slide pages, and sequentially traversing and reading the text information contained in each slide page. It also includes an AI agent node, which is bound to a specific large language model. This AI agent node combines the text information of the PPT page with preset prompts and calls the large language model via an API interface to generate the explanatory text information for each PPT page. Furthermore, the second preset workflow includes a code execution node, which adds the explanatory text information generated by the large language model to the notes section of the PPT page. Finally, the second preset workflow includes a file writing node, which saves the PPT file content.
[0016] Furthermore, the preprocessed PPT presentation document is parsed to obtain the set of graphic elements and explanatory text information contained in each slide of the PPT presentation document. The specific steps include:
[0017] Load the uploaded PPT document, read the PPT document structure information, and obtain the set of slide pages contained therein;
[0018] Iterate through the collection of slideshow pages, generate a thumbnail for each slideshow page, and save it as a local image file;
[0019] Read the set of graphic elements contained in each slide page, obtain the contained dynamic graphic elements, including dynamic graphic elements such as animated images, audio, and video, obtain the coordinates, width, height, and embedding file of the dynamic graphic elements in the page, obtain the embedding file of the dynamic graphic elements from the PPT document structure and save it to the specified local directory.
[0020] Retrieve the preset explanatory text information contained in each slide page. If a page does not have preset explanatory text information, set the preset text information of the corresponding page to empty.
[0021] Display the PowerPoint slides as a list within the video scene editing unit.
[0022] Furthermore, the video scenes in the parsed PPT presentation document are edited. Specific steps include:
[0023] Create a new video scene and select a preset scene template. The preset scene template defines a set of pre-set PPT presentation scene templates, including background images, slides and digital human display positions. It also supports user-defined settings for background fill images, resolution, frame rate and video format parameters.
[0024] In the slide display settings, select the PPT slide page and adjust its position, width, and height within the video scene.
[0025] Digital human character settings: Select a suitable digital human character and adjust its position, width, and height in the video scene;
[0026] Digital human voice settings: Select a suitable preset voice, and adjust the volume, speech rate, and tone of the voice.
[0027] The presentation text settings allow for secondary editing of the presentation text on the slides, and support for inserting a pause at a specific position in the presentation text;
[0028] Dynamic element settings allow you to configure playback control parameters for dynamic elements such as videos, audio, and animated images on the page. When the page contains multiple video elements, you need to set their playback order. When the duration of the digital human's narration video differs from the playback duration of other video elements on the page, select a preset playback control strategy. When the page contains multiple audio elements, you need to set their playback order and whether each audio element should loop. When the page contains multiple animated image elements, you need to set the playback order and playback time parameters for each animated image element.
[0029] Furthermore, using the explanatory text information, digital human image, and digital human voice from a single slide in the PowerPoint presentation as input data, the first target video is generated. Specific steps include:
[0030] The system performs speech synthesis on the input PowerPoint presentation text and digital human voice for a single page, generating an audio file. The speech synthesis process includes: performing word segmentation, text analysis, and normalization preprocessing on the input text to generate phoneme sequences, sentiment indicators, and prosodic feature annotations; encoding the text using a pre-trained large-scale text model to generate semantic vectors; inputting the semantic vectors into an autoregressive transformer to generate speech tags, which are discrete representations after vector quantization, with each tag representing a basic unit in the speech signal; optimizing the speech tags using conditional flow matching technology to adjust and optimize their acoustic features, converting the optimized speech tags into Mel spectra; inputting the Mel spectra into a pre-trained vocoder to generate continuous speech waveforms; performing volume adjustment postprocessing on the generated speech waveforms; and outputting the processed speech waveforms as the final audio presentation file.
[0031] Based on the input digital human image and audio narration file, a digital lip-syncing operation is performed to generate a virtual digital human narration video. The digital lip-syncing operation is implemented using a multi-lip-syncing algorithm framework to control the virtual digital human's lip movements and facial expressions to align and match with the input audio file, achieving the effect of lip-syncing between the digital human's narration audio and lip movements.
[0032] Further, a digital human foreground target is extracted from the first target video to generate a second target video. Specific steps include:
[0033] The input digital human explanation video is decoded to obtain image data for each frame;
[0034] A pre-trained image segmentation model is used to segment the foreground of each frame of the image to obtain a segmentation mask for the foreground object. The segmentation mask is a binary image used to distinguish between the foreground and the background.
[0035] Foreground objects are extracted from the original image based on the segmentation mask, and the contours of the foreground objects are adjusted to make them smoother and more refined using morphological operations and contour fitting algorithms.
[0036] Create a transparent background image with the same resolution as the original image, and overlay the extracted foreground object image onto the transparent background image to generate a digital human image with a transparent background;
[0037] Each frame, after removing the background, is re-merged into a single complete video file. Finally, the video file is merged with the audio of the original video file to output a digital human narration video with a transparent background. The digital human narration video file with a transparent background is in the WebM video container format encoded with VP9.
[0038] Furthermore, the second target video undergoes audio-video separation processing to generate a subtitle file in standard subtitle format. Specific steps include:
[0039] The input audio files are preprocessed using noise reduction, filtering, and gain control techniques to improve audio quality;
[0040] A pre-trained speech recognition model is used to perform speech recognition on the pre-processed audio, converting the speech into text and generating a timestamp for each word.
[0041] Extract the start time, end time, and text content of each paragraph;
[0042] The extracted timestamps and text content are formatted according to the SRT format to generate a subtitle file in standard subtitle format.
[0043] Furthermore, taking the set of graphic elements contained in a single slide and the digital human narration video as input, a third target video is generated by overlaying clipping layers. Specific steps include:
[0044] Add a background layer to the clip layer, and load a background image or background color into the background layer;
[0045] Add a static PPT element rendering layer to the clipping layer to load PPT thumbnails;
[0046] Add a PPT motion graphics rendering layer to the clipping layer. The number of PPT motion graphics rendering layers should be consistent with the number of motion graphics contained in the PPT page. Each motion graphics rendering layer loads and displays one motion graphics element contained in the PPT page. Based on the coordinates, width, height, embedded files, and playback control information of the set of motion graphics elements in the PPT page, load the corresponding embedded file resources in the video, audio, and animated image rendering layers, calculate the coordinate transformation, move the motion graphics rendering layer to the corresponding position in the video output scene, and adjust the display width and height of the motion graphics rendering layer.
[0047] Add a digital human rendering layer to the editing layer to load the digital human narration video file;
[0048] Add a foreground graphic rendering layer to the clipping layer. The number of foreground graphic rendering layers should match the number of foreground graphics in the PPT page, and they should be stacked and displayed in ascending order of index.
[0049] Add a subtitle rendering layer to the clip layer to load the subtitle file and display the video subtitles.
[0050] Secondly, this invention also discloses a system for creating virtual digital human presentation PPTs, comprising: a PPT document generation unit, a document preprocessing unit, a document parsing and processing unit, a video scene editing unit, a first video generation unit, a second video generation unit, a subtitle generation unit, a third video generation unit, and a fourth video generation unit; wherein:
[0051] The PPT document generation unit is used to generate PPT presentation documents according to user needs;
[0052] The document preprocessing unit is used to preprocess the PPT presentation document and generate explanatory text information for each page of the PPT presentation document;
[0053] The document parsing and processing sheet is used to parse and process the pre-processed PPT presentation document to obtain the set of graphic elements and explanatory text information contained in each slide of the PPT presentation document;
[0054] The video scene editing unit allows editing of video scenes in the parsed PPT presentation document, including at least creating new video scenes, setting preset scene templates, setting slide display settings, setting digital human appearance, setting digital human voice, setting narration text, and setting video output.
[0055] The first video generation unit is used to generate a first target video by taking the explanatory text information, digital human image and digital human voice of a single slide page in a PPT presentation document as input data. The first target video is a virtual digital human explanation video file generated by combining the virtual digital human image with sound samples and single-page explanatory text.
[0056] The second video generation unit is used to extract the digital human foreground target from the first target video, remove the video background, and generate a second target video, wherein the second target video is a digital human explanation video file with a transparent background.
[0057] The subtitle generation unit is used to perform audio-video separation processing on the second target video, extract the audio file and input it into the subtitle generation unit to generate a subtitle file in standard subtitle format;
[0058] The third video generation unit is used to take the set of graphic elements contained in a single slide page and the digital human explanation video as input, perform a video layer compositing operation, and generate a third target video, wherein the third target video is a video file of a virtual digital human explaining a single slide page;
[0059] The fourth video generation unit is used to take the third target video set and video transition animation set generated from all slide pages of the PPT presentation document as input, perform a video segment merging operation, and generate a fourth target video. The fourth target video is a video file in which a virtual digital human explains the entire PPT presentation document.
[0060] This invention provides a method for creating a virtual digital human PowerPoint presentation. The method involves: generating a PowerPoint presentation document based on user requirements; preprocessing the PowerPoint presentation document to generate explanatory text information for each page; parsing the preprocessed PowerPoint presentation document to obtain the set of graphic elements and explanatory text information contained in each slide; editing the video scenes of the parsed PowerPoint presentation document; generating a first target video using the explanatory text information, digital human image, and digital human voice of a single slide as input data; extracting the digital human foreground target from the first target video to generate a second target video; performing audio-video separation processing on the second target video to generate a subtitle file in standard subtitle format; generating a third target video using the set of graphic elements contained in a single slide and the digital human PowerPoint presentation video as input; and generating a fourth target video using the set of third target videos generated from all slides of the PowerPoint presentation document and a set of video transition animations as input.
[0061] This invention can solve the problem that existing technical solutions can only display static PPT page elements and cannot display dynamic elements such as audio, video, and animated images contained in PPT pages. It has better compatibility with PPT presentation documents and can effectively enhance the presentation effect of PPT documents.
[0062] This invention can solve the problem of low generation efficiency in existing publicly available technical solutions. By adopting a hardware deployment method of one machine with multiple cards or graphics card clusters, and by splitting PPT documents and using a multi-task parallel processing combined with multi-threaded concurrent processing operation for video synthesis, it can give full play to the advantages of parallel computing, effectively improve the generation speed and performance of PPT presentation videos, and at the same time have good scalability. Attached Figure Description
[0063] Figure 1 A flowchart illustrating a method for creating a virtual digital human to present a PowerPoint presentation, as provided in an embodiment of the present invention.
[0064] Figure 2 This is a structural block diagram of a system for creating virtual digital human presentation slides, provided as an embodiment of the present invention. Detailed Implementation
[0065] To enable those skilled in the art to better understand the technical solutions of the present invention, exemplary embodiments of the present invention are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0066] Where there is no conflict, the various embodiments of the present invention and the features thereof may be combined with each other.
[0067] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0068] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0069] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art and the invention, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.
[0070] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information all comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example: appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely locating a specific individual.
[0071] To address at least one of the technical problems existing in the aforementioned related technologies, the present invention provides a method and system for creating virtual digital human presentation slides.
[0072] This embodiment discloses a method for creating a virtual digital human to present a PowerPoint presentation, such as... Figure 1 ,include:
[0073] S100. Generate a PPT presentation document according to user requirements; In this embodiment, generating a PPT presentation document according to user requirements includes the following steps: Executing a first preset workflow according to the PPT theme information and generation conditions input by the user to generate a PPT presentation document that meets the user's requirements; The first preset workflow is based on a workflow automation tool, defines multiple workflow nodes, and executes them in an automated workflow manner to achieve automatic generation of the PPT presentation document; The first preset workflow includes a first input node, which is used to receive the PPT theme information and generation conditions input by the user; The first preset workflow also includes an AI agent node, which is bound to a specific large language model, uses the information from the first input node as prompt words for the large language model, calls the large language model via an API interface, and generates PPT outline and content data in a specified format; The first preset workflow also includes a data conversion node, which is used to parse the PPT content generated by the large model and convert its data format into a PPT document structure; The first preset workflow also includes a code execution node, which inputs the PPT document data into the PPT document generation tool to complete the PPT document generation operation.
[0074] In some preferred embodiments, users may also use common PPT creation software (such as Microsoft PowerPoint or WPS software) to create PPT presentation documents, or use PPT documents from other sources.
[0075] S200. Preprocess the PPT presentation document to generate explanatory text information for each page of the PPT presentation document;
[0076] In this embodiment, the PPT presentation document is preprocessed to generate explanatory text information for each page of the PPT presentation document. Specific steps include: using the PPT presentation document as input, executing a second preset workflow to generate explanatory text information for each page of the PPT, and saving the generated explanatory text information to the notes section of the PPT page; the second preset workflow is based on a workflow automation tool, defining multiple workflow nodes, and using automated workflow execution to automatically generate the explanatory text for the PPT pages; the second preset workflow includes a file reading node, which is used to read the content of the PPT file; the second preset workflow also includes a code execution node, which... The second preset workflow involves parsing the document structure of the read PPT file to obtain a set of PPT slide pages, and sequentially traversing and reading the text information contained in each slide page. It also includes an AI agent node, which is bound to a specific large language model. This AI agent node combines the text information of the PPT page with preset prompts and calls the large language model via an API interface to generate the explanatory text information for each PPT page. Furthermore, the second preset workflow includes a code execution node, which adds the explanatory text information generated by the large language model to the notes section of the PPT page. Finally, the second preset workflow includes a file writing node, which saves the PPT file content.
[0077] In some preferred embodiments, users can also manually enter the text content to be explained on each page in the notes section of the PPT page, or check and modify the explanation text information generated with the assistance of the large language model. In order to distinguish it from the original notes in the notes section of the PPT page, the explanation text information needs to be marked with custom identifiers as the start and end markers of the explanation text.
[0078] S300. Parse the pre-processed PPT presentation document to obtain the set of graphic elements and explanatory text information contained in each slide of the PPT presentation document;
[0079] In this embodiment, the preprocessed PPT presentation document is parsed to obtain the set of graphic elements and explanatory text information contained in each slide of the PPT presentation document. The specific steps include:
[0080] Load the uploaded PPT document, read the PPT document structure information, and obtain the set of slide pages contained therein;
[0081] Iterate through the collection of slideshow pages, generate a thumbnail for each slideshow page, and save it as a local image file;
[0082] Read the set of graphic elements contained in each slide page, obtain the contained dynamic graphic elements, including dynamic graphic elements such as animated images, audio, and video, obtain the coordinates, width, height, and embedding file of the dynamic graphic elements in the page, obtain the embedding file of the dynamic graphic elements from the PPT document structure and save it to the specified local directory.
[0083] Retrieve the preset explanatory text information contained in each slide page. If a page does not have preset explanatory text information, set the preset text information of the corresponding page to empty.
[0084] Display the PowerPoint slides as a list within the video scene editing unit.
[0085] S400. Edit the video scenes of the parsed PPT presentation document, including at least creating new video scenes, setting preset scene templates, setting slide display settings, setting digital human image, setting digital human voice, setting narration text, and setting video output.
[0086] In this embodiment, the video scenes of the parsed PPT presentation document are edited. The specific steps include:
[0087] Create a new video scene and select a preset scene template. The preset scene template defines a set of pre-set PPT presentation scene templates, including background images, slides and digital human display positions. It also supports user-defined settings for background fill images, resolution, frame rate and video format parameters.
[0088] In the slide display settings, select the PPT slide page and adjust its position, width, and height within the video scene.
[0089] Digital human character settings: Select a suitable digital human character and adjust its position, width, and height in the video scene;
[0090] Digital human voice settings: Select a suitable preset voice, and adjust the volume, speech rate, and tone of the voice.
[0091] The presentation text settings allow for secondary editing of the presentation text on the slides, and support for inserting a pause at a specific position in the presentation text;
[0092] Dynamic element settings allow you to configure playback control parameters for dynamic elements such as videos, audio, and animated images on the page. When the page contains multiple video elements, you need to set their playback order. When the duration of the digital human's narration video differs from the playback duration of other video elements on the page, select a preset playback control strategy. When the page contains multiple audio elements, you need to set their playback order and whether each audio element should loop. When the page contains multiple animated image elements, you need to set the playback order and playback time parameters for each animated image element.
[0093] S500. Using the explanatory text information, digital human image, and digital human voice of a single slide in a PPT presentation document as input data, generate a first target video. The first target video is a virtual digital human explanation video file generated by combining the virtual digital human image with sound samples and single-page explanatory text.
[0094] In this embodiment, the first target video is generated using the explanatory text information, digital human image, and digital human voice of a single slide in a PPT presentation document as input data. Specific steps include:
[0095] The system performs speech synthesis on the input PowerPoint presentation text and digital human voice for a single page, generating an audio file. The speech synthesis process includes: performing word segmentation, text analysis, and normalization preprocessing on the input text to generate phoneme sequences, sentiment indicators, and prosodic feature annotations; encoding the text using a pre-trained large-scale text model to generate semantic vectors; inputting the semantic vectors into an autoregressive transformer to generate speech tags, which are discrete representations after vector quantization, with each tag representing a basic unit in the speech signal; optimizing the speech tags using conditional flow matching technology to adjust and optimize their acoustic features, converting the optimized speech tags into Mel spectra; inputting the Mel spectra into a pre-trained vocoder to generate continuous speech waveforms; performing volume adjustment postprocessing on the generated speech waveforms; and outputting the processed speech waveforms as the final audio presentation file.
[0096] Based on the input digital human image and audio narration file, a digital lip-syncing operation is performed to generate a virtual digital human narration video. The digital lip-syncing operation is implemented using a multi-lip-syncing algorithm framework to control the virtual digital human's lip movements and facial expressions to align and match with the input audio file, achieving the effect of lip-syncing between the digital human's narration audio and lip movements.
[0097] S600. Extract the digital human foreground target from the first target video, remove the video background, and generate a second target video, wherein the second target video is a digital human explanation video file with a transparent background;
[0098] In this embodiment, a digital human foreground target is extracted from the first target video to generate a second target video. Specific steps include:
[0099] The input digital human explanation video is decoded to obtain image data for each frame;
[0100] A pre-trained image segmentation model is used to segment the foreground of each frame of the image to obtain a segmentation mask for the foreground object. The segmentation mask is a binary image used to distinguish between the foreground and the background.
[0101] Foreground objects are extracted from the original image based on the segmentation mask, and the contours of the foreground objects are adjusted to make them smoother and more refined using morphological operations and contour fitting algorithms.
[0102] Create a transparent background image with the same resolution as the original image, and overlay the extracted foreground object image onto the transparent background image to generate a digital human image with a transparent background;
[0103] Each frame, after removing the background, is re-merged into a single complete video file. Finally, the video file is merged with the audio of the original video file to output a digital human narration video with a transparent background. The digital human narration video file with a transparent background is in the WebM video container format encoded with VP9.
[0104] S700. Perform audio-video separation processing on the second target video, extract the audio file and input it into the subtitle generation unit to generate a subtitle file in standard subtitle format;
[0105] In this embodiment, the SRT subtitle file format is used to perform audio-video separation processing on the second target video to generate a subtitle file in standard subtitle format. The specific steps include:
[0106] The input audio files are preprocessed using noise reduction, filtering, and gain control techniques to improve audio quality;
[0107] A pre-trained speech recognition model is used to perform speech recognition on the pre-processed audio, converting the speech into text and generating a timestamp for each word.
[0108] Extract the start time, end time, and text content of each paragraph;
[0109] The extracted timestamps and text content are formatted according to the SRT format to generate a subtitle file in standard subtitle format.
[0110] S800. Taking the set of graphic elements contained in a single slide page and the digital human explanation video as input, perform a video layer compositing operation to generate a third target video, wherein the third target video is a video file of a virtual digital human explaining a single slide page;
[0111] The solution implemented in this embodiment can support the rendering and display of static graphic elements such as background images, background colors, and slide thumbnails contained in the slide page, and can also support the rendering and display of dynamic graphic elements such as dynamic images, videos, audio, and subtitles. The third video generation unit generates a third target video by overlaying clipping layers, including the following steps:
[0112] Add a background layer to the clip layer, and load a background image or background color into the background layer;
[0113] Add a static PPT element rendering layer to the clipping layer to load PPT thumbnails;
[0114] Add PPT motion graphics rendering layers to the clipping layer. The number of PPT motion graphics rendering layers should match the number of motion graphics contained in the PPT page. Each motion graphics rendering layer loads and displays one motion graphics element contained in the PPT page, such as video, audio, and animated images. Specifically, based on the coordinates, width, height, embedded files, playback controls, and other information of the set of motion graphics elements in the PPT page, load the corresponding embedded file resources in the video, audio, and animated image rendering layers, calculate coordinate transformations, move the motion graphics rendering layers to the corresponding positions in the video output scene, and adjust the display width and height of the motion graphics rendering layers.
[0115] Add a digital human rendering layer to the editing layer to load the digital human narration video file;
[0116] Add a foreground graphic rendering layer to the clipping layer. The number of foreground graphic rendering layers should match the number of foreground graphics in the PPT page, and they should be stacked and displayed in ascending order of index.
[0117] Add a subtitle rendering layer to the editing layer to load the subtitle file and display the video subtitles;
[0118] The above video synthesis process is completed automatically by the program and does not require manual operation by the user.
[0119] When creating animated graphic elements in a PowerPoint presentation, the animated graphic elements need to be processed using video compositing control based on the user's pre-set editing options.
[0120] Specifically, when performing video compositing on video elements contained in a PPT page, multiple video elements are spliced and composited in sequence according to the playback parameters set by the user, or multiple video elements are superimposed and composited at specified coordinate positions; when the duration of the digital human's presentation video is inconsistent with the playback duration of the video elements contained on the page, the corresponding video compositing operation is executed according to the playback control strategy set by the user.
[0121] Specifically, when performing video compositing on audio elements contained in a PPT page, it is necessary to control the playback order of the audio elements and whether each audio element loops, based on the editing options pre-selected by the user.
[0122] Specifically, when compositing videos of animated image elements in a PPT presentation, it is necessary to control the playback order and playback time of each animated image element according to the editing options selected by the user.
[0123] S900. Taking the third target video set and video transition animation set generated from all slide pages of the PPT presentation document as input, perform a video segment merging operation to generate a fourth target video, wherein the fourth target video is a video file in which a virtual digital human explains the entire PPT presentation document.
[0124] Specifically, the input third target video set is first used to generate video editing units. Intro and outro transition animations are added to each video editing unit. Then, the video clips are merged, and a complete video file of the virtual digital human presenting the PPT document is output. Preferably, the fourth video generation unit can be configured with multi-threading parameters to perform the video clip merging operation in a parallel processing manner, improving the speed and efficiency of video synthesis.
[0125] This embodiment provides a method for creating a virtual digital human presentation PPT. The method involves: generating a PPT presentation document based on user requirements; preprocessing the PPT presentation document to generate explanation text information for each page; parsing the preprocessed PPT presentation document to obtain the set of graphic elements and explanation text information contained in each slide; editing the video scenes of the parsed PPT presentation document; generating a first target video using the explanation text information, digital human image, and digital human voice of a single slide as input data; extracting the digital human foreground target from the first target video to generate a second target video; performing audio-video separation processing on the second target video to generate a subtitle file in standard subtitle format; generating a third target video using the set of graphic elements and the digital human explanation video of a single slide as input; and generating a fourth target video using the set of third target videos generated from all slides of the PPT presentation document and the set of video transition animations as input.
[0126] This embodiment can solve the problem that existing publicly available technical solutions can only display static PPT page elements and cannot display dynamic elements such as audio, video, and animated images contained in PPT pages. It has better compatibility with PPT presentation documents and can effectively enhance the presentation effect of PPT documents.
[0127] This embodiment can solve the problem of low generation efficiency in existing publicly available technical solutions. By adopting a hardware deployment method of one machine with multiple cards or a graphics card cluster, and by splitting the PPT document and using a multi-task parallel processing combined with multi-threaded concurrent processing operation for video synthesis, it can give full play to the advantages of parallel computing, effectively improve the generation speed and performance of PPT presentation videos, and at the same time have good scalability.
[0128] Based on the same inventive concept, embodiments of the present invention also provide a system for creating virtual digital human presentation slides, such as... Figure 2 It includes: a PPT document generation unit, a document preprocessing unit, a document parsing and processing unit, a video scene editing unit, a first video generation unit, a second video generation unit, a subtitle generation unit, a third video generation unit, and a fourth video generation unit; among which:
[0129] The PPT document generation unit is used to generate PPT presentation documents according to user needs;
[0130] The document preprocessing unit is used to preprocess the PPT presentation document and generate explanatory text information for each page of the PPT presentation document;
[0131] The document parsing and processing sheet is used to parse and process the pre-processed PPT presentation document to obtain the set of graphic elements and explanatory text information contained in each slide of the PPT presentation document;
[0132] The video scene editing unit allows editing of video scenes in the parsed PPT presentation document, including at least creating new video scenes, setting preset scene templates, setting slide display settings, setting digital human appearance, setting digital human voice, setting narration text, and setting video output.
[0133] The first video generation unit is used to generate a first target video by taking the explanatory text information, digital human image and digital human voice of a single slide page in a PPT presentation document as input data. The first target video is a virtual digital human explanation video file generated by combining the virtual digital human image with sound samples and single-page explanatory text.
[0134] The second video generation unit is used to extract the digital human foreground target from the first target video, remove the video background, and generate a second target video, wherein the second target video is a digital human explanation video file with a transparent background.
[0135] The subtitle generation unit is used to perform audio-video separation processing on the second target video, extract the audio file and input it into the subtitle generation unit to generate a subtitle file in standard subtitle format;
[0136] The third video generation unit is used to take the set of graphic elements contained in a single slide page and the digital human explanation video as input, perform a video layer compositing operation, and generate a third target video, wherein the third target video is a video file of a virtual digital human explaining a single slide page;
[0137] The fourth video generation unit is used to take the third target video set and video transition animation set generated from all slide pages of the PPT presentation document as input, perform a video segment merging operation, and generate a fourth target video. The fourth target video is a video file in which a virtual digital human explains the entire PPT presentation document.
[0138] The specific working methods of the PPT document generation unit, document preprocessing unit, document parsing and processing unit, video scene editing unit, first video generation unit, second video generation unit, subtitle generation unit, third video generation unit, and fourth video generation unit have been described in detail in the above-mentioned method for creating a virtual digital human to explain PPT, and will not be repeated here in this embodiment.
[0139] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0140] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0141] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0142] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0143] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0144] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0145] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0146] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0148] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.
Claims
1. A method for making a virtual digital human to explain a PPT, characterized in that, The method comprises the following steps: According to user requirements, generate a PPT presentation document; Pretreat the PPT presentation document to generate the explanation text information of each page of the PPT presentation document; Pretreated PPT presentation document is analyzed and processed to obtain the set of graphic elements and explanation text information contained in each slide page of the PPT presentation document; Edit the video scene of the PPT presentation document after analysis, including at least new video scene, preset scene template setting, slide display setting, digital human image setting, digital human voice setting, explanation text setting and video output setting; Take the explanation text information, digital human image and digital human voice of a single slide page in the PPT presentation document as input data to generate a first target video, which is a virtual digital human explanation video file generated by combining a virtual digital human image with a sound sample and a single page explanation text; Extract the digital human foreground target from the first target video to remove the video background and generate a second target video, which is a digital human explanation video file with a transparent background; Perform audio and video separation processing on the second target video, extract the audio file and input it into the subtitle generation unit to generate a subtitle file in standard subtitle format; Take the set of graphic elements contained in a single slide page and the digital human explanation video as input, perform video layer synthesis operation to generate a third target video, which is a video file of a virtual digital human explaining a single slide page; Take the set of third target videos generated by all slide pages of the PPT presentation document and the video transition animation set as input, perform video segment merging operation to generate a fourth target video, which is a video file of a virtual digital human explaining the entire PPT presentation document.
2. The method of claim 1, wherein, According to user requirements, generate a PPT presentation document, the specific steps include: according to the PPT theme information and generation conditions input by the user, execute the first preset workflow to generate the PPT presentation document meeting the user's requirements; The first preset workflow is based on a workflow automation tool, defines multiple workflow nodes, and is executed in a workflow automation manner to realize the automatic generation of PPT presentation document; The first preset workflow includes a first input node, which is used to receive the PPT theme information and generation conditions input by the user; The first preset workflow also includes an AI agent node, which is bound with a specific large language model, inputs the information of the first input node as the prompt word of the large language model, calls the large language model through API interface, and generates PPT outline and content data in a specified format; The first preset workflow also includes a data conversion node, which is used to analyze the PPT content generated by the large model and convert its data format into PPT document structure; The first preset workflow also includes a code execution node, which inputs the PPT document data into a PPT document generation tool to complete the generation of PPT document.
3. The method of claim 1, wherein, The PPT presentation document is preprocessed to generate the explanation text information of each page of the PPT presentation document, and the specific steps include: taking the PPT presentation document as input, executing a second preset workflow, generating the explanation text information of each page of the PPT, and saving the generated explanation text information to the note bar position of the PPT page; the second preset workflow is based on a workflow automation tool, defines a plurality of workflow nodes, and realizes the automatic generation of the PPT page explanation text in the form of workflow automation execution; the second preset workflow includes a file reading node, and the file reading node is used to read the PPT file content; the second preset workflow also includes a code execution node, which parses the document structure of the read PPT file content, obtains a PPT slide page set, and iteratively traverses and reads the text information contained in each slide page; the second preset workflow also includes an AI agent node, which is bound with a specific large language model, inputs the text information of the PPT page combined with a preset prompt word, calls the large language model in the form of an API interface, and generates the explanation text information of each PPT page; the second preset workflow also includes a code execution node, which adds the explanation text information generated by the large language model to the note bar position of the PPT page; the second preset workflow also includes a file writing node, and the file writing node is used to save the PPT file content.
4. The method of claim 1, wherein, The preprocessed PPT presentation document is parsed and processed to obtain a set of graphic elements and explanation text information contained in each slide page of the PPT presentation document, and the specific steps include: Load the uploaded PPT document, read the PPT document structure information, and obtain a set of slide pages contained; Traverse the slide page set, generate a thumbnail for each slide page, and save it as a local picture file; Read the set of graphic elements contained in each slide page, obtain the dynamic graphic elements, which include dynamic pictures, audio, video and other dynamic graphic elements, obtain the coordinate, width, height, embedded file and other attribute information of the dynamic graphic elements in the page, and save the embedded file of the dynamic graphic elements to a specified local directory from the PPT document structure; Obtain the preset explanation text information contained in each slide page, and if the page does not have preset explanation text information, set the preset text information of the corresponding page to be empty; Display the PPT slide pages in the form of a list in the video scene editing unit.
5. The method of claim 1, wherein, Edit the video scene of the parsed PPT presentation document, and the specific steps include: New video scene, select a preset scene template, the preset scene template defines a set of pre-set PPT explanation scene templates, including background picture, slide and digital person display position, and also supports user-defined setting of background fill picture, resolution, frame rate and video format parameters; Slide display setting, select PPT slide page, adjust the position, width and height of PPT slide page in the video scene; Digital human image setting, selecting a suitable digital human image, adjusting the position, width and height of the digital human image in the video scene; Digital human voice setting, selecting a suitable preset voice, supporting setting the volume, speed and tone of the voice; Explanation text setting, supporting secondary editing of the explanation text of the slide page, and supporting inserting a pause time at a specific position of the explanation text; Dynamic element setting, setting the play control parameters of dynamic elements such as video, audio and dynamic pictures contained in the page; when the page contains multiple video elements, the play order of the multiple video elements needs to be set; when the digital human explanation video length and the video element play length contained in the page are inconsistent, a preset play control strategy is selected; when the page contains multiple audio elements, the play order of the multiple audio elements and whether each audio element is played in a loop are set; when the page contains multiple dynamic picture elements, the play order and play time parameters of each dynamic picture element are set.
6. The method of claim 1, wherein, Taking the explanation text information, digital human image and digital human voice of a single slide page in a PPT presentation document as input data, a first target video is generated, and the specific steps include: Performing speech synthesis operation on the input explanation text and digital human voice of the PPT single page to generate an explanation audio file; the speech synthesis operation includes: performing word segmentation, text analysis and normalization preprocessing operation on the input explanation text to generate phoneme sequence, sentiment tendency and prosody feature labeling information; using a pre-trained text large model to encode the text to generate a semantic vector; inputting the semantic vector into a self-attention transformer to generate a speech mark, which is a discrete representation after vector quantization processing, and each mark represents a basic unit in the speech signal; using conditional flow matching technology to optimize the speech mark, adjusting and optimizing its acoustic characteristics, and converting the optimized speech mark into a mel spectrum; inputting the mel spectrum into a pre-trained vocoder to generate continuous speech waveform; performing volume adjustment post-processing operation on the generated speech waveform; outputting the processed speech waveform as the final audio explanation file; Performing digital human mouth shape synchronization operation according to the input digital human image and audio explanation file to generate a virtual digital human explanation video; the digital human mouth shape synchronization operation is implemented by using multiple mouth shape synchronization algorithm frameworks, which is used to control the alignment and matching processing of the virtual digital human's lip movement and facial expression with the input audio file, and realize the effect of synchronization of digital human explanation audio and lip shape.
7. The method of claim 1, wherein, Extracting the digital human foreground target from the first target video to generate a second target video, and the specific steps include: Performing video decoding on the input digital human explanation video to obtain each frame of image data; Using a pre-trained image segmentation model to perform foreground segmentation on each frame of image to obtain a segmentation mask of the foreground object, which is a binary image for distinguishing foreground and background; According to the segmentation mask, the foreground object is extracted from the original image, and morphological operation and contour fitting algorithm are used to adjust the contour of the foreground object to make it more smooth and fine; A transparent background image with the same resolution as the original image is created, and the extracted foreground object image is superimposed onto the transparent background image to generate a digital human image with a transparent background; The background-removed images of each frame are recombined into a complete video file, and finally the video file is merged with the audio of the original video file to output a digital human explanation video with a transparent background. The digital human explanation video file format uses the VP9 video encoding webm video container format.
8. The method of claim 1, wherein, The second target video is subjected to audio and video separation processing to generate a subtitle file in standard subtitle format, including the following steps: Using noise reduction, filtering, and gain control preprocessing techniques to improve the audio quality of the input explanation audio file; Using a pre-trained speech recognition model to perform speech recognition on the preprocessed audio, converting speech to text while generating timestamps for each word; Extracting the start time, end time, and text content of each paragraph; The extracted timestamps and text content are formatted according to the SRT format to generate a subtitle file in standard subtitle format.
9. The method of claim 1, wherein, The set of graphical elements contained in a single slide page and the digital human explanation video are used as input to generate a third target video using the clip layer superposition method, including the following steps: Add a background layer to the clip layer and load a background picture or background color in the background layer; Add a PPT static element rendering layer to the clip layer to load PPT thumbnails; Add a PPT dynamic graphics rendering layer to the clip layer, and the number of PPT dynamic graphics rendering layers should be consistent with the number of dynamic graphics contained in the PPT page. Each dynamic graphics rendering layer loads and displays a dynamic graphic element contained in a PPT page. According to the coordinates, width, height, embedded files, and playback control information of the dynamic graphic element set obtained from the PPT page, load the corresponding embedded file resources in the video, audio, and dynamic picture rendering layers. Calculate the coordinate transformation and move the dynamic graphics rendering layer to the corresponding position in the video output scene. Adjust the display width and height of the dynamic graphics rendering layer; Add a digital human rendering layer to the clip layer to load the digital human explanation video file; Add a foreground graphics rendering layer to the clip layer, and the number of foreground graphics rendering layers should be consistent with the number of foreground graphics contained in the PPT page. They are displayed in order from small to large according to the index. Add a subtitle rendering layer to the clip layer to load and display the video subtitles.
10. A system for making a virtual digital human explain a PPT, adopting any of the methods in claims 1-9, characterized in that, It includes: A PPT document generation unit, a document preprocessing unit, a document analysis processing unit, a video scene editing unit, a first video generation unit, a second video generation unit, a subtitle generation unit, a third video generation unit, and a fourth video generation unit. Among them: The PPT document generation unit is used to generate a PPT presentation document according to user requirements; The document preprocessing unit is used to preprocess the PPT presentation document to generate explanation text information for each page of the PPT presentation document; The document analysis processing unit is configured to analyze a preprocessed PPT presentation document to obtain a set of graphical elements and a lecture text information contained in each slide page of the PPT presentation document. The video scene editing unit is configured to edit a video scene of the analyzed PPT presentation document, including at least new video scene creation, preset scene template setting, slide display setting, digital human image setting, digital human voice setting, lecture text setting, and video output setting. The first video generation unit is configured to generate a first target video by taking the lecture text information, the digital human image, and the digital human voice of a single slide page of the PPT presentation document as input data, wherein the first target video is a virtual digital human lecture video file generated by combining a virtual digital human image with a sound sample and a single-page lecture text. The second video generation unit is configured to extract a digital human foreground object from the first target video, remove a video background, and generate a second target video, wherein the second target video is a digital human lecture video file with a transparent background. The subtitle generation unit is configured to perform audio-video separation processing on the second target video, extract an audio file, and input the audio file into the subtitle generation unit to generate a subtitle file in a standard subtitle format. The third video generation unit is configured to take the set of graphical elements and the digital human lecture video as input, perform video layer synthesis, and generate a third target video, wherein the third target video is a video file in which a virtual digital human explains a single slide page. The fourth video generation unit is configured to take a set of third target videos generated by all slide pages of the PPT presentation document and a set of video transition animations as input, perform video segment merging, and generate a fourth target video, wherein the fourth target video is a video file in which a virtual digital human explains the entire PPT presentation document.
Citation Information
Cited By
WebM protocol low-delay video and audio translation and subtitle optimization method and system
CN121585844A