A method and related equipment for videoizing presentations

By generating corresponding explanatory text and audio information for the slides, and combining it with digital human information, the problem of mechanical and monotonous presentation video methods is solved, thereby improving learning efficiency and interactivity.

CN115630177BActive Publication Date: 2026-03-13SHENZHEN SHANJIAN INTELLIGENT SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-20
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing methods for converting presentations into videos often result in mechanical and monotonous presentations that fail to improve students' learning efficiency.

Method used

By generating narration text for each slide, combining audio and digital human information, the system controls the digital human to give narration, and adds transition information when switching slides to improve the coherence and interactivity of the narration.

Benefits of technology

It makes presentations more engaging and dynamic, enhancing students' learning interest and efficiency. The digital human's transition information during slide changes makes the presentation process more coherent, allowing students to better connect the knowledge points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115630177B_ABST
    Figure CN115630177B_ABST
Patent Text Reader

Abstract

This invention discloses a method and related equipment for videoizing presentations. The method includes: acquiring a presentation document to be videoized; generating explanatory text for each slide in the presentation document; generating explanation information for each slide based on the explanation text, wherein the explanation information includes first audio information and first digital human information; generating transition information between slides based on the playback order of the slides and the explanation text, wherein the transition information includes second audio information and second digital human information; and playing the presentation document when a playback command is detected, and controlling a preset digital human to provide explanations based on the first audio information, the second audio information, the first digital human information, and the second digital human information. This invention can increase the flexibility of presentation playback and improve the user's learning efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method and related equipment for videoizing presentations. Background Technology

[0002] In current teaching and presentation processes, presentation documents have become a primary tool for teachers. These documents, rich in text and images, can vividly and flexibly display more information, enhancing students' absorption of knowledge.

[0003] After creating a presentation document, teachers still need to personally explain it, which requires a significant amount of time and effort. Therefore, some have suggested converting the notes on each page of the presentation document into audio, and then generating a video by matching the audio with the slides. However, audio generally just mechanically reads out the text and notes in the presentation document, which is too mechanical and fails to create a sense of engagement for students, resulting in low learning efficiency. Summary of the Invention

[0004] The technical problem this invention aims to solve is that the resulting video presentations are mechanical and monotonous, failing to improve students' learning efficiency. To address the shortcomings of existing technologies, this invention provides a method and related equipment for video presentations.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] A method for videoizing a presentation, the method comprising:

[0007] Obtain the presentation document to be video-based;

[0008] Based on the presentation document, generate the explanatory text for each slide in the presentation document;

[0009] Based on the explanatory text, generate the corresponding explanation information for the slide, wherein the explanation information includes first audio information and first digital human information;

[0010] Based on the playback order of the slides and the narration text, transition information between the slides is generated, wherein the transition information includes second audio information and second digital human information;

[0011] When a playback command is detected, the presentation document is played, and a preset digital human is controlled to provide narration based on the first audio information, the second audio information, the first digital human information, and the second digital human information.

[0012] The method for videoizing a presentation document, wherein generating the explanatory text for each slide in the presentation document includes:

[0013] For each slide in the presentation document, generate a first text based on the text information in the slide, and / or generate a second text based on the image information in the slide;

[0014] Based on the display order of the text information and the image information, the first copy and the second copy are organized to obtain the explanatory copy.

[0015] The method for videoizing a presentation, wherein generating the second text based on the image information in the slide includes:

[0016] The image information is subjected to image recognition to obtain the image type;

[0017] Based on the image type, the text information is matched to obtain the target word;

[0018] A second copy is generated based on the target words and the image type.

[0019] The method for videoizing the presentation slides, wherein the first digital human information includes first performance data and first lip-sync data; and the step of generating narration information corresponding to the slides based on the narration text includes:

[0020] For each of the aforementioned explanatory texts, sentiment recognition is performed on the text to obtain the text sentiment; and,

[0021] The explanatory text is converted into audio to obtain initial audio information corresponding to the explanatory text.

[0022] Based on the emotional content of the text, the first performance data corresponding to the explanatory text is determined, and the pitch of the initial audio information is adjusted to obtain the first audio information;

[0023] Based on the first audio information, determine the first lip-sync data corresponding to the narration text.

[0024] The method for videoizing a presentation includes generating transition information between slides based on the playback order of the slides and the narration text, which includes:

[0025] Based on the text sentiment corresponding to the Nth slide and the text sentiment corresponding to the (N+1)th slide, determine the transition sentiment corresponding to the Nth slide, where N is a positive integer;

[0026] Based on the explanatory text corresponding to the Nth slide and the explanatory text corresponding to the N+1th slide, determine the transition text corresponding to the Nth slide;

[0027] Based on the transition mood and the transition text, generate second audio information and second digital human information corresponding to the Nth slide.

[0028] The method for videoizing the presentation document, wherein, after playing the presentation document when a playback command is detected, and controlling a preset digital human to provide narration based on the first audio information, the second audio information, the first digital human information, and the second digital human information, further includes:

[0029] When a question is detected, semantic recognition is performed on the question to obtain the question semantics;

[0030] Information retrieval is performed on the semantics of the question to obtain feedback information;

[0031] Based on the feedback information, the digital human is controlled to answer the question.

[0032] The method for videoizing the presentation, wherein the step of performing information retrieval on the semantics of the question to obtain feedback information includes:

[0033] Based on the semantics of the question, the explanatory text is retrieved to obtain the target information;

[0034] Based on the target information, feedback information is generated.

[0035] The method for videoizing the presentation document, wherein, when a playback command is detected, the presentation document is played, and a preset digital human is controlled to provide narration based on the first audio information, the second audio information, the first digital human information, and the second digital human information, the method includes:

[0036] When a playback command is detected, the presentation document, the first audio information, and the second audio information are played; and...

[0037] Based on the first performance data and the second performance data, control the digital human to perform corresponding actions; and,

[0038] Based on the second performance data, the digital human is controlled to display the corresponding lip movements.

[0039] A computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the presentation videoization method as described above.

[0040] A terminal device includes: a processor, a memory, and a communication bus; the memory stores a computer-readable program that can be executed by the processor;

[0041] The communication bus enables communication between the processor and the memory;

[0042] When the processor executes the computer-readable program, it implements the steps in the presentation videoization method described above.

[0043] Beneficial Effects: This invention provides a method for videoizing presentations. First, based on the presentation document to be videoized, a script is generated to explain the content. Then, based on the script, audio information and digital human information are generated. The audio information is used to play the script, while the digital human information complements the presentation, making the explanation more lively and dynamic. Simultaneously, transition information is added to slide transitions. When switching from one slide to the next, transitional audio and digital human information are added, making the transitions smoother and allowing students to connect the preceding and following knowledge points. Upon detecting a play command, the presentation document is played, and simultaneously, the digital human is controlled to play the first and second audio information and execute the actions corresponding to the first and second digital human information. Attached Figure Description

[0044] Figure 1 A flowchart illustrating the method for videoizing a presentation provided by this invention.

[0045] Figure 2 This is a schematic diagram illustrating the generation of explanatory text in the video-based presentation method provided by the present invention.

[0046] Figure 3 This is a schematic diagram of the overall process of the video-based presentation method provided by the present invention.

[0047] Figure 4 The structural schematic diagram of the terminal device provided by the present invention. Detailed Implementation

[0048] This invention provides a method for videoizing presentations. To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the following detailed description, with reference to the accompanying drawings and embodiments, further illustrates the invention. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of the invention.

[0049] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0050] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0051] like Figure 1 As shown, this embodiment provides a method for videoizing presentations. For ease of explanation, a common server is used as the execution subject in the description. The server here can be replaced by a tablet, computer, or other device with data processing capabilities. The method for videoizing presentations includes the following steps:

[0052] S10. Obtain the presentation document to be video-converted.

[0053] Specifically, the first step is to obtain the presentation to be video-converted, which includes multiple slides.

[0054] S20. Based on the presentation document, generate the explanatory text corresponding to each slide in the presentation document.

[0055] Specifically, after obtaining the presentation slides, for each slide, which contains text and / or image information, the corresponding explanatory text is generated based on the text and / or image information.

[0056] For example, the text information in the slides can be used directly as the explanatory text. This text information may include the text displayed directly on the slides or the text content annotated by the creator for the slides.

[0057] Alternatively, the presentation document may contain an outline file that includes the content to be explained. After obtaining the presentation document, the text and image information are first extracted. Then, the outline file is compared with the text and / or image information to determine the explanation content corresponding to each slide in the outline file, thus obtaining the explanation text for each slide.

[0058] Furthermore, the former approach simply reads aloud the content of the presentation document, converting visual content into audio content, without significantly improving information reception and absorption. The latter approach requires the creator to write a detailed outline beforehand, which is time-consuming and labor-intensive. Therefore, this embodiment proposes another method for generating presentation text, specifically including:

[0059] A10. For each slide in the presentation document, generate a first text based on the text information in the slide, and / or generate a second text based on the image information in the slide.

[0060] Specifically, for each slide, extract the text information from that slide. For example... Figure 2 As shown, the text information in a certain slide is "color", "red", "yellow" and "blue". Extract this text information.

[0061] In one method of generating the first text, the extracted text information is directly used as the first text. While simple, this method is overly mechanical. In another method of generating the first text in this embodiment, a preset template document is modified based on the text attribute corresponding to each text information and the text information corresponding to that text attribute to generate the first text. For example, the text attribute corresponding to "red" is "item list". For the text attribute "item list", the template document includes conjunctions, numbering words, etc. Conjunctions include "first", "second", "then", "finally", etc. Numbering words include "first", "first step", "first one", etc. For the item list, the final generated first text is "The first one is red, the second one is yellow, and the third one is blue".

[0062] In one method of generating second text, the slide includes annotations for the images, such as "The image is a schematic diagram of the completed presentation." The second text is generated directly based on these annotations. However, not all images have annotations; currently, images without corresponding annotations are generally ignored. This doesn't effectively convey the meaning of the image to students. Therefore, in a second method of generating second text, the images on the slide are first extracted, and then the images are categorized to determine their corresponding categories. For example, in this case, the images include pure red, pure yellow, and pure blue images. After identification, the corresponding categories are "red image," "yellow image," and "blue image," respectively.

[0063] When an image is present in a slideshow, it should be associated with the text information within the slideshow. First, the text information is segmented into several words. Then, the category corresponding to the image is compared with the words to determine the word corresponding to that category as the "target word." In this example, the determined target words are "red," "yellow," and "blue." Then, based on the target words, a second text corresponding to the image information is generated. The second text can be generated by modifying a second template document based on the target words. For example, the second template document might be "The image is a 'category' diagram for the 'target word'," and in this embodiment, the second text could be "The image is a red image for the red color."

[0064] A20. Based on the display order of the text information and the image information, organize the first copy and the second copy to obtain the explanatory copy.

[0065] Specifically, the display order is generally consistent with the presenter's presentation order. Therefore, based on the display order, the first and second documents are organized to obtain presentation documents that are consistent with the presentation order of the slides during subsequent demonstrations.

[0066] When the slide contains animation tags, the order in which the animation tags appear is used as the display order to organize the first and second text. For example, in this embodiment, the first animation tag corresponds to text information, and the second animation tag corresponds to image information. Therefore, the first text is arranged before the second text to obtain the explanatory text.

[0067] When the slide does not have animation tags, the first and second texts are arranged in the default order to form the explanatory text. When the slide contains multiple pieces of different text information, such as the first method of creating the presentation document on the left and the second method on the right, the first text is arranged in the order of general reading habits, from top to bottom and from left to right.

[0068] In addition, the first copy, the second copy, and the explanatory copy will be displayed on the screen after generation, allowing users to make modifications.

[0069] S30. Generate the corresponding explanation information for the slide based on the explanation text.

[0070] Specifically, a digital human model is pre-defined, capable of performing actions, changing facial expressions, and outputting speech. During the subsequent playback of a presentation document, the digital human model performs in sync with the presentation. Different performance parameters are set within the digital human model, including action parameters, facial expression parameters, etc.

[0071] After obtaining the narration script, it is first transcribed into audio to obtain the first audio information. Simultaneously, the narration script is matched against pre-defined facial expression data for the digital human model to obtain the first digital human information corresponding to the narration script. One matching method involves keyword matching, where different keywords are pre-mapped to different performance data. The narration script is then matched against these keywords; when a match is successful, the performance data corresponding to the matched keywords is used as the first digital human information.

[0072] In addition, the lip movements of the digital human need to change when giving a presentation. This requires control based on the first lip movement data, which can be determined from the first audio data.

[0073] Furthermore, keywords can only respond to words with obvious meanings, such as the keyword "hello" which corresponds to the action data "waving." They cannot accurately reflect more subtle changes. Therefore, in this embodiment, emotion recognition can also be used to determine the first data person information, specifically including:

[0074] B10. For each of the explanatory texts, perform emotion recognition on the explanatory texts to obtain text emotion; and perform audio conversion on the explanatory texts to obtain initial audio information corresponding to the explanatory texts.

[0075] Specifically, on the one hand, sentiment recognition is performed on the explanatory text to obtain the corresponding textual sentiment. This sentiment recognition of the explanatory text involves using algorithms to analyze and extract the emotions expressed in the text, such as happiness or sadness. Naive Bayes-based sentiment recognition methods and vector machine-based sentiment recognition methods are both used to achieve sentiment recognition of the explanatory text and obtain the textual sentiment.

[0076] Simultaneously, the explanatory text is converted from text to speech to obtain the initial audio information corresponding to the explanatory text. For example, TTS (Text-To-Speech) is used for speech synthesis, and TensorFlow is used for text-to-speech conversion.

[0077] B20. Based on the emotional content of the text, determine the first performance data corresponding to the explanatory text, and adjust the pitch of the initial audio information to obtain the first audio information.

[0078] Specifically, different performance parameters are pre-set for different emotion types. For example, if the emotion type is "happy," the corresponding performance parameters include "clapping" for the action and "smiling" for the facial expression. After obtaining the text emotion, the performance parameters corresponding to the emotional type of the text emotion are determined as the first performance data.

[0079] Different emotions result in different tones of voice; for example, a rising tone indicates happiness, while a low tone indicates sadness. In addition to action parameters, different tone parameters are assigned to different emotional types. The tone parameters are determined based on the text's emotion, and then adjusted according to these parameters to obtain the initial audio information.

[0080] B30. Based on the first audio information, determine the first lip-sync data corresponding to the narration text.

[0081] Specifically, the character's lip movements differ depending on the spoken content. To increase the consistency between the digital human's lip movements and the content of the played audio, first lip movement data corresponding to the narration can be generated based on the first audio information. Methods such as voice-driven lip movement and phoneme-driven lip movement can be used.

[0082] S40. Generate transition information between the slides based on the playback order of the slides and the narration text.

[0083] Specifically, the explanation information refers to the information provided for each slide. However, if there is a lack of transition content when switching between slides, the explanation process becomes disjointed, causing a significant sense of fragmentation for students. Therefore, it is necessary to add transition information between slide transitions. As mentioned earlier, the explanation method used in this embodiment is implemented through digital humans. Therefore, the transition information may include the second audio information played by the digital human and the second digital human information of the digital human performing.

[0084] In the first method of generating transition information, multiple transition words or phrases with connecting functions are pre-defined, such as "Next, let's see," "Then," etc. Based on the transition words or phrases, the second audio information and the second digital human information are generated. This process is similar to the method of generating the first audio information and the first digital human information mentioned above, so it will not be described in detail here.

[0085] In the second method of generating transition information, the presentation contains attributes with hierarchical relationships, such as "sections" and "headings," "subheadings," and "body text." Therefore, the presentation text can be organized into a relationship tree based on the hierarchical relationships between and within slides. Then, the position of each slide in the relationship tree is determined according to the playback order, and the corresponding transition text is selected from a preset transition template. For example, if the previous slide contains a heading and the next slide also contains a heading, the corresponding second audio information could be "We previously explained 'A,' now we will explain 'B.'" "A" represents the content of the heading on the previous slide, and "B" represents the content of the heading on the next slide. The transition text is then converted to audio to generate the second audio information. Additionally, multiple transition actions for digital humans can be pre-set, and the second digital human information can be determined based on the second audio information or randomly.

[0086] In the second method of generating transition information, to make the transition smoother, emotional elements can be added to the transition information. This process includes:

[0087] C10. Determine the transition mood corresponding to the Nth slide based on the text sentiment corresponding to the Nth slide and the text sentiment corresponding to the (N+1)th slide.

[0088] Specifically, during the generation of the explanatory text, emotion recognition is performed on the explanatory text corresponding to each slide to obtain its corresponding text emotion. There are also differences between text emotions; small differences, such as "happy" and "joyful," and large differences, such as "sad" and "happy." For cases with large differences, the emotion type to be used during the transition can be determined based on the preceding and following text emotions, i.e., the transition emotion. For example, for "sad" and "happy," the transition emotion would be "peaceful."

[0089] C20. Based on the explanatory text corresponding to the Nth slide and the explanatory text corresponding to the N+1th slide, determine the transition text corresponding to the Nth slide.

[0090] Specifically, based on the explanatory text for the Nth slide and the explanatory text for the (N+1)th slide, the transition text can be generated using the aforementioned method for generating transition information, which will not be elaborated further here. Here, N is a positive integer.

[0091] In addition, for the last slide, a pre-set closing remark can be used as the transition text, such as "That's all for today's lesson. I hope you all studied hard."

[0092] C30. Based on the transition mood and the transition text, generate second audio information and second digital human information corresponding to the Nth slide.

[0093] Specifically, after obtaining the transition mood and transition text, the corresponding intonation parameters are determined based on the emotion type of the transition mood. Then, based on the intonation parameters, the intonation of the audio converted from the transition text is adjusted to obtain the second audio information. Simultaneously, based on the emotion type, the corresponding first performance data is determined from preset performance parameters, and the corresponding second lip-sync data is determined based on the second audio information.

[0094] S50. When a playback command is detected, the presentation document is played, and a preset digital human is controlled to give an explanation based on the first audio information, the second audio information, the first digital human information, and the second digital human information.

[0095] Specifically, when a playback command for a presentation is detected, the presentation document is played. Simultaneously, based on the first audio information, the second audio information, the first digital human information, and the second digital human information, a preset digital human is controlled to explain the content of the presentation document. Therefore, as follows... Figure 3 As shown, in one playback mode, based on the first audio information, the second audio information, the first digital human information, and the second digital human information, a preset digital human is controlled to provide narration and simultaneously present the slides, generating a demonstration video. When a playback command is detected, the demonstration video is played.

[0096] This method produces large files. To reduce file size, links containing the first audio information, second audio information, first digital human information, and second digital human information can be embedded in the presentation document. When a playback command for the presentation document is detected, the links in the presentation document are automatically triggered. These embedded links can be placed on every slide. When it's necessary to start from the middle slide, a playback command is sent to that slide, simultaneously triggering the link embedded in that slide, activating the digital human, and controlling the digital human to give a presentation based on the corresponding first audio information, second audio information, first digital human information, and second digital human information.

[0097] Furthermore, when a playback command is detected, the presentation document, the first audio information, and the second audio information are played. At the same time, based on the first performance data and the second performance data, the digital human is controlled to perform corresponding actions, and based on the second performance data, the digital human is controlled to display corresponding lip movements.

[0098] This embodiment enables flexible and engaging slideshow presentations, enhancing user learning interest and efficiency. Furthermore, in slideshow creation, different methods can be used to prepare the presentation text based on the information provided by the creator, allowing for quick and accurate preparation even without an outline file. Regarding the digital human, different emotions, actions, and expressions are given to the digital human based on the content of the presentation text, providing users with a stronger sense of realism.

[0099] When a presentation is played, users may have questions. Typically, users need to take notes and then look them up, which is inefficient. Therefore, to address this situation, this embodiment also proposes an interactive method, including:

[0100] D10. When a question is detected, semantic recognition is performed on the question to obtain the question semantics.

[0101] Specifically, users can send questions to the device playing the presentation via input devices such as microphones and keyboards. To correctly understand the questions, semantic recognition is first performed on them to obtain the meaning the question intends to express, thus obtaining the question semantics. For example, if the question is "What are the three primary colors?", the semantic recognition result might be "The definition of the three primary colors".

[0102] D20. Perform information retrieval on the semantics of the question to obtain feedback information.

[0103] Specifically, for a user's question, the semantics of the question can be retrieved first to obtain an answer that responds to the user's question, and this answer is used as feedback information. To enable the digital human to answer the question, the feedback information also includes third-party audio information and third-party digital human information. Since the retrieved answer is a standard answer, the intonation and performance parameters corresponding to the third-party audio information can be determined based on the emotion category used by the user before initiating the question.

[0104] When conducting a search, since this solution explains a presentation, user questions raised during the explanation of earlier content may be answered at the end of the text. Therefore, we can first search all the explanation text based on the semantics of the question to obtain the answer corresponding to that question, and use that answer as the target information. Then, based on the aforementioned method, we can generate feedback information corresponding to the target information. For example, if the definition of the three primary colors appears at the end of the presentation, we can directly use the retrieved definition of the three primary colors as the target information. In addition, for search results with excessively long target information, to avoid affecting the normal learning sequence, we can use a pre-set feedback statement as the target information, such as "We will explain the three primary colors later." If no result corresponding to the semantics of the question is found, we can search through a pre-set database or the internet to obtain the search results, and then generate feedback information based on the search results.

[0105] D30. Based on the feedback information, control the digital human to answer the question.

[0106] Specifically, since the feedback information contains third-party audio information and third-party motion information, it is possible to control the digital human to perform certain behaviors and answer the user's questions.

[0107] When students lack a teacher's guidance, such as during self-study, they may struggle to understand the knowledge points presented in the demonstration document. In such cases, students need to consult a teacher or conduct extensive research. This solution, however, can explain the demonstration document while answering user questions, improving the efficiency of students' learning and absorption of knowledge.

[0108] Based on the above-described method for videoizing presentations, this invention also provides a terminal device, such as... Figure 4 As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logical commands stored in the memory 22 to execute the methods described in the above embodiments.

[0109] Furthermore, the logical commands in the aforementioned memory 22 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0110] The memory 22, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program commands or modules corresponding to the methods in the embodiments of this disclosure. The processor 20 executes functional applications and data processing by running the software programs, commands, or modules stored in the memory 22, thereby implementing the methods in the above embodiments.

[0111] The memory 22 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 22 may include high-speed random access memory (RAM) and non-volatile memory. Examples include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks; it may also be a transient computer-readable storage medium.

[0112] Furthermore, the specific process of loading and executing multiple command processors in the aforementioned computer-readable storage medium and terminal device has been described in detail in the above method, and will not be repeated here.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method of videoing a presentation, characterized by, The method comprises: acquiring a presentation document to be videoed; generating a script corresponding to each slide in the presentation document according to the presentation document; generating explanation information corresponding to the slide according to the script, wherein the explanation information comprises first audio information and first digital human information; generating transition information between the slides according to the playing order of the slides and the script, wherein the transition information comprises second audio information and second digital human information; when a playing instruction is detected, playing the presentation document and controlling a preset digital human to give an explanation according to the first audio information, the second audio information, the first digital human information and the second digital human information; wherein the generating of the script corresponding to each slide in the presentation document according to the presentation document comprises: for each slide in the presentation document, generating a first script according to text information in the slide and / or generating a second script according to image information in the slide; arranging the first script and the second script according to a display order corresponding to the text information and the image information to obtain the script; wherein the generating of the second script according to the image information in the slide comprises: performing image recognition on the image information to obtain an image type; matching the text information according to the image type to obtain a target word; generating the second script according to the target word and the image type; wherein the first digital human information comprises first performance data and first lip movement data; the generating of the explanation information corresponding to the slide according to the script comprises: for each script, performing sentiment recognition on the script to obtain a text sentiment; and performing audio conversion on the script to obtain initial audio information corresponding to the script; determining first performance data corresponding to the script according to the text sentiment and adjusting the initial audio information to obtain first audio information; determining first lip movement data corresponding to the script according to the first audio information.

2. The method of claim 1, wherein the video presentation of the presentation is further characterized by: The generating of the transition information between the slides according to the playing order of the slides and the script comprises: determining a transition emotion corresponding to the Nth slide according to a text sentiment corresponding to the Nth slide and a text sentiment corresponding to the N+1th slide, wherein N is a positive integer; determining transition text corresponding to the Nth slide according to a script corresponding to the Nth slide and a script corresponding to the N+1th slide; generating second audio information and second digital human information corresponding to the Nth slide according to the transition emotion and the transition text.

3. The method of claim 1, wherein the video presentation of the presentation is further characterized by, The playing of the presentation document and the controlling of the preset digital human to give an explanation according to the first audio information, the second audio information, the first digital human information and the second digital human information when the playing instruction is detected further comprises: when a question information is detected, performing semantic recognition on the question information to obtain question semantics; Information retrieval is performed on the question semantics to obtain feedback information; According to the feedback information, the digital person controls the question information to answer.

4. The method of claim 3, wherein the video presentation is a slideshow. The information retrieval on the question semantics to obtain feedback information includes: According to the question semantics, the explanation script is retrieved to obtain target information; Based on the target information, feedback information is generated.

5. The method of Claim 1, wherein the video presentation of the presentation is further characterized by, When the play instruction is detected, the presentation document is played, and based on the first audio information, the second audio information, the first digital person information and the second digital person information, a preset digital person is controlled to explain, including: When the play instruction is detected, the presentation document, the first audio information and the second audio information are played; and, According to the first performance data and the second performance data, the digital person is controlled to perform corresponding actions; and, According to the second performance data and the second performance data, the digital person is controlled to display corresponding mouth shapes.

6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores one or more programs which can be executed by one or more processors to implement the steps in the presentation video method of any one of claims 1-5.

7. A terminal device, comprising: Including: Processor, memory and communication bus; The memory stores computer readable programs which can be executed by the processor; The communication bus realizes the connection communication between the processor and the memory; The processor executes the computer readable programs to realize the steps in the presentation video method of any one of claims 1-5.

Citation Information

Patent Citations

  • Method and system for automatically generating demonstration video, equipment and storage medium

    CN111538851A

  • Method and device for interaction control of AI digital human to PPT based on templated editing

    CN114741541A