Slideshow video generation method, device, equipment and storage medium

By automatically processing slide files and generating mind map animations and subtitle files, the problem of time-consuming and labor-intensive manual operations in traditional PPT-to-video conversion is solved, efficient and synchronous video production is achieved, and teaching effects and production efficiency are improved.

CN120017926BActive Publication Date: 2025-09-23深圳市深圳通有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510495501.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-09-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the traditional PPT-to-video method, manually adding subtitles and adjusting the subtitle display time is time-consuming and labor-intensive, and prone to errors, making it difficult to achieve efficient large-scale video production.

Method used

By obtaining the slide file, performing entity extraction and association recognition to generate a mind map animation file, generating the target subtitles and audio files based on the explanation notes information, and performing video synthesis, the addition and time adjustment of subtitles and audio are automatically processed.

Benefits of technology

It realizes automatic video generation, reduces the complexity of manual operations, improves video production efficiency, ensures the synchronization and consistency of subtitles and audio, and enhances teaching effectiveness and the vividness of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017926B_ABST
    Figure CN120017926B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device, and storage medium for generating a slide video, relating to the field of video generation technology. The method comprises: obtaining a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides; performing entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides; generating a target subtitle file and a target audio file based on each of the explanation notes; and performing video synthesis based on each of the slides, each of the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file. The present application can solve the time-consuming, labor-intensive, and error-prone problem of manually adding subtitles and adjusting the display time of subtitles during the process of converting slide files to videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video generation technology, and in particular to a slide video generation method, device, equipment and storage medium. Background Art

[0002] In the current field of multimedia teaching and online training, PPT (slide presentation) is widely used to produce teaching courseware and training materials. With the popularity of video content, converting PPT to video has become a common demand, especially when producing online courses, training videos and remote teaching content.

[0003] However, traditional PPT-to-video conversion methods usually require manual addition of subtitles and adjustment of subtitle display time. This process is not only time-consuming and labor-intensive, but also prone to errors, making it difficult to achieve efficient large-scale video production. Summary of the Invention

[0004] The main purpose of this application is to provide a slide video generation method, device, equipment and storage medium, aiming to solve the problems of manually adding subtitles and adjusting the display time of subtitles in the process of converting PPT to video, which is time-consuming, labor-intensive and error-prone.

[0005] To achieve the above objectives, the present application proposes a method for generating a slideshow video, the method comprising:

[0006] Obtaining a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides;

[0007] Perform entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides;

[0008] Based on the explanation note information, generate a target subtitle file and a target audio file;

[0009] Video synthesis is performed based on each of the slides, each of the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file.

[0010] In one embodiment, the performing entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides includes:

[0011] Inputting each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model;

[0012] Extracting design style information corresponding to the slide in the slide file;

[0013] Based on the entity information, the entity association relationship and the design style information, a mind map animation file corresponding to each of the slides is generated.

[0014] In one embodiment, after inputting each of the slides into a key information recognition model and obtaining the entity information and entity association relationship corresponding to each of the slides output by the key information recognition model, the method further includes:

[0015] Inputting the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtaining the slide value score corresponding to each of the slides output by the knowledge point scoring model;

[0016] For any of the slides, if the slide value score is greater than a preset value score, the step of extracting the design style information in the slide file is performed.

[0017] In one embodiment, generating a target subtitle file and a target audio file based on each of the explanation note information includes:

[0018] Get target timbre information;

[0019] For any of the explanation note information, input the explanation note information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model;

[0020] The explanation note information is parsed to generate a target subtitle file.

[0021] In one embodiment, parsing the explanation note information to generate a target subtitle file includes:

[0022] Parsing and splitting the explanation note information according to a preset splitting rule to obtain a plurality of subtitle short sentences, and determining the subtitle sequence relationship corresponding to each of the subtitle short sentences;

[0023] Inputting each of the subtitle sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files;

[0024] Determining the subtitle playback duration corresponding to each subtitle sentence based on the total playback duration and the audio playback duration corresponding to each subtitle audio file;

[0025] The subtitle short sentences, the subtitle sequence relationship corresponding to the subtitle short sentences, and the subtitle playback time are synthesized to generate a target subtitle file.

[0026] In one embodiment, the performing video synthesis based on the slides, the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file includes:

[0027] Obtaining the slide show order corresponding to each of the slides;

[0028] Determining the animation playback node corresponding to each of the mind map animation files based on the target audio file;

[0029] According to the slide play sequence and the animation play node, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized to generate a target video corresponding to the slide file.

[0030] In one embodiment, determining the animation playback node corresponding to each mind map animation file based on the target audio file includes:

[0031] For any of the mind map animation files, obtaining a first timestamp corresponding to each entity information in the mind map animation file and a relationship between entity playback orders;

[0032] Determine a second timestamp at which each entity information appears in the target audio file;

[0033] Based on the entity playback sequence relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a slide video generation device, which includes:

[0035] An information acquisition module is used to acquire a slide file and a lecturer's voice audio file, wherein the slide file includes a plurality of slides and the lecture notes corresponding to each slide;

[0036] A first generating module is used to perform entity extraction and entity association recognition on each of the slides, and generate a mind map animation file corresponding to each of the slides;

[0037] A second generating module is used to generate a target subtitle file and a target audio file based on each of the lecture note information and the lecturer's timbre audio file;

[0038] The third generating module is used to perform video synthesis on the slides, the mind map animation file, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes a slide video generation device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the slide video generation method as described above.

[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the slide video generation method described above are implemented.

[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the slide video generation method as described above.

[0042] The present application provides a slide video generation method, apparatus, device and storage medium. The slide video generation method obtains a slide file, wherein the slide file includes a plurality of slides and explanation note information corresponding to each of the slides, and then performs entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides. Based on each of the explanation note information, a target subtitle file and a target audio file are generated. Then, based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file, video synthesis is performed to generate a target video corresponding to the slide file, thereby solving the problem of time-consuming, labor-intensive and error-prone manual addition of subtitles and adjustment of subtitle display time in the process of converting slide files to videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0045] Figure 1 A flowchart of the first embodiment of the method for generating a slide video of the present application is provided;

[0046] Figure 2 A flowchart of the second embodiment of the method for generating a slide video of the present application is provided;

[0047] Figure 3 A flowchart of the third embodiment of the method for generating a slide video of the present application is provided;

[0048] Figure 4 A flowchart of the fourth embodiment of the method for generating a slide video of the present application is provided;

[0049] Figure 5 This is a schematic diagram of the module structure of the slide video generation device according to an embodiment of the present application;

[0050] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the slide video generation method in the embodiment of the present application.

[0051] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0052] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0053] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0054] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a big data service platform, a slideshow video generation system, etc. The slideshow video generation system is used as an example to illustrate this embodiment and the following embodiments.

[0055] Based on this, the embodiment of the present application provides a method for generating a slide video. Figure 1 , Figure 1 This is a flow chart of the first embodiment of the method for generating a slide video of this application.

[0056] In this embodiment, the slide video generation method includes steps S11 to S14:

[0057] Step S11, obtaining a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each slide;

[0058] It should be noted that the slide file refers to a presentation file containing several slides, which is usually stored in PPT (PowerPoint), PPTX or other similar presentation formats. It not only contains the visual content of each slide (such as text, charts, pictures, etc.), but also may contain explanation notes related to each slide.

[0059] It should be further clarified that a slide refers to a single page within a slide file, typically used to present a specific theme or knowledge point. Each slide may contain elements such as text, charts, images, and graphics to convey specific information and showcase the content of a particular theme. Explanation notes refer to the textual content associated with each slide within a slide file, used to assist in explaining or illustrating the content of the slide. These notes are typically stored in the slide's notes column and include detailed text descriptions, key points, and supplementary information.

[0060] Step S12, performing entity extraction and entity association recognition on each of the slides, and generating a mind map animation file corresponding to each of the slides;

[0061] It should be noted that the mind map animation file refers to a dynamic display file of the mind map generated based on the content of the slide, which is used to present the knowledge points in the slide in a structured manner, including the nodes, branches and dynamic effects of the mind map (such as node expansion, contraction, highlighting, etc.), which are used to show the logical relationship between knowledge points.

[0062] Specifically, each of the slides is input into the key information recognition model to obtain the entity information and entity association relationship corresponding to each of the slides output by the key information recognition model, and then the design style information corresponding to the slide in the slide file is extracted, so as to generate a mind map animation file corresponding to each of the slides based on the entity information, the entity association relationship and the design style information.

[0063] Step S13, generating a target subtitle file and a target audio file based on each of the explanation note information;

[0064] It should be noted that the target subtitle file refers to a subtitle file generated based on the slide's explanation notes, which is used to display the text of the explanation content in the video. It contains text content synchronized with the explanation audio, and usually uses a timestamp to mark the display time of each subtitle. For example, an SRT (SubRip Subtitle Format, subtitle extraction format) subtitle file contains text content synchronized with the explanation audio.

[0065] It should be further explained that the target audio file refers to an audio file generated based on the slide's explanation notes, which is used to play the explanation content in the video and includes the explanation voice, which is usually generated by text-to-speech technology, or can be a recorded real-person explanation voice.

[0066] Specifically, the target timbre information is obtained, and then for any of the explanation notes information, the explanation notes information and the target timbre information are input into the text-to-speech model to obtain the target audio file output by the text-to-speech model, thereby parsing the explanation notes information and generating a target subtitle file.

[0067] Step S14: Perform video synthesis based on the slides, the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file.

[0068] It should be noted that the target video refers to the final video file generated by combining the slide content, mind map animation, subtitles and audio. It contains the visual content of the slide, mind map animation, subtitles and explanation audio, and is a complete teaching or demonstration video. For example, the target video is a video file in MP4 format, which is used to display the slide content, and uses subtitles and audio to assist in the explanation, while displaying the logical relationship of the knowledge points in the form of a mind map animation.

[0069] Specifically, the slide playback order corresponding to each of the slides is obtained, and then based on the target audio file, the animation playback node corresponding to each of the mind map animation files is determined, so that according to the slide playback order and the animation playback node, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized into a video to generate the target video corresponding to the slide file.

[0070] This embodiment obtains a slide file, wherein the slide file includes several slides and explanation note information corresponding to each slide, then performs entity extraction and entity association recognition on each slide, generates a mind map animation file corresponding to each slide, and then generates a target subtitle file and a target audio file based on each explanation note information. Then, video synthesis is performed based on each slide, each mind map animation file, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file. By generating a mind map animation, complex knowledge points in the slide are presented in a structured manner, helping the audience to better understand and remember the content, clearly showing the logical relationship between knowledge points (such as cause and effect, progression, and parallelism), making it easier for the audience to grasp the overall framework, thereby improving teaching effectiveness and making the video content more vivid and interesting. At the same time, the automatic generation of videos reduces the complexity of manual operations and improves the efficiency of video production, thereby solving the time-consuming and error-prone problems of manually adding subtitles and adjusting the display time of subtitles in the process of converting slide files to videos.

[0071] Based on this, the embodiment of the present application provides a method for generating a slide video. Figure 2 , Figure 2 This is a flow chart of the second embodiment of the slide video generation method of this application.

[0072] In a feasible implementation, the performing entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides includes:

[0073] Step S21, inputting each of the slides into a key information recognition model, and obtaining entity information and entity association relationships corresponding to each of the slides output by the key information recognition model;

[0074] It should be noted that the key information recognition model refers to a model based on artificial intelligence and natural language processing technology, which is used to extract key information from text, such as entities, relationships, concepts, etc., to automatically identify important information in slides, and provide structured data for subsequent mind map generation and video production. It is usually based on deep learning algorithms (such as BERT model, Transformer, etc.) or traditional natural language processing technologies (such as named entity recognition NER, dependency syntax analysis, etc.).

[0075] It should be further explained that the entity information refers to nouns or noun phrases with specific meanings extracted from the slide text, which usually represent specific concepts, objects or themes, such as names of people, places, organizations, terms, concepts, etc. In addition, the entity association relationship refers to the logical relationship between entities, describing the interaction or connection between entities, such as causal relationships, parallel relationships, progressive relationships, and subordinate relationships. For example, the relationship between "artificial intelligence" and "deep learning" can be a "containment relationship" (deep learning is a branch of artificial intelligence), which is used to construct the branch structure of the mind map and display the logical connection between knowledge points.

[0076] Specifically, each of the slides is input into a key information recognition model to obtain the entity information and entity association relationships corresponding to each of the slides output by the key information recognition model. The training process of the key information recognition model includes: collecting and annotating a large amount of slide text data, clarifying the entity information (such as names, terms, concepts, etc.) and entity association relationships (such as causality, subordination, etc.) therein, and then using natural language processing (NLP) technology, such as named entity recognition (NER) and dependency syntax analysis, to preprocess and extract features of the text, so as to select a suitable deep learning architecture (such as BERT, Transformer or its variants), input the annotated data into the model for training, and adjust the model parameters through optimization algorithms (such as Adam or SGD) to minimize the error between the predicted results and the actual annotations to obtain the key information recognition model. At the same time, during the training process, data enhancement, transfer learning, and multi-task learning are used to improve the generalization ability and recognition accuracy of the model.

[0077] Step S22, extracting the design style information corresponding to the slide in the slide file;

[0078] It should be noted that the design style information refers to the visual design features extracted from the slide file, which is used to maintain the consistency of the generated mind map animation with the original slide in visual style. The design style information includes but is not limited to: color, such as the main color, background color, text color, etc. used in the slide; font, such as the font type, font size, bold, italic and other styles used in the slide; layout, such as the layout method of the slide, such as symmetrical, asymmetrical, column, etc.; icons and graphics, such as icons, graphics, charts and other elements used in the slide; animation effects, such as animation effects used in the slide, such as fade in and out, zoom, etc., so as to generate a mind map animation that is consistent with the visual style of the original slide, enhancing the overall sense and professionalism of the video.

[0079] Specifically, in order to reduce computational complexity and improve video generation efficiency, usually only the main color, font type, font size, and other basic elements in the slide file are extracted to generate a mind map. Among them, image processing tools (such as OpenCV or Pillow) can be used to extract the main color of the slide, and then analyze the color distribution of the text, charts and background in the slide. The main color is extracted and determined through a clustering algorithm (such as K-Means), and then a special library (such as python-pptx) is used to parse the PPT file, extract the font information of each slide, and analyze the frequency of font usage to determine the main font and determine the design style information.

[0080] Step S23: generating a mind map animation file corresponding to each slide based on the entity information, the entity association relationship, and the design style information.

[0081] Specifically, the entity information extracted from the slides is used as the core nodes of the mind map to form the basic framework of the mind map. Then, the entity association relationships (such as cause and effect, subordination, parallelism, etc.) are used to determine the logical connections between these nodes. The hierarchy and interaction between knowledge points are clearly displayed through branches and lines, thereby constructing the logical structure of the mind map.

[0082] On this basis, combined with the design style information extracted from the slides (including visual elements such as color, fonts, layout, and icons), the mind map is given a visual style consistent with the original slides, ensuring that the generated animation file maintains a high degree of consistency in appearance with the slides, enhancing the overall visual coherence and professionalism. Furthermore, animation techniques (such as SVG animation, CSS animation, or specialized animation libraries) are used to add dynamic effects to the mind map, such as node expansion, contraction, highlighting, and animated transition effects, making it present in a vivid and intuitive form, helping the audience better understand and remember the slide content. In addition, to keep the expansion of the mind map consistent with the audio, timestamps are added to each node to facilitate subsequent adjustment of the expansion speed and maintain the synchronization of the mind map, audio, and subtitles.

[0083] This embodiment inputs each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model, and then extracts design style information corresponding to the slides in the slide file, thereby generating a mind map animation file corresponding to each of the slides based on the entity information, the entity association relationships and the design style information, thereby clearly displaying the logical relationship between knowledge points (such as cause and effect, progression, parallelism, etc.), helping the audience to better understand and remember the content, especially for complex or information-rich slide content, improving the structuring, logic and comprehensibility of the content, enhancing the visualization effect of the information, and maintaining the consistency of the overall visual style through a unified design style, making the video content more personalized and professional, while enhancing the audience's visual experience, thereby improving production efficiency and quality.

[0084] Based on this, the embodiment of the present application provides a method for generating a slide video. Figure 3 , Figure 3 This is a flow chart of Example 3 of the slide video generation method of this application.

[0085] In a feasible implementation manner, after inputting each of the slides into a key information recognition model and obtaining the entity information and entity association relationship corresponding to each of the slides output by the key information recognition model, the method further includes:

[0086] Step S31, inputting the explanation notes information, entity information and entity association relationship corresponding to each slide into a knowledge point scoring model, so that the knowledge point scoring model outputs a slide value score corresponding to each slide;

[0087] It should be noted that the knowledge point scoring model is used to evaluate the quality and value of slide content. The slide value score refers to a numerical value or grade output by the knowledge point scoring model, which is used to indicate the quality and importance of the slide content. It can be presented in numerical form or graded form. For example, it can be a value between 0 and 1, or a qualitative grade such as "high," "medium," or "low." There is no limitation here and it can be set according to actual circumstances.

[0088] Specifically, the explanation notes, entity information, and entity associations corresponding to each slide are input into a knowledge point scoring model, and the knowledge point scoring model outputs a slide value score corresponding to each slide. The training process of the knowledge point scoring model includes collecting a large amount of labeled slide data, including the slide content and corresponding value scores (e.g., high, medium, low, or specific numerical values). Key features are then extracted from these slides, such as the number of entities, the complexity of entity associations, text length, keyword density, and information entropy, to reflect the knowledge content and logical structure of the slides. This allows the selection of an appropriate machine learning algorithm to construct the model. Common algorithms include deep learning methods (e.g., BERT and Transformer architectures for processing text data), traditional machine learning algorithms (e.g., random forests and support vector machines), or hybrid models. During training, cross-validation and other techniques are used to evaluate model performance, and model parameters are adjusted through optimization algorithms to minimize the error between the predicted scores and the true annotations.

[0089] Furthermore, in order to improve the accuracy of the scoring, more training samples are generated by performing synonym replacement and sentence reorganization on the slide content, and pre-trained language models (such as BERT) are used to extract text features to better capture the semantic information of the slide content. The model is then trained to complete multiple related tasks (such as entity recognition, relationship recognition, and value scoring) at the same time to enhance the model's overall understanding of the slide content. External knowledge bases, such as domain expert knowledge or industry standards, are introduced to provide the model with additional contextual information to help it more accurately evaluate the value of the slide, thereby more accurately identifying the key information and logical structure in the slide, enabling the model to output more accurate value scores and provide a reliable basis for subsequent content screening and processing.

[0090] Step S32: For any of the slides, if the slide value score is greater than a preset value score, then the step of extracting the design style information from the slide file is executed.

[0091] It should be noted that the preset value score refers to a pre-set threshold used to determine whether to further process a slide. The preset value score can be adjusted based on actual needs and application scenarios. For example, for high-quality online courses, a higher threshold (such as 0.8) can be set; for ordinary presentations, a lower threshold (such as 0.5) can be set. The threshold can also be dynamically adjusted based on actual usage to optimize screening results.

[0092] Specifically, for any of the slides, if the slide value score is higher than or equal to a preset value score, the slide is deemed to have sufficient value, and the steps of extracting the design style information in the slide file and generating a mind map animation are performed. Conversely, if the score is lower than a preset threshold, the slide is ignored or skipped. For example, if the preset value score is set to 0.6, if the value score of slide A is 0.9, subsequent processing is performed; if the value score of slide B is 0.3, subsequent processing is skipped and the scoring process for the next slide is entered.

[0093] This embodiment inputs the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtains the slide value score corresponding to each of the slides output by the knowledge point scoring model. Then, for any of the slides, if the slide value score is greater than the preset value score, the step of extracting the design style information in the slide file is executed, so as to evaluate the value of the slide through the knowledge point scoring model, quickly screen out high-value content, avoid unnecessary processing of low-value slides, save time and computing resources, and improve the efficiency of content screening. Only when the value score of the slide is higher than the preset threshold will the design style information be extracted and the mind map animation be generated, thereby ensuring the high quality and high value of the generated content.

[0094] Based on this, the embodiment of the present application provides a method for generating a slide video. Figure 4 , Figure 4 This is a flow chart of Example 4 of the slide video generation method of this application.

[0095] In a feasible implementation manner, generating a target subtitle file and a target audio file based on each of the explanation note information includes:

[0096] Step S41, obtaining target timbre information;

[0097] It should be noted that the target timbre information is used to define data on specific timbre features in the speech synthesis process, including pitch, timbre, speaking speed, intonation and other features of the speech, so that the generated speech sounds closer to a specific speaker and the speech generated by the text-to-speech model is adjusted to a specific timbre, such as imitating the voice of a lecturer or a specific speech style.

[0098] Specifically, in one embodiment, a 20-minute, clear, high-quality audio recording of a lecturer, free of accompaniment and noise, is obtained and an audio file is generated. Based on the audio file, acoustic features are extracted to replicate the lecturer's timbre and obtain target timbre information. Furthermore, users can adjust the target timbre information through audio software to meet their personalized customization needs.

[0099] Step S42: for any of the explanation notes, input the explanation notes and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model;

[0100] It should be noted that the text-to-speech (TTS) model refers to a model based on artificial intelligence and machine learning technologies that can convert text content into speech output and generate natural and fluent speech by analyzing the semantic and grammatical structure of the text.

[0101] Specifically, for any of the explanation note information, the explanation note information and the target timbre information are input into a text-to-speech model to obtain a target audio file output by the text-to-speech model.

[0102] In one embodiment, the characteristic voice of the training instructor is used to convert the text of the explanation notes of each slide into audio through the text-to-speech function of the KAN-TTS framework in the Sambert-Hifigan model, thereby obtaining a set of target audio files in order and the playback time T[i] of each target audio file.

[0103] Step S43: parsing the explanation note information to generate a target subtitle file.

[0104] Specifically, according to the preset splitting rules, the explanation notes information is parsed and split to obtain several subtitle sentences, and the subtitle order relationship corresponding to each subtitle sentence is determined, and then each subtitle sentence is input into the text-to-speech model to obtain the subtitle audio file corresponding to each subtitle sentence output by the text-to-speech model, and the total playback time of each subtitle audio file is determined, so as to determine the subtitle playback time corresponding to each subtitle sentence based on the total playback time and the audio playback time corresponding to each subtitle audio file, and then perform subtitle synthesis on each subtitle sentence and the subtitle order relationship and subtitle playback time corresponding to each subtitle sentence to generate a target subtitle file.

[0105] This embodiment obtains target timbre information, and then for any of the explanation note information, inputs the explanation note information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model, thereby parsing the explanation note information and generating a target subtitle file, thereby enhancing the realism and professionalism of the video content through personalized voice, making it easier for the audience to accept and understand the explanation content, and maintaining consistency between the subtitles and the voice, because the subtitle file is generated based on the explanation note information and fully matches the voice content, ensuring that the audience can accurately understand the content of the voice explanation through the subtitles when watching the video, improving the personalization and consistency of the voice and subtitles, and enhancing the comprehensibility and accessibility of the video.

[0106] In a feasible implementation manner, parsing the explanation note information to generate a target subtitle file includes:

[0107] Step S51, parsing and splitting the explanation note information according to a preset splitting rule to obtain a plurality of subtitle short sentences, and determining the subtitle sequence relationship corresponding to each of the subtitle short sentences;

[0108] It should be noted that the preset splitting rules refer to predefined logic and standards used to split the commentary information into several concise and clear subtitle phrases. Subtitle phrases are short sentences separated from the commentary information using the preset splitting rules and displayed as subtitles in the video. Each subtitle phrase typically contains a complete unit of meaning and is of moderate length (generally no more than 40 characters) to facilitate quick reading by viewers, ensuring clear and coherent subtitles during playback.

[0109] It should be further explained that the subtitle sequence relationship refers to the display order of subtitle sentences during video playback, which reflects the logical relationship between the subtitle sentences, that is, it is consistent with the logical order of the explanation content, ensuring that the audience can understand the content in the correct order when watching the video and the subtitles are logically clear and coherent during playback, avoiding confusion among the audience due to confusion in the subtitle order.

[0110] Specifically, in one embodiment, according to a preset splitting rule, the explanation notes of slide X in the training courseware PPT are parsed, and the explanation notes of the slide on this page are split into several shortest sentences according to pause symbols (e.g., period, comma, semicolon, question mark, line break, etc.), and then the shortest sentences are pieced together into subtitle short sentences. The preset splitting rule is as follows:

[0111] a) If the length of the shortest sentence exceeds 40 characters, it will be directly upgraded to a subtitle short sentence;

[0112] b) If the current shortest sentence does not exceed 40 characters, try to combine it with the next shortest sentence to form a concatenated sentence;

[0113] c) If the length of the spliced ​​short sentence exceeds 40 characters, the current spliced ​​short sentence will be upgraded to a subtitle short sentence;

[0114] d) If the current spelling does not exceed 40 characters, try to combine it with the next shortest sentence to form a spelling to see if the new spelling exceeds 40 characters. If it exceeds 40 characters, the current spelling will be upgraded to a subtitle sentence; if it does not exceed 40 characters, it will be combined with the next shortest sentence to form a new spelling, and this process will continue until the latest spelling exceeds 40 characters and is upgraded to a subtitle sentence.

[0115] e) Record the correspondence and order between the subtitles and PPT pages obtained in the previous step.

[0116] f) Based on the list of subtitle sentences obtained in the previous step, call the TextClip tool to generate subtitle images with transparent backgrounds, and record the correspondence between subtitle images and subtitle sentences, that is, the subtitle order.

[0117] In addition, the explanation notes can be translated into other languages, and then the subtitle audio and subtitle files in the corresponding languages ​​can be generated through the text-to-speech model to support the needs of multilingual audiences.

[0118] Step S52: inputting each of the subtitle sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files;

[0119] It should be noted that the subtitle audio file refers to the audio file generated by converting subtitle sentences through the text-to-speech (TTS) model. Each subtitle sentence corresponds to a subtitle audio file, which contains the voice data of the sentence, and is used to ensure that the display time of the subtitles accurately matches the voice playback time, thereby improving the video viewing experience.

[0120] It should be further explained that the total playback time refers to the total time required to complete the playback of all subtitle audio files, which is obtained by adding the playback time of each subtitle audio file.

[0121] Specifically, continuing with the above example, using the generated characteristic timbre of the training instructor (target timbre information), through the text-to-speech function of the KAN-TTS framework in the Sambert-Hifigan model, the sequential subtitle short sentences generated by splitting each page of the courseware are converted into audio, and a set of sequential subtitle audio files corresponding to the slide on that page, the playback time t[j] of each subtitle audio file, and the total playback time t of all subtitle audio files on that page are obtained.

[0122] Step S53, determining the subtitle playback duration corresponding to each subtitle sentence based on the total playback duration and the audio playback duration corresponding to each subtitle audio file;

[0123] It should be noted that the subtitle playback duration refers to the length of time each subtitle sentence is displayed in the video.

[0124] Specifically, according to a preset duration calculation formula, based on the total playback duration and the audio playback duration corresponding to each of the subtitle audio files, the subtitle playback duration corresponding to each of the subtitle sentences is determined. The preset duration calculation formula is as follows:

[0125] A=T[i] *(t[j] / t)

[0126] Among them, A is the subtitle playback time of each subtitle image, T[i] is the playback time of the target audio file corresponding to the current slide explanation note information, t[j] is the audio playback time of the current subtitle audio file, and t is the total playback time corresponding to each of the subtitle audio files, so that the calculated subtitle playback time can accurately match the voice playback.

[0127] Step S54 , synthesizing the subtitles according to the subtitle sentences, the subtitle sequence relationship corresponding to the subtitle sentences, and the subtitle playback duration to generate a target subtitle file.

[0128] This embodiment parses and splits the explanation and remark information according to preset splitting rules to obtain a plurality of subtitle sentences, determines the subtitle sequence relationship corresponding to each subtitle sentence, and then inputs each subtitle sentence into the text-to-speech model to obtain a subtitle audio file corresponding to each subtitle sentence output by the text-to-speech model. The total playback time of each subtitle audio file is determined, and then the subtitle playback time corresponding to each subtitle sentence is determined based on the total playback time and the audio playback time corresponding to each subtitle audio file. Subtitles are then synthesized based on each subtitle sentence, the subtitle sequence relationship corresponding to each subtitle sentence, and the subtitle playback time to generate a target subtitle file, thereby ensuring that the semantics of each subtitle sentence are complete and concise, improving the accuracy and readability of the subtitles, and determining the display time of the subtitles based on the actual playback time of the audio, ensuring accurate matching of the subtitles with the speech, avoiding errors caused by manually setting the subtitle duration, ensuring that the subtitles are always synchronized with the speech, achieving accurate synchronization of the subtitles with the speech, and improving the efficiency and quality of subtitle production.

[0129] In a feasible implementation manner, performing video synthesis based on each of the slides, each of the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file includes:

[0130] Step S61, obtaining the slide show order corresponding to each slide;

[0131] It should be noted that the slide playback order refers to the order in which the slide pages are displayed in the video, which determines the expansion logic of the slide content when the audience watches the video. The slide playback order is usually consistent with the page order in the slide file, but can also be adjusted according to actual needs (for example, skipping certain pages or rearranging the page order) to ensure that the video content is expanded in a logical order so that the audience can follow the explanation smoothly.

[0132] Specifically, the slide file is converted into a PDF document format, and each page of the generated PDF file is converted into an image format (such as JPEG format) to obtain a set of sequential courseware static images to obtain the slide play order corresponding to each of the slides.

[0133] Step S62: determining the animation play node corresponding to each mind map animation file based on the target audio file;

[0134] It should be noted that the animation playback node refers to the display time point of the mind map animation in the video, which determines when each mind map animation segment starts playing and the duration of the playback. The animation playback node usually corresponds to the voice explanation timestamp in the target audio file to ensure that the rhythm of the animation and the voice explanation are completely matched.

[0135] Specifically, for any of the mind map animation files, the first timestamp corresponding to each entity information in the mind map animation file and the entity playback order relationship are obtained, and then the second timestamp at which each entity information appears in the target audio file is determined, and based on the entity playback order relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

[0136] Step S63 , performing video synthesis on the slides, the mind map animation files, the target subtitle files and the target audio files according to the slide play order and the animation play node, to generate a target video corresponding to the slide file.

[0137] Specifically, the order in which each slide is displayed in the video is determined based on the slide playback sequence. For example, if a slide file contains three slides, the playback order is page 1, page 2, and page 3. Furthermore, based on the animation playback node, the specific display time of the mind map animation on each slide (including only slides that can generate mind map files, i.e., slides whose value score meets a preset threshold) is determined. Assuming that the explanation for slide 1 starts at second 0 and lasts for 10 seconds, the corresponding mind map animation will also start at second 0 and complete within 10 seconds. Simultaneously, the subtitles in the target subtitle file will be displayed synchronously with the timestamps of the audio explanation in the target audio file. For example, if the first half of the explanation (0-5 seconds) corresponds to the first subtitle, and the second half (5-10 seconds) corresponds to the second subtitle, the subtitles will be displayed sequentially during this time period. Finally, these elements (slides, mind map animation, subtitles, and audio) are synthesized using a video editing tool (such as Adobe Premiere, Final Cut Pro, or the Python moviepy library). During the synthesis process, the playback time and order of each element are ensured to be completely consistent with the preset playback order and playback nodes. For example, when the first slide plays to the 5th second, the corresponding mind map animation node and subtitle content will also be updated synchronously, providing a consistent visual and auditory experience for the audience. The resulting target video is not only coherent in content, but also has a precise match of visual and auditory elements, greatly enhancing the educational value and viewing experience of the video.

[0138] The display position of the mind map can be determined based on the current slide's position. For example, a mind map can be fixed in a corner (such as the lower right corner), edge (such as the top or bottom), or halfway across the screen (such as the left half of the screen) during the video to facilitate the continuous presentation of structured information without interfering with the main content of the slide. For example, displaying a mind map in the lower right corner allows viewers to readily review the logical structure of knowledge points; displaying a mind map at the top can provide an overall framework. Alternatively, the position of the mind map can be dynamically adjusted based on the content of the presentation. For example, the mind map can be enlarged and centered during key points; during non-key points, it can be faded out or moved to the edge through animation (such as sliding or zooming) to avoid interfering with the main slide content. The mind map's position can also be intelligently adjusted based on slide content density and key areas. For example, if there is less content on the right side, the mind map can be placed on the right. The mind map's display position can be flexibly adjusted based on the video content, teaching objectives, and audience needs, without any restrictions. As you can understand, fixed position display is suitable for continuous information presentation, dynamic position adjustment can highlight key points, and intelligent layout can optimize display quality. Combined with the slide layout, it can enhance content relevance.

[0139] This embodiment obtains the slide play order corresponding to each of the slides, and then determines the animation play node corresponding to each of the mind map animation files based on the target audio file, thereby synthesizing the slides, the mind map animation files, the target subtitle file and the target audio file according to the slide play order and the animation play node to generate a target video corresponding to the slide file. The orderly combination makes the video content more organized and clear, avoids content fragmentation and jumpiness, helps the audience to better understand and remember knowledge points, and enhances the synchronization of vision and hearing. At the same time, by accurately synthesizing slides, mind map animations, subtitles and audio, the synchronization of visual elements (slides and animations) and auditory elements (audio) is ensured, so that when watching the video, the audience can simultaneously see the slide content and mind map animation corresponding to the voice explanation, enhance the learning effect, and thus improve the efficiency and quality, attractiveness and interactivity of video production.

[0140] In a feasible implementation manner, determining the animation playback node corresponding to each mind map animation file based on the target audio file includes:

[0141] Step S71: for any of the mind map animation files, obtaining a first timestamp corresponding to each entity information in the mind map animation file and an entity playback order relationship;

[0142] It should be noted that the first timestamp refers to the time point when each entity information (such as nodes, branches, etc.) appears in the mind map animation file. It marks the specific display time of each entity information in the animation, thereby determining the display order and time of each entity information in the mind map animation, and providing a basis for synchronization with the voice explanation. For example, if the "artificial intelligence" node appears at the 5th second in the mind map animation, then the first timestamp of the "artificial intelligence" node is 5 seconds.

[0143] It should be further explained that the entity playback order relationship refers to the relationship in which the entity information in the mind map animation is arranged in a logical order. It describes the hierarchical structure and logical order between the entity information, thereby ensuring that the display logic of the mind map animation is clear, and the entity information appears in the correct order, which is consistent with the logic of the voice explanation. For example, the mind map contains three nodes: "Artificial Intelligence", "Machine Learning" and "Deep Learning". The logical order between them is "Artificial Intelligence" → "Machine Learning" → "Deep Learning", then the entity playback order relationship is to display each node in this order.

[0144] Step S72, determining a second timestamp where each entity information appears in the target audio file;

[0145] It should be noted that the second timestamp refers to the time point when each entity information (such as voice paragraphs, keywords, etc.) first appears in the target audio file, that is, the specific time when each entity information first appears in the audio, thereby providing a basis for synchronization with the mind map animation. For example, if the word "artificial intelligence" appears for the first time at the 10th second in the target audio file, then the second timestamp of the word "artificial intelligence" is the 10th second.

[0146] Step S73: Based on the entity playback sequence relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

[0147] Specifically, for example, there are three nodes in the mind map: "Artificial Intelligence", "Machine Learning", and "Deep Learning". The entity playback order relationship is "Artificial Intelligence" → "Machine Learning" → "Deep Learning". The corresponding first timestamps are the 5th second, 10th second, and 15th second, and the corresponding second timestamps are the 10th second, 20th second, and 30th second. According to the entity playback order relationship, the first timestamp is associated with the second timestamp, that is, the animation playback node of the "Artificial Intelligence" node is set to start at the 10th second of the audio and last for 5 seconds (consistent with the mention time in the audio); the animation playback node of the "Machine Learning" node starts from the 20th second of the audio and lasts for 5 seconds; the "Deep Learning" node starts from the 30th second of the audio and lasts for 5 seconds, so that the display time of the mind map animation matches the time of the voice explanation, ensuring that when the audience hears a certain knowledge point, the corresponding node animation in the mind map is also displayed at the same time, thereby achieving precise synchronization of vision and hearing, and improving teaching effects and user experience.

[0148] This embodiment obtains, for any of the mind map animation files, the first timestamp corresponding to each entity information in the mind map animation file and the entity playback order relationship, and then determines the second timestamp at which each entity information appears in the target audio file. Based on the entity playback order relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file. Furthermore, by associating the entity information in the mind map animation (such as node expansion, highlighting, etc.) with the voice explanation timestamp in the target audio file, the precise synchronization of the visual element (animation) and the auditory element (voice) is ensured. When the audience hears the voice explanation of a certain knowledge point, the corresponding entity information in the mind map will be displayed at the same time, avoiding the disconnection between visual and auditory information, allowing the audience to see and hear related content at the same time, thereby improving the teaching effect and knowledge transfer efficiency, and increasing the attractiveness and professionalism of the video content.

[0149] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0150] This application also provides a slide video generation device, please refer to Figure 5 , the slide video generating device includes:

[0151] The information acquisition module 51 is used to acquire a slide file and a lecturer's voice audio file, wherein the slide file includes a plurality of slides and the lecture notes corresponding to each slide;

[0152] A first generating module 52 is configured to perform entity extraction and entity association recognition on each of the slides, and generate a mind map animation file corresponding to each of the slides;

[0153] A second generating module 53 is configured to generate a target subtitle file and a target audio file based on the lecture note information and the lecturer's voice audio file;

[0154] The third generating module 54 is configured to perform video synthesis on the slides, the mind map animation file, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file.

[0155] The slide video generating device is also used for:

[0156] Inputting each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model;

[0157] Extracting design style information corresponding to the slide in the slide file;

[0158] Based on the entity information, the entity association relationship and the design style information, a mind map animation file corresponding to each of the slides is generated.

[0159] The slide video generating device is also used for:

[0160] Inputting the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtaining the slide value score corresponding to each of the slides output by the knowledge point scoring model;

[0161] For any of the slides, if the slide value score is greater than a preset value score, the step of extracting the design style information in the slide file is performed.

[0162] The slide video generating device is also used for:

[0163] Get target timbre information;

[0164] For any of the explanation note information, input the explanation note information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model;

[0165] The explanation note information is parsed to generate a target subtitle file.

[0166] The slide video generating device is also used for:

[0167] Parsing and splitting the explanation note information according to a preset splitting rule to obtain a plurality of subtitle short sentences, and determining the subtitle sequence relationship corresponding to each of the subtitle short sentences;

[0168] Inputting each of the subtitle sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files;

[0169] Determining the subtitle playback duration corresponding to each subtitle sentence based on the total playback duration and the audio playback duration corresponding to each subtitle audio file;

[0170] The subtitle short sentences, the subtitle sequence relationship corresponding to the subtitle short sentences, and the subtitle playback time are synthesized to generate a target subtitle file.

[0171] The slide video generating device is also used for:

[0172] Obtaining the slide show order corresponding to each of the slides;

[0173] Determining the animation playback node corresponding to each of the mind map animation files based on the target audio file;

[0174] According to the slide play sequence and the animation play node, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized to generate a target video corresponding to the slide file.

[0175] The slide video generating device is also used for:

[0176] For any of the mind map animation files, obtaining a first timestamp corresponding to each entity information in the mind map animation file and a relationship between entity playback orders;

[0177] Determine a second timestamp at which each entity information appears in the target audio file;

[0178] Based on the entity playback sequence relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

[0179] The slideshow video generation device provided in this application utilizes the slideshow video generation method of the aforementioned embodiment to resolve the technical problems described in the background art. Compared to the prior art, the slideshow video generation device provided in this application achieves the same beneficial effects as the slideshow video generation method of the aforementioned embodiment. Other technical features of the slideshow video generation device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.

[0180] The present application provides a slide video generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the slide video generation method in the above-mentioned embodiment 1.

[0181] Reference below Figure 6 , which shows a schematic diagram of the structure of a slideshow video generation device suitable for implementing the embodiments of the present application. The slideshow video generation device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The slideshow video generating device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0182] like Figure 6As shown, the slideshow video generation device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the slideshow video generation device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. The communication device 1009 can allow the slideshow video generation device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a slideshow video generation device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or have alternatively.

[0183] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0184] The slideshow video generation device provided in this application utilizes the slideshow video generation method of the aforementioned embodiment to resolve the technical problems described in the background art. Compared to the prior art, the slideshow video generation device provided in this application achieves the same beneficial effects as the slideshow video generation method of the aforementioned embodiment. Other technical features of the slideshow video generation device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.

[0185] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0186] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0187] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the slide video generation method in the above embodiment.

[0188] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0189] The computer-readable storage medium may be included in the slideshow video generating device; or may exist independently without being assembled into the slideshow video generating device.

[0190] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the slideshow video generating device, the slideshow video generating device:

[0191] Obtaining a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides;

[0192] Perform entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides;

[0193] Based on the explanation note information, generate a target subtitle file and a target audio file;

[0194] Video synthesis is performed based on each of the slides, each of the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file.

[0195] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0196] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0197] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0198] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned slideshow video generation method, thereby resolving the technical problems described in the background art. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the slideshow video generation method provided in the aforementioned embodiments, and are not further elaborated here.

[0199] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned slide video generation method.

[0200] The computer program product provided in this application can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of this application are the same as the beneficial effects of the slide video generation method provided in the above embodiment, which will not be repeated here.

[0201] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for generating a slide video, characterized in that: include: Obtaining a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides; Perform entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides; Based on the explanation note information, generate a target subtitle file and a target audio file; Performing video synthesis based on the slides, the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file; The performing video synthesis based on the slides, the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file includes: Obtaining a slide play order corresponding to each of the slides; determining an animation play node corresponding to each of the mind map animation files based on the target audio file; synthesizing each of the slides, each of the mind map animation files, the target subtitle file, and the target audio file according to the slide play order and the animation play node to generate a target video corresponding to the slide file; wherein, according to the slide play order, the display order of each page of the slides in the target video is determined, and according to the animation play node, the specific display time of the mind map animation in each page of the slides is determined, and the subtitle content in the target subtitle file will be synchronously displayed according to the voice explanation timestamp in the target audio file, so as to synthesize the slides, the mind map animation files, the target subtitle file and the target audio file through a video editing tool, and during the synthesis process, ensure that the playback time and display order of each element are completely consistent with the preset playback order and playback node, wherein the elements include slides, mind map animation files, target subtitle files and target audio files; The step of determining the animation playback node corresponding to each mind map animation file based on the target audio file includes: For any of the mind map animation files, obtain the first timestamp corresponding to each entity information in the mind map animation file and the entity playback order relationship; determine the second timestamp at which each entity information appears in the target audio file; based on the entity playback order relationship, associate the first timestamp corresponding to each entity information with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

2. The slideshow video generation method according to claim 1, wherein: The performing entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides includes: Inputting each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model; Extracting design style information corresponding to the slide in the slide file; Based on the entity information, the entity association relationship and the design style information, a mind map animation file corresponding to each of the slides is generated.

3. The slideshow video generation method according to claim 2, wherein: After inputting each of the slides into the key information recognition model and obtaining the entity information and entity association relationship corresponding to each of the slides output by the key information recognition model, the method further includes: Inputting the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtaining the slide value score corresponding to each of the slides output by the knowledge point scoring model; For any of the slides, if the slide value score is greater than a preset value score, the step of extracting the design style information in the slide file is performed.

4. The method for generating a slideshow video according to claim 1, wherein: The generating of target subtitle files and target audio files based on the explanation note information includes: Get target timbre information; For any of the explanation note information, input the explanation note information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model; The explanation note information is parsed to generate a target subtitle file.

5. The method for generating a slideshow video according to claim 4, wherein: The step of parsing the explanation note information to generate a target subtitle file includes: Parsing and splitting the explanation note information according to a preset splitting rule to obtain a plurality of subtitle short sentences, and determining the subtitle sequence relationship corresponding to each of the subtitle short sentences; Inputting each of the subtitle sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files; Determining the subtitle playback duration corresponding to each subtitle sentence based on the total playback duration and the audio playback duration corresponding to each subtitle audio file; The subtitle short sentences, the subtitle sequence relationship corresponding to the subtitle short sentences, and the subtitle playback time are synthesized to generate a target subtitle file.

6. A slideshow video generating device, characterized in that: include: An information acquisition module is used to acquire a slide file and a lecturer's voice audio file, wherein the slide file includes a plurality of slides and the lecture notes corresponding to each slide; A first generating module is used to perform entity extraction and entity association recognition on each of the slides, and generate a mind map animation file corresponding to each of the slides; A second generating module is used to generate a target subtitle file and a target audio file based on each of the lecture note information and the lecturer's timbre audio file; A third generating module is used to perform video synthesis on the slides, the mind map animation file, the target subtitle file and the target audio file to generate a target video corresponding to the slide file; The third generation module is further used to obtain the slide play order corresponding to each of the slides; based on the target audio file, determine the animation play node corresponding to each of the mind map animation files; according to the slide play order and the animation play node, perform video synthesis on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file; wherein, according to the slide play order, the display order of each page of the slides in the target video is determined, and according to the animation play node, the specific display time of the mind map animation in each page of the slides is determined, and the subtitle content in the target subtitle file will be synchronously displayed according to the voice explanation timestamp in the target audio file, so as to be displayed through the video editor The tool performs video synthesis on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file, and during the synthesis process, ensures that the playback time and display order of each element are completely consistent with the preset playback order and playback node, wherein the elements include slides, mind map animation files, target subtitle files and target audio files; for any of the mind map animation files, obtain the first timestamp corresponding to each entity information in the mind map animation file and the entity playback order relationship; determine the second timestamp at which each entity information appears in the target audio file; based on the entity playback order relationship, associate the first timestamp corresponding to each entity information with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

7. A slideshow video generating device, characterized in that: The slide video generating device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the slide video generating method according to any one of claims 1 to 5.

8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the slide video generation method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method and system for automatically generating demonstration video, equipment and storage medium

    CN111538851A

  • Powerpoint construction method and device, electronic equipment and storage medium

    CN117272944A