Slide video generation method and device, equipment and storage medium

Through automated slide video generation methods, including entity extraction, subtitle and audio generation and video synthesis, the problem of manual subtitles being time-consuming and laborious in the traditional PPT to video method is solved, and efficient and accurate video production is achieved.

CN120017926AActive Publication Date: 2025-05-16深圳市深圳通有限公司

Patent Information

Application Number
CN202510495501.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The traditional PPT video conversion method requires manual subtitles and adjustment of the display time of subtitles, which is time-consuming and labor-intensive and error-prone, making it difficult to achieve efficient large-scale video production.

Method used

By obtaining the slide file, entity extraction and association recognition are performed, mind map animation files are generated; target subtitles and audio files are generated based on explanation and notes; then video synthesis is performed to generate target videos.

Benefits of technology

The subtitle addition and time adjustment process is automated, which improves the efficiency and accuracy of video production and reduces the complexity of manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017926A_ABST
    Figure CN120017926A_ABST
Patent Text Reader

Abstract

The invention discloses a slide video generation method and device, equipment and a storage medium, and relates to the technical field of video generation, and the method comprises the steps: obtaining a slide file which comprises a plurality of slides and explanation remark information corresponding to the slides; performing entity extraction and entity association identification on each slide, and generating a mind map animation file corresponding to each slide; generating a target subtitle file and a target audio file based on the explanation remark information; and performing video synthesis based on each slide, each mind map animation file, the target subtitle file and the target audio file to generate a target video corresponding to the slide file. According to the method, the problems that in the process of converting the slide file into the video, time and labor are consumed during manual subtitle adding and subtitle display adjusting, and errors are prone to occurring can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video generation technology, and in particular to a slide video generation method, device, equipment and storage medium. Background Art

[0002] In the current field of multimedia teaching and online training, PPT (slide presentation) is widely used to produce teaching courseware and training materials. With the popularity of video content, converting PPT to video has become a common demand, especially when producing online courses, training videos and distance learning content.

[0003] However, traditional PPT-to-video methods usually require manual addition of subtitles and adjustment of subtitle display time. This process is not only time-consuming and labor-intensive, but also prone to errors, making it difficult to achieve efficient large-scale video production. Summary of the invention

[0004] The main purpose of the present application is to provide a slide video generation method, device, equipment and storage medium, aiming to solve the time-consuming, labor-intensive and error-prone problems in the process of converting PPT to video, such as manually adding subtitles and adjusting the display time of subtitles.

[0005] To achieve the above purpose, the present application proposes a method for generating a slide video, the method comprising: Obtain a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides; Perform entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides; Based on each of the explanation note information, generate a target subtitle file and a target audio file; Video synthesis is performed based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

[0006] In one embodiment, the step of performing entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides includes: Input each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model; Extracting design style information corresponding to the slide in the slide file; Based on the entity information, the entity association relationship and the design style information, a mind map animation file corresponding to each of the slides is generated.

[0007] In one embodiment, after inputting each of the slides into a key information recognition model and obtaining the entity information and entity association relationship corresponding to each of the slides output by the key information recognition model, the method further includes: Inputting the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtaining the slide value score corresponding to each of the slides output by the knowledge point scoring model; For any of the slides, if the slide value score is greater than a preset value score, the step of extracting the design style information in the slide file is performed.

[0008] In one embodiment, generating a target subtitle file and a target audio file based on each of the explanation note information includes: Get target timbre information; For any of the explanation notes information, input the explanation notes information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model; The explanation note information is parsed to generate a target subtitle file.

[0009] In one embodiment, the parsing of the explanation note information to generate a target subtitle file includes: According to a preset splitting rule, the explanation note information is parsed and split to obtain a plurality of subtitle short sentences, and a subtitle sequence relationship corresponding to each of the subtitle short sentences is determined; Inputting each of the subtitle short sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle short sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files; Determine the subtitle playback duration corresponding to each of the subtitle sentences based on the total playback duration and the audio playback duration corresponding to each of the subtitle audio files; The subtitle short sentences, the subtitle sequence relationship corresponding to the subtitle short sentences, and the subtitle playback time are synthesized to generate a target subtitle file.

[0010] In one embodiment, the video synthesis is performed based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file, including: Obtaining the slide show order corresponding to each of the slides; Based on the target audio file, determining the animation playback node corresponding to each of the mind map animation files; According to the slide play sequence and the animation play node, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized to generate a target video corresponding to the slide file.

[0011] In one embodiment, determining the animation playback node corresponding to each mind map animation file based on the target audio file includes: For any of the mind map animation files, obtaining a first timestamp corresponding to each entity information in the mind map animation file and a relationship between entity playback orders; Determine a second timestamp at which each entity information appears in the target audio file; Based on the entity playback sequence relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a slide video generation device, the slide video generation device comprising: An information acquisition module is used to acquire a slide file and a lecturer's voice audio file, wherein the slide file includes a plurality of slides and explanation notes information corresponding to each of the slides; A first generating module is used to extract entities and identify entity associations for each of the slides, and generate a mind map animation file corresponding to each of the slides; A second generating module is used to generate a target subtitle file and a target audio file based on each of the lecture notes information and the lecturer's timbre audio file; The third generation module is used to perform video synthesis on the slides, the mind map animation file, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a slide video generation device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the slide video generation method described above.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the slide video generation method described above are implemented.

[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the slide video generation method described above are implemented.

[0016] The present application provides a method, apparatus, device and storage medium for generating a slide video. The slide video generating method obtains a slide file, wherein the slide file includes a plurality of slides and explanation note information corresponding to each of the slides, and then performs entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides, thereby generating a target subtitle file and a target audio file based on each of the explanation note information, and then performs video synthesis based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file, thereby solving the problem of time-consuming, labor-intensive and error-prone manual addition of subtitles and adjustment of subtitle display time in the process of converting slide files to videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A flowchart of the first embodiment of the method for generating a slide video of the present application is provided; Figure 2 A flowchart diagram of the second embodiment of the method for generating a slide video of the present application; Figure 3 A flowchart diagram of the third embodiment of the method for generating a slide video of the present application; Figure 4 A flowchart diagram of the fourth embodiment of the method for generating a slide video of the present application; Figure 5 This is a schematic diagram of the module structure of the slide video generating device according to an embodiment of the present application; Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the slide video generation method in the embodiment of the present application.

[0020] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0022] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0023] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a big data service platform, a slide video generation system, etc. The slide video generation system is taken as an example to illustrate this embodiment and the following embodiments.

[0024] Based on this, the present application embodiment provides a method for generating a slide video. Figure 1 , Figure 1 A flowchart diagram of the first embodiment of the method for generating a slide video of the present application is provided.

[0025] In this embodiment, the slide video generation method includes steps S11 to S14: Step S11, obtaining a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides; It should be noted that the slide file refers to a presentation file containing several slides, which is usually stored in PPT (PowerPoint), PPTX or other similar presentation formats. It not only contains the visual content of each slide (such as text, charts, pictures, etc.), but also may contain explanation notes related to each slide.

[0026] It should be further explained that the slide refers to a single page in a slide file, which is usually used to present a specific theme or knowledge point. Each slide may contain text, charts, pictures, graphics and other elements to convey specific information and present the content of a certain theme. The explanation notes information refers to the text content associated with each slide in the slide file, which is used to assist in explaining or illustrating the content of the slide, and is usually stored in the notes column of the slide, including detailed text descriptions, explanation points, supplementary information, etc.

[0027] Step S12, performing entity extraction and entity association recognition on each of the slides, and generating a mind map animation file corresponding to each of the slides; It should be noted that the mind map animation file refers to a dynamic display file of the mind map generated according to the content of the slide, which is used to present the knowledge points in the slide in a structured manner, including the nodes, branches and dynamic effects of the mind map (such as node expansion, contraction, highlighting, etc.), which is used to show the logical relationship between knowledge points.

[0028] Specifically, each of the slides is input into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model, and then the design style information corresponding to the slides in the slide file is extracted, thereby generating a mind map animation file corresponding to each of the slides based on the entity information, the entity association relationships and the design style information.

[0029] Step S13, generating a target subtitle file and a target audio file based on each of the explanation note information; It should be noted that the target subtitle file refers to a subtitle file generated based on the slide's explanation notes information, which is used to display the text of the explanation content in the video, including text content synchronized with the explanation audio, and usually uses a timestamp to mark the display time of each subtitle. For example, an SRT (SubRip Subtitle Format, subtitle extraction format) subtitle file contains text content synchronized with the explanation audio.

[0030] It should be further explained that the target audio file refers to an audio file generated according to the explanation notes of the slides, which is used to play the explanation content in the video, including the explanation voice, which is usually generated by text-to-speech technology, and can also be a recorded real-person explanation voice.

[0031] Specifically, the target timbre information is obtained, and then for any of the explanation notes information, the explanation notes information and the target timbre information are input into a text-to-speech model to obtain a target audio file output by the text-to-speech model, thereby parsing the explanation notes information and generating a target subtitle file.

[0032] Step S14, performing video synthesis based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

[0033] It should be noted that the target video refers to the final video file generated by combining the slide content, mind map animation, subtitles and audio, which contains the visual content of the slide, mind map animation, subtitles and explanation audio, and is a complete teaching or demonstration video. For example, the target video is a video file in MP4 format, which is used to display the slide content, and uses subtitles and audio to assist in the explanation, while displaying the logical relationship of the knowledge points in the form of a mind map animation.

[0034] Specifically, the slide playback order corresponding to each of the slides is obtained, and then based on the target audio file, the animation playback nodes corresponding to each of the mind map animation files are determined, so that according to the slide playback order and the animation playback nodes, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized into a video to generate the target video corresponding to the slide file.

[0035] This embodiment obtains a slide file, wherein the slide file includes a plurality of slides and explanation notes information corresponding to each of the slides, and then performs entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides, thereby generating a target subtitle file and a target audio file based on each of the explanation notes information, and then performs video synthesis based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file, thereby presenting complex knowledge points in the slide in a structured manner by generating a mind map animation, helping the audience to better understand and remember the content, clearly showing the logical relationship between the knowledge points (such as cause and effect, progression, parallelism, etc.), making it easier for the audience to grasp the overall framework, thereby improving the teaching effect and making the video content more vivid and interesting, and automatically generating videos to reduce the complexity of manual operations and improve the efficiency of video production, thereby solving the time-consuming, labor-intensive and error-prone problems of manually adding subtitles and adjusting the display time of subtitles in the process of converting slide files to videos.

[0036] Based on this, the present application embodiment provides a method for generating a slide video. Figure 2 , Figure 2 A flowchart diagram of the second embodiment of the slide video generation method of the present application is provided.

[0037] In a feasible implementation manner, the entity extraction and entity association recognition are performed on each of the slides to generate a mind map animation file corresponding to each of the slides, including: Step S21, inputting each of the slides into a key information recognition model, and obtaining entity information and entity association relationships corresponding to each of the slides output by the key information recognition model; It should be noted that the key information identification model refers to a model based on artificial intelligence and natural language processing technology, which is used to extract key information from text, such as entities, relationships, concepts, etc., to automatically identify important information in slides, and to provide structured data for subsequent mind map generation and video production. It is usually based on deep learning algorithms (such as BERT model, Transformer, etc.) or traditional natural language processing technology (such as named entity recognition NER, dependency syntax analysis, etc.).

[0038] It should be further explained that the entity information refers to nouns or noun phrases with specific meanings extracted from the slide text, which usually represent specific concepts, objects or themes, such as names of people, places, organization names, terms, concepts, etc. In addition, the entity association relationship refers to the logical relationship between entities, which describes the interaction or connection between entities, such as causal relationship, parallel relationship, progressive relationship, subordinate relationship, etc. For example, the relationship between "artificial intelligence" and "deep learning" can be a "containment relationship" (deep learning is a branch of artificial intelligence), which is used to construct the branch structure of the mind map and show the logical connection between knowledge points.

[0039] Specifically, each of the slides is input into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model, wherein the training process of the key information recognition model includes: collecting and annotating a large amount of slide text data, clarifying the entity information (such as names, terms, concepts, etc.) and entity association relationships (such as causality, subordination, etc.) therein, and then using natural language processing (NLP) technology, such as named entity recognition (NER) and dependency syntax analysis, to pre-process and extract features of the text, thereby selecting a suitable deep learning architecture (such as BERT, Transformer or its variants), inputting the annotated data into the model for training, adjusting the model parameters through an optimization algorithm (such as Adam or SGD) to minimize the error between the predicted results and the actual annotations, and obtaining the key information recognition model. At the same time, during the training process, data enhancement, transfer learning, multi-task learning and other technologies are used to improve the generalization ability and recognition accuracy of the model.

[0040] Step S22, extracting the design style information corresponding to the slide in the slide file; It should be noted that the design style information refers to the visual design features extracted from the slide file, which is used to maintain the consistency of the generated mind map animation with the original slide in visual style. The design style information includes but is not limited to: color, such as the main color, background color, text color, etc. used in the slide; font, such as the font type, font size, bold, italic and other styles used in the slide; layout, such as the layout method of the slide, such as symmetry, asymmetry, column, etc.; icons and graphics, such as icons, graphics, charts and other elements used in the slide; animation effects, such as animation effects used in the slide, such as fade in and out, zoom, etc., so as to generate a mind map animation consistent with the visual style of the original slide, enhancing the overall sense and professionalism of the video.

[0041] Specifically, in order to reduce computational complexity and improve video generation efficiency, usually only the main color tone, font type, font size, and other basic elements for generating mind maps in the slide file are extracted. Among them, image processing tools (such as OpenCV or Pillow) can be used to extract the main color tone of the slide, and then analyze the color distribution of the text, charts, and background in the slide. The main color is extracted and determined through a clustering algorithm (such as K-Means), and then a special library (such as python-pptx) is used to parse the PPT file, extract the font information of each slide, and analyze the frequency of font usage to determine the main font and the design style information.

[0042] Step S23, generating a mind map animation file corresponding to each of the slides based on the entity information, the entity association relationship and the design style information.

[0043] Specifically, the entity information extracted from the slides is used as the core nodes of the mind map to form the basic framework of the mind map. Then, the entity association relationships (such as cause and effect, subordination, parallelism, etc.) are used to determine the logical connections between these nodes. The hierarchy and interaction between knowledge points are clearly displayed through branches and lines, thereby constructing the logical structure of the mind map.

[0044] On this basis, combined with the design style information extracted from the slides (including visual elements such as color, font, layout, and icons), the mind map is given a visual style consistent with the original slides, ensuring that the generated animation file is highly consistent with the slides in appearance, enhancing the overall visual coherence and professionalism. Furthermore, dynamic effects are added to the mind map through animation technology (such as SVG animation, CSS animation, or a dedicated animation library), such as node expansion, contraction, highlighting, and animation transition effects, so that it is presented in a vivid and intuitive form, helping the audience to better understand and remember the content of the slides. In addition, in order to keep the expansion of the mind map consistent with the audio, timestamps are added to each node to facilitate subsequent adjustment of the expansion speed and maintain the synchronization of the mind map, audio, and subtitles.

[0045] This embodiment inputs each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model, and then extracts design style information corresponding to the slides in the slide file, thereby generating a mind map animation file corresponding to each of the slides based on the entity information, the entity association relationship and the design style information, thereby clearly displaying the logical relationship between knowledge points (such as cause and effect, progression, parallelism, etc.), helping the audience to better understand and remember the content, especially for complex or information-rich slide content, improving the structuring, logic and comprehensibility of the content, and enhancing the visualization of information. At the same time, the consistency of the overall visual style is maintained through a unified design style, making the video content more personalized and professional, while enhancing the audience's visual experience, thereby improving production efficiency and quality.

[0046] Based on this, the present application embodiment provides a method for generating a slide video. Figure 3 , Figure 3 This is a flow chart of Example 3 of the slide video generation method of this application.

[0047] In a feasible implementation manner, after inputting each of the slides into a key information recognition model and obtaining the entity information and entity association relationship corresponding to each of the slides output by the key information recognition model, the method further includes: Step S31, inputting the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtaining the slide value score corresponding to each of the slides output by the knowledge point scoring model; It should be noted that the knowledge point scoring model is used to evaluate the quality and value of the slide content. The slide value score refers to a value or level output by the knowledge point scoring model, which is used to indicate the quality and importance of the slide content. It can be presented in numerical form or level form, for example, it can be a value between 0 and 1, or a qualitative level such as "high", "medium", "low", etc., which is not limited here and can be set according to actual conditions.

[0048] Specifically, the explanation notes information, entity information and entity association relationship corresponding to each of the slides are input into the knowledge point scoring model, and the knowledge point scoring model outputs the slide value score corresponding to each of the slides, wherein the training process of the knowledge point scoring model includes: collecting a large amount of labeled slide data, including the content of the slide and the corresponding value score (such as high, medium, low or specific numerical value), and then extracting key features from these slides, such as the number of entities, the complexity of entity associations, text length, keyword density, information entropy, etc., to reflect the knowledge content and logical structure of the slides, so as to select a suitable machine learning algorithm to build a model. Common algorithms include deep learning methods (such as BERT, Transformer architecture for processing text data), traditional machine learning algorithms (such as random forests, support vector machines) or hybrid models. During the training process, cross-validation and other techniques are used to evaluate the performance of the model, and the model parameters are adjusted through optimization algorithms to minimize the error between the predicted score and the actual annotation.

[0049] Furthermore, in order to improve the accuracy of the scoring, more training samples are generated by performing synonym replacement and sentence reorganization on the slide content, and pre-trained language models (such as BERT) are used to extract text features to better capture the semantic information of the slide content. The model is then trained to complete multiple related tasks (such as entity recognition, relationship recognition, and value scoring) at the same time to enhance the model's overall understanding of the slide content. External knowledge bases, such as domain expert knowledge or industry standards, are introduced to provide the model with additional contextual information to help it more accurately evaluate the value of the slide, thereby more accurately identifying the key information and logical structure in the slide, allowing the model to output more accurate value scores, providing a reliable basis for subsequent content screening and processing.

[0050] Step S32: for any of the slides, if the slide value score is greater than a preset value score, then the step of extracting the design style information in the slide file is executed.

[0051] It should be noted that the preset value score refers to a pre-set threshold value used to determine whether to further process the slides, wherein the preset value score can be adjusted according to actual needs and application scenarios. For example, for high-quality online courses, a higher threshold value (such as 0.8) can be set; while for ordinary presentations, a lower threshold value (such as 0.5) can be set, and it can also be dynamically adjusted according to the actual use effect to optimize the screening effect.

[0052] Specifically, for any of the slides, if the slide value score is higher than or equal to the preset value score, the slide is considered to have sufficient value, and then the step of extracting the design style information in the slide file is performed to generate a mind map animation, etc. On the contrary, if it is lower than the preset threshold, the slide is ignored or skipped. For example, if the preset value score is set to 0.6, if the value score of slide A is 0.9, subsequent processing is performed; if the value score of slide B is 0.3, subsequent processing is skipped and the scoring process of the next slide is entered.

[0053] In this embodiment, the explanation notes information, entity information and entity association relationship corresponding to each of the slides are input into the knowledge point scoring model, and the knowledge point scoring model outputs the slide value score corresponding to each of the slides. Then, for any of the slides, if the slide value score is greater than the preset value score, the step of extracting the design style information in the slide file is executed, so that the slide value is evaluated through the knowledge point scoring model, and high-value content is quickly screened out to avoid unnecessary processing of low-value slides, saving time and computing resources, and improving the efficiency of content screening. Only when the value score of the slide is higher than the preset threshold, the design style information will be extracted and the mind map animation will be generated, thereby ensuring the high quality and high value of the generated content.

[0054] Based on this, the present application embodiment provides a method for generating a slide video. Figure 4 , Figure 4 A flowchart diagram of the fourth embodiment of the slide video generation method of the present application is provided.

[0055] In a feasible implementation manner, generating a target subtitle file and a target audio file based on each of the explanation note information includes: Step S41, obtaining target timbre information; It should be noted that the target timbre information is used to define data on specific timbre features in the speech synthesis process, including the pitch, timbre, speaking speed, intonation and other features of the speech, so that the generated speech sounds closer to a specific speaker and the speech generated by the text-to-speech model is adjusted to a specific timbre, such as imitating the voice of a lecturer or a specific voice style.

[0056] Specifically, in one embodiment, a clear, high-quality audio of about 20 minutes without accompaniment and noise recorded by a lecturer is obtained, an audio file is generated, and acoustic features are extracted based on the audio file to replicate the lecturer's timbre to obtain target timbre information. In addition, users can also adjust the target timbre information through audio software to meet the user's personalized customization needs.

[0057] Step S42, for any of the explanation notes information, input the explanation notes information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model; It should be noted that the text-to-speech (TTS) model refers to a model based on artificial intelligence and machine learning technology, which can convert text content into voice output and generate natural and fluent speech by analyzing the semantic and grammatical structure of the text.

[0058] Specifically, for any of the explanation notes information, the explanation notes information and the target timbre information are input into a text-to-speech model to obtain a target audio file output by the text-to-speech model.

[0059] In one embodiment, the characteristic voice of the training instructor is used to convert the text of the explanation notes of each slide into audio through the text-to-speech function of the KAN-TTS framework in the Sambert-Hifigan model, thereby obtaining a set of target audio files in order and the playback time T[i] of each target audio file.

[0060] Step S43, parsing the explanation note information to generate a target subtitle file.

[0061] Specifically, according to preset splitting rules, the explanation notes information is parsed and split to obtain a number of subtitle sentences, and the subtitle sequence relationship corresponding to each of the subtitle sentences is determined, and then each of the subtitle sentences is input into the text-to-speech model to obtain the subtitle audio files corresponding to each of the subtitle sentences output by the text-to-speech model, and the total playback time of each of the subtitle audio files is determined, so as to determine the subtitle playback time corresponding to each of the subtitle sentences based on the total playback time and the audio playback time corresponding to each of the subtitle audio files, and then perform subtitle synthesis on each of the subtitle sentences and the subtitle sequence relationship corresponding to each of the subtitle sentences and the subtitle playback time to generate a target subtitle file.

[0062] This embodiment obtains the target timbre information, and then for any of the explanation notes information, inputs the explanation notes information and the target timbre information into the text-to-speech model to obtain the target audio file output by the text-to-speech model, thereby parsing the explanation notes information and generating a target subtitle file, thereby enhancing the realism and professionalism of the video content through personalized voice, making it easier for the audience to accept and understand the explanation content, and maintaining the consistency of the subtitles and voice, because the subtitle file is generated based on the explanation notes information and fully matches the voice content, ensuring that the audience can accurately understand the content of the voice explanation through the subtitles when watching the video, improving the personalization and consistency of the voice and subtitles, and enhancing the comprehensibility and accessibility of the video.

[0063] In a feasible implementation manner, the parsing of the explanation note information to generate a target subtitle file includes: Step S51, parsing and splitting the explanation note information according to a preset splitting rule to obtain a plurality of subtitle short sentences, and determining the subtitle sequence relationship corresponding to each of the subtitle short sentences; It should be noted that the preset splitting rules refer to predefined logic and standards for splitting the explanation notes information into several concise and clear subtitle sentences. The subtitle sentences refer to short sentences split from the explanation notes information by the preset splitting rules, which are used to be displayed as subtitles in the video. Each subtitle sentence usually contains a complete meaning unit and has a moderate length (generally not more than 40 characters) to facilitate the audience to read quickly, so as to ensure that the subtitles are clear and coherent during playback.

[0064] It should be further explained that the subtitle sequence relationship refers to the display order of subtitle sentences in the video playback, which reflects the logical relationship between the subtitle sentences, that is, it is consistent with the logical order of the explanation content, ensuring that the audience can understand the content in the correct order when watching the video and the subtitles are logically clear and coherent when playing, avoiding confusion among the audience due to confusion in the subtitle order.

[0065] Specifically, in one embodiment, according to a preset splitting rule, the explanation notes information of the slide on page X in the training courseware PPT is parsed, and the explanation notes information of the slide on page X is split into several shortest sentences according to pause symbols (for example, period, comma, semicolon, question mark, line break, etc.), and then the shortest sentences are pieced together into subtitle short sentences. The preset splitting rule is as follows: a) If the length of the shortest sentence exceeds 40 characters, it will be directly upgraded to a subtitle short sentence; b) If the current shortest sentence does not contain more than 40 characters, try to combine it with the next shortest sentence to form a concatenated short sentence; c) If the length of the spelling exceeds 40 characters, the current spelling is upgraded to a subtitle sentence; d) If the current spelling does not exceed 40 characters, try to combine it with the next shortest sentence to see if the new spelling exceeds 40 characters. If it exceeds 40 characters, the current spelling will be upgraded to a subtitle sentence; if it does not exceed 40 characters, it will be combined with the next shortest sentence to form a new spelling, until the latest spelling exceeds 40 characters and is upgraded to a subtitle sentence; e) Record the correspondence and order between the subtitles and PPT pages obtained in the previous step.

[0066] f) Based on the subtitle sentence list obtained in the previous step, call the TextClip tool to generate a subtitle image with a transparent background, and record the correspondence between the subtitle image and the subtitle sentence, that is, the subtitle order.

[0067] In addition, the commentary information can be translated into other languages, and then the subtitle audio and subtitle files in the corresponding languages ​​can be generated through the text-to-speech model to support the needs of multilingual audiences.

[0068] Step S52, inputting each of the subtitle short sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle short sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files; It should be noted that the subtitle audio file refers to an audio file generated by converting subtitle sentences through a text-to-speech (TTS) model. Each subtitle sentence corresponds to a subtitle audio file, which contains the voice data of the sentence, and is used to ensure that the display time of the subtitles accurately matches the voice playback time, thereby improving the video viewing experience.

[0069] It should be further explained that the total playback time refers to the total time required to complete the playback of all subtitle audio files, which is obtained by adding the playback time of each subtitle audio file.

[0070] Specifically, continuing with the above example, using the generated characteristic timbre of the training instructor (target timbre information), through the text-to-speech function of the KAN-TTS framework in the Sambert-Hifigan model, the sequential subtitle short sentences generated by splitting each page of the courseware are converted into audio, and a set of sequential subtitle audio files corresponding to the page of slides, the playback time t[j] of each subtitle audio file, and the total playback time t of all subtitle audio files on the page are obtained.

[0071] Step S53, determining the subtitle playback duration corresponding to each of the subtitle sentences based on the total playback duration and the audio playback duration corresponding to each of the subtitle audio files; It should be noted that the subtitle playback duration refers to the length of time each subtitle phrase is displayed in the video.

[0072] Specifically, according to a preset duration calculation formula, based on the total playback duration and the audio playback duration corresponding to each of the subtitle audio files, the subtitle playback duration corresponding to each of the subtitle sentences is determined. The preset duration calculation formula is as follows: A=T[i] *(t[j] / t) Among them, A is the subtitle playback time of each subtitle image, T[i] is the playback time of the target audio file corresponding to the current slide explanation note information, t[j] is the audio playback time of the current subtitle audio file, and t is the total playback time corresponding to each of the subtitle audio files, so that the calculated subtitle playback time can accurately match the voice playback.

[0073] Step S54, synthesizing the subtitles of the subtitle short sentences, the subtitle sequence relationship corresponding to the subtitle short sentences, and the subtitle playback duration to generate a target subtitle file.

[0074] In this embodiment, the explanation note information is parsed and split according to a preset splitting rule to obtain a plurality of subtitle sentences, and the subtitle sequence relationship corresponding to each subtitle sentence is determined, and then each subtitle sentence is input into the text-to-speech model to obtain the subtitle audio file corresponding to each subtitle sentence output by the text-to-speech model, and the total playback time of each subtitle audio file is determined, so as to determine the subtitle playback time corresponding to each subtitle sentence based on the total playback time and the audio playback time corresponding to each subtitle audio file, and then the subtitles are synthesized for each subtitle sentence and the subtitle sequence relationship corresponding to each subtitle sentence and the subtitle playback time to generate a target subtitle file, so as to ensure that the semantics of each subtitle sentence are complete and concise, improve the accuracy and readability of the subtitles, and determine the display time of the subtitles according to the actual playback time of the audio, ensure the accurate matching of the subtitles and the voice, avoid the error caused by manually setting the subtitle duration, ensure that the subtitles are always synchronized with the voice, achieve accurate synchronization of the subtitles and the voice, and improve the efficiency and quality of subtitle production.

[0075] In a feasible implementation manner, the video synthesis is performed based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file, including: Step S61, obtaining the slide show order corresponding to each of the slides; It should be noted that the slide playback order refers to the order in which the slide pages are displayed in the video, which determines the expansion logic of the slide content when the audience watches the video. The slide playback order is usually consistent with the page order in the slide file, but can also be adjusted according to actual needs (for example, skipping certain pages or rearranging the page order) to ensure that the video content is expanded in a logical order so that the audience can follow the explanation smoothly.

[0076] Specifically, the slide file is converted into a PDF document format, and then each page of the generated PDF file is converted into an image format (such as a JPEG format) to obtain a set of sequential static images of the courseware, so as to obtain the slide play order corresponding to each of the slides.

[0077] Step S62, determining the animation play nodes corresponding to each of the mind map animation files based on the target audio file; It should be noted that the animation playback node refers to the display time point of the mind map animation in the video, which determines when each mind map animation segment starts playing and the duration of the playback. The animation playback node usually corresponds to the voice explanation timestamp in the target audio file to ensure that the rhythm of the animation and the voice explanation are fully matched.

[0078] Specifically, for any of the mind map animation files, the first timestamp corresponding to each entity information in the mind map animation file and the entity playback order relationship are obtained, and then the second timestamp at which each entity information appears in the target audio file is determined, and based on the entity playback order relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file. Step S63, according to the slide play sequence and the animation play node, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized into a video to generate a target video corresponding to the slide file.

[0079] Specifically, according to the slide play order, the display order of each slide in the video is determined. For example, if the slide file contains three slides, the play order is page 1, page 2, and page 3. Further, according to the animation play node, the specific display time of the mind map animation in each slide (including only the slides that can generate mind map files, that is, the value score reaches the preset threshold) is determined. Assuming that the explanation of the first slide starts from the 0th second and lasts for 10 seconds, the corresponding mind map animation will also start playing at the 0th second and complete the display within 10 seconds. At the same time, the subtitle content in the target subtitle file will be displayed synchronously according to the voice explanation timestamp in the target audio file. For example, if the first half of the explanation (0 to 5 seconds) corresponds to the first subtitle, and the second half (5 to 10 seconds) corresponds to the second subtitle, then the subtitles will be displayed in sequence during this period. Finally, these elements (slides, mind map animations, subtitles, and audio) are synthesized through video editing tools (such as Adobe Premiere, Final Cut Pro, or Python's moviepy library). During the synthesis process, ensure that the playback time and order of each element are completely consistent with the preset playback order and playback nodes. For example, when the first slide is played to the 5th second, the corresponding mind map animation node and subtitle content will also be updated synchronously, providing the audience with a consistent visual and auditory experience, so that the generated target video is not only coherent in content, but also has precise matching of visual and auditory elements, greatly enhancing the educational value and viewing experience of the video.

[0080] Among them, the display position of the mind map can be confirmed according to the position of the current slide. For example, the mind map can be fixedly displayed in the corner (such as the lower right corner), edge (such as the top or bottom) or half screen position (such as the left half screen) of the screen in the video to adapt to the continuous display of structured information without interfering with the main content of the slide. For example, displaying a mind map in the lower right corner allows the audience to view the logical structure of the knowledge point at any time; displaying a mind map at the top can be used to display the overall framework. Alternatively, the position of the mind map can be changed dynamically according to the content of the explanation. For example, the mind map can be enlarged and displayed in the center when explaining the key points; it can be faded out or moved to the edge through animation effects (such as sliding, zooming) when it is not the key point to avoid interfering with the main slide content. The position of the mind map can also be intelligently adjusted by analyzing the density and key areas of the slide content. For example, when there is less content on the right side, the mind map can be placed on the right side. The display position of the mind map can be flexibly set according to the video content, teaching objectives and audience needs. There is no restriction here. It can be understood that fixed position display is suitable for continuous display of information, dynamic position adjustment can highlight the key points, intelligent layout can optimize the display effect, and combined with the slide layout can enhance the relevance of the content.

[0081] This embodiment obtains the slide play order corresponding to each of the slides, and then determines the animation play nodes corresponding to each of the mind map animation files based on the target audio file, so as to perform video synthesis on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file according to the slide play order and the animation play node to generate a target video corresponding to the slide file, and then the orderly combination makes the video content more organized and clear, avoids content fragmentation and jumpiness, helps the audience to better understand and remember knowledge points, and enhances the synchronization of vision and hearing. At the same time, by accurately synthesizing slides, mind map animations, subtitles and audio, the synchronization of visual elements (slides and animations) and auditory elements (audio) is ensured, so that when watching the video, the audience can simultaneously see the slide content and mind map animation corresponding to the voice explanation, enhance the learning effect, and thus improve the efficiency and quality, attractiveness and interactivity of video production.

[0082] In a feasible implementation manner, determining the animation playback node corresponding to each mind map animation file based on the target audio file includes: Step S71, for any of the mind map animation files, obtaining a first timestamp corresponding to each entity information in the mind map animation file and a relationship between entity playback orders; It should be noted that the first timestamp refers to the time point when each entity information (such as nodes, branches, etc.) appears in the mind map animation file. It marks the specific display time of each entity information in the animation, thereby determining the display order and time of each entity information in the mind map animation, and providing a basis for synchronization with voice explanations. For example, the "artificial intelligence" node appears in the 5th second in the mind map animation, so the first timestamp of the "artificial intelligence" node is 5 seconds.

[0083] It should be further explained that the entity playback order relationship refers to the relationship in which the entity information in the mind map animation is arranged in a logical order. It describes the hierarchical structure and logical order between the entity information, thereby ensuring that the display logic of the mind map animation is clear, and the entity information appears in the correct order, consistent with the logic of the voice explanation. For example, the mind map contains three nodes: "Artificial Intelligence", "Machine Learning" and "Deep Learning". The logical order between them is "Artificial Intelligence" → "Machine Learning" → "Deep Learning", then the entity playback order relationship is to display each node in this order.

[0084] Step S72, determining a second timestamp at which each entity information appears in the target audio file; It should be noted that the second timestamp refers to the time point when each entity information (such as voice paragraphs, keywords, etc.) appears for the first time in the target audio file, that is, the specific time when each entity information appears for the first time in the audio, thereby providing a basis for synchronization with the mind map animation. For example, the word "artificial intelligence" appears for the first time in the target audio file at the 10th second, then the second timestamp of the word "artificial intelligence" is the 10th second.

[0085] Step S73: Based on the entity playback sequence relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

[0086] Specifically, for example, there are three nodes in the mind map: "Artificial Intelligence", "Machine Learning", and "Deep Learning", and the entity playback order relationship is "Artificial Intelligence" → "Machine Learning" → "Deep Learning", and the corresponding first timestamps are the 5th second, 10th second, and 15th second, and the corresponding second timestamps are the 10th second, 20th second, and 30th second. According to the entity playback order relationship, the first timestamp is associated with the second timestamp, that is, the animation playback node of the "Artificial Intelligence" node is set to start at the 10th second of the audio and last for 5 seconds (consistent with the mention time in the audio); the animation playback node of the "Machine Learning" node starts from the 20th second of the audio and lasts for 5 seconds; the "Deep Learning" node starts from the 30th second of the audio and lasts for 5 seconds, so that the display time of the mind map animation matches the time of the voice explanation, ensuring that when the audience hears a certain knowledge point, the corresponding node animation in the mind map is also displayed at the same time, thereby achieving precise synchronization of vision and hearing, and improving teaching effectiveness and user experience.

[0087] This embodiment obtains the first timestamp corresponding to each entity information in the mind map animation file and the entity playback order relationship for any of the mind map animation files, and then determines the second timestamp at which each entity information appears in the target audio file, and then associates the first timestamp corresponding to each entity information with the second timestamp based on the entity playback order relationship to determine the animation playback node corresponding to the mind map animation file, and then associates the entity information in the mind map animation (such as node expansion, highlighting, etc.) with the voice explanation timestamp in the target audio file to ensure the precise synchronization of visual elements (animation) and auditory elements (voice), so that when the audience hears the voice explanation of a certain knowledge point, the corresponding entity information in the mind map will be displayed at the same time, avoiding the disconnection between visual and auditory information, allowing the audience to see and hear related content at the same time, thereby improving teaching effectiveness and knowledge transfer efficiency, and improving the attractiveness and professionalism of video content.

[0088] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0089] This application also provides a slide video generation device, please refer to Figure 5 , the slide video generating device comprises: The information acquisition module 51 is used to acquire a slide file and a lecturer's voice audio file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides; A first generating module 52 is used to extract entities and identify entity associations for each of the slides, and generate a mind map animation file corresponding to each of the slides; A second generating module 53 is used to generate a target subtitle file and a target audio file based on each of the lecture notes information and the lecturer's timbre audio file; The third generating module 54 is used to perform video synthesis on the slides, the mind map animation file, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

[0090] The slide video generating device is also used for: Input each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model; Extracting design style information corresponding to the slide in the slide file; Based on the entity information, the entity association relationship and the design style information, a mind map animation file corresponding to each of the slides is generated.

[0091] The slide video generating device is also used for: Inputting the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtaining the slide value score corresponding to each of the slides output by the knowledge point scoring model; For any of the slides, if the slide value score is greater than a preset value score, the step of extracting the design style information in the slide file is performed.

[0092] The slide video generating device is also used for: Get target timbre information; For any of the explanation notes information, input the explanation notes information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model; The explanation note information is parsed to generate a target subtitle file.

[0093] The slide video generating device is also used for: According to a preset splitting rule, the explanation note information is parsed and split to obtain a plurality of subtitle short sentences, and a subtitle sequence relationship corresponding to each of the subtitle short sentences is determined; Inputting each of the subtitle short sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle short sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files; Determine the subtitle playback duration corresponding to each of the subtitle sentences based on the total playback duration and the audio playback duration corresponding to each of the subtitle audio files; The subtitle short sentences, the subtitle sequence relationship corresponding to the subtitle short sentences, and the subtitle playback time are synthesized to generate a target subtitle file.

[0094] The slide video generating device is also used for: Obtaining the slide show order corresponding to each of the slides; Based on the target audio file, determining the animation playback node corresponding to each of the mind map animation files; According to the slide play sequence and the animation play node, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized to generate a target video corresponding to the slide file.

[0095] The slide video generating device is also used for: For any of the mind map animation files, obtaining a first timestamp corresponding to each entity information in the mind map animation file and a relationship between entity playback orders; Determine a second timestamp at which each entity information appears in the target audio file; Based on the entity playback sequence relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

[0096] The slide video generation device provided by the present application adopts the slide video generation method in the above embodiment, which can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the slide video generation device provided by the present application are the same as the beneficial effects of the slide video generation method provided by the above embodiment, and other technical features in the slide video generation device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0097] The present application provides a slide video generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the slide video generation method in the above-mentioned embodiment 1.

[0098] Reference below Figure 6 , which shows a schematic diagram of the structure of a slide video generating device suitable for implementing the embodiment of the present application. The slide video generating device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The slide video generating device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0099] like Figure 6 As shown, the slide video generating device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 to the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the slide video generating device are also stored. The processing device 1001, the read-only memory 1002 and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the slide video generation device to communicate with other devices wirelessly or by wire to exchange data. Although the slide video generation device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0100] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0101] The slide video generation device provided by the present application adopts the slide video generation method in the above embodiment, which can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the slide video generation device provided by the present application are the same as the beneficial effects of the slide video generation method provided by the above embodiment, and other technical features in the slide video generation device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0102] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0103] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0104] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the slide video generation method in the above-mentioned embodiment.

[0105] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.

[0106] The computer-readable storage medium may be included in the slide video generating device; or may exist independently without being assembled into the slide video generating device.

[0107] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the slide video generating device, the slide video generating device: Obtain a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides; Perform entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides; Based on the explanation note information, generate a target subtitle file and a target audio file; Video synthesis is performed based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

[0108] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0109] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0110] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0111] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned slide video generation method, and can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the slide video generation method provided in the above-mentioned embodiment, and will not be repeated here.

[0112] An embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the above-mentioned slide video generation method when executed by a processor.

[0113] The computer program product provided in this application can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of this application are the same as the beneficial effects of the slide video generation method provided in the above embodiment, which will not be repeated here.

[0114] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for generating a slide video, characterized in that: include: Obtain a slide file, wherein the slide file includes a plurality of slides and explanation notes corresponding to each of the slides; Perform entity extraction and entity association recognition on each of the slides to generate a mind map animation file corresponding to each of the slides; Based on each of the explanation note information, generate a target subtitle file and a target audio file; Video synthesis is performed based on each of the slides, each of the mind map animation files, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

2. The method for generating a slide video according to claim 1, wherein: The extracting entities and identifying entity associations of the slides to generate a mind map animation file corresponding to the slides includes: Input each of the slides into a key information recognition model to obtain entity information and entity association relationships corresponding to each of the slides output by the key information recognition model; Extracting design style information corresponding to the slide in the slide file; Based on the entity information, the entity association relationship and the design style information, a mind map animation file corresponding to each of the slides is generated.

3. The method for generating a slide video according to claim 2, wherein: After inputting each of the slides into the key information recognition model to obtain the entity information and entity association relationship corresponding to each of the slides output by the key information recognition model, the method further includes: Inputting the explanation notes information, entity information and entity association relationship corresponding to each of the slides into the knowledge point scoring model, and obtaining the slide value score corresponding to each of the slides output by the knowledge point scoring model; For any of the slides, if the slide value score is greater than a preset value score, the step of extracting the design style information in the slide file is performed.

4. The method for generating a slide video according to claim 1, wherein: The generating of a target subtitle file and a target audio file based on each of the explanation note information includes: Get target timbre information; For any of the explanation notes information, input the explanation notes information and the target timbre information into a text-to-speech model to obtain a target audio file output by the text-to-speech model; The explanation note information is parsed to generate a target subtitle file.

5. The method for generating a slide video according to claim 4, wherein: The step of parsing the explanation note information to generate a target subtitle file includes: According to a preset splitting rule, the explanation note information is parsed and split to obtain a plurality of subtitle short sentences, and a subtitle sequence relationship corresponding to each of the subtitle short sentences is determined; Inputting each of the subtitle short sentences into the text-to-speech model, obtaining a subtitle audio file corresponding to each of the subtitle short sentences output by the text-to-speech model, and determining a total playback time of each of the subtitle audio files; Determine the subtitle playback duration corresponding to each of the subtitle sentences based on the total playback duration and the audio playback duration corresponding to each of the subtitle audio files; The subtitle short sentences, the subtitle sequence relationship corresponding to the subtitle short sentences, and the subtitle playback time are synthesized to generate a target subtitle file.

6. The method for generating a slide video according to claim 1, wherein: The step of performing video synthesis based on each of the slides, each of the mind map animation files, the target subtitle file, and the target audio file to generate a target video corresponding to the slide file includes: Obtaining the slide show order corresponding to each of the slides; Based on the target audio file, determining the animation playback node corresponding to each of the mind map animation files; According to the slide play sequence and the animation play node, each of the slides, each of the mind map animation files, the target subtitle file and the target audio file are synthesized to generate a target video corresponding to the slide file.

7. The method for generating a slide video according to claim 6, wherein: The step of determining the animation playback nodes corresponding to the mind map animation files based on the target audio file includes: For any of the mind map animation files, obtaining a first timestamp corresponding to each entity information in the mind map animation file and a relationship between entity playback orders; Determine a second timestamp at which each entity information appears in the target audio file; Based on the entity playback sequence relationship, the first timestamp corresponding to each entity information is associated with the second timestamp to determine the animation playback node corresponding to the mind map animation file.

8. A slide video generating device, characterized in that: include: An information acquisition module is used to acquire a slide file and a lecturer's voice audio file, wherein the slide file includes a plurality of slides and explanation notes information corresponding to each of the slides; A first generating module is used to extract entities and identify entity associations for each of the slides, and generate a mind map animation file corresponding to each of the slides; A second generating module is used to generate a target subtitle file and a target audio file based on each of the lecture notes information and the lecturer's timbre audio file; The third generation module is used to perform video synthesis on the slides, the mind map animation file, the target subtitle file and the target audio file to generate a target video corresponding to the slide file.

9. A slide video generating device, characterized in that: The slide video generating device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the slide video generating method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the slide video generation method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and system for automatically generating demonstration video, equipment and storage medium

    CN111538851A

  • Generation association method of online course knowledge tree

    CN112231522A

  • Presentation file display method and device, storage medium and electronic device

    CN113220200A

  • Powerpoint construction method and device, electronic equipment and storage medium

    CN117272944A

  • Structured content processing method and system oriented to narrative media data

    CN117312588A

Cited By

  • Transforming content across visual mediums using artificial intelligence and user generated media

    US12579710B2

  • Transforming Content Across Visual Mediums Using Artificial Intelligence and User Generated Media

    US20240257420A1