Slide explanation video generation method and device, computer equipment and storage medium
Through the in-depth processing and synthesis of slide files by multimodal large models, the problems of low efficiency and crude effects of slide-to-video tools are solved, and the generation of slide explanation videos with clear logic and personalized features is achieved, meeting the professional needs of the financial technology and medical health care fields.
Patent Information
- Application Number
- CN202510839867.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-19
AI Technical Summary
Existing slide-to-video tools have low production efficiency in the fields of financial technology and medical health care, lacking deep understanding and dynamic optimization capabilities. As a result, the generated videos are stiff, the rhythm of voice and picture does not match, and they cannot meet the complex needs of professional fields.
A multimodal large model is used to deeply process the slide files to generate explanation script data, and the slide explanation video is synthesized through the visual material library and text-to-speech model to achieve precise alignment and dynamic adaptation of visual elements and voice content.
The generated slideshow explanation videos are more in line with the knowledge dissemination needs in professional fields, with accurate logic, detailed and concise content, which improves the video quality and user experience.
Smart Images

Figure CN120672919A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multimodal large models, and in particular to a method, apparatus, computer equipment, and storage medium for generating a slideshow video. Background Art
[0002] In the fields of fintech and healthcare, slideshow videos are crucial tools for knowledge dissemination and business promotion. In fintech, complex charts and data require precise interpretation, and to aid understanding, diverse visuals and a uniquely designed presentation hierarchy are essential. In healthcare, specialized terminology and flowcharts must be clearly presented, and specific medical images require specialized annotation and interpretation.
[0003] However, most software tools simply perform static conversions based on fixed templates and rules, lacking dynamic optimization capabilities. Some simply stitch together PowerPoint pages in sequence or simply convert text to speech, requiring manual addition of animations, dubbing, and transitions. This is inefficient and lacks in-depth understanding and analysis of the content, let alone designing and improving the hierarchical structure and audiovisual effects of the presentation, thus failing to meet the complex demands of professional fields. In particular, the videos generated by these tools lack sync between the audio and the image, and the text summary lacks a logical connection with the visual elements, resulting in a stilted video presentation.
[0004] Therefore, in the fields of financial technology and medical care, there is an urgent need for a PPT-to-video tool that can deeply understand PPT content, dynamically adapt to multimodal fusion, and support personalization and interactivity to improve video quality and user experience. Summary of the Invention
[0005] The present application discloses a method, device, computer equipment and storage medium for generating a slideshow explanation video, which solves the problems of low efficiency, blunt effect and inconvenient operation in the related art of generating a slideshow explanation video.
[0006] In a first aspect, the present application provides a method for generating a slideshow video, comprising:
[0007] Acquire a slide file, process the slide file, and obtain processed file data;
[0008] Inputting the file data into the first multimodal large model to obtain explanation script data, wherein the explanation script data includes visual element information and key text information;
[0009] Inputting the explanation script data into the second multimodal large model, aligning and generating fused instruction information, wherein the fused instruction information includes visual content instruction information and voice content instruction information;
[0010] According to the visual content instruction information, calling the visual material library to obtain animation data;
[0011] Inputting the voice content instruction information into a text-to-speech model to obtain dubbing data;
[0012] A slideshow explanation video is synthesized based on the animation data and the dubbing data.
[0013] In a second aspect, the present application provides a device for generating a slideshow video, comprising:
[0014] A file acquisition module, used to acquire a slide file, process the slide file, and obtain processed file data;
[0015] A script generation module, configured to input the file data into the first multimodal large model to obtain explanation script data;
[0016] An instruction generation module, configured to input the explanation script data into a second multimodal large model, align and generate fused instruction information, wherein the fused instruction information includes visual content instruction information and voice content instruction information;
[0017] A visual effect generation module, configured to call a visual material library according to the visual content instruction information to obtain animation data;
[0018] A sound effect generation module, configured to input the voice content instruction information into a text-to-speech model to obtain dubbing data;
[0019] The video synthesis module is used to synthesize the slideshow explanation video according to the animation data and the dubbing data.
[0020] In a third aspect, the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the method provided in any embodiment of the present application.
[0021] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor implements a method as provided in any embodiment of the present application.
[0022] The slideshow explanation video generation method proposed in this application deeply processes slideshow files and uses a multimodal large model to generate explanation script data, thereby obtaining an accurate, reliable, in-depth, logically accurate, and detailed explanation script that is more in line with the real needs of knowledge dissemination and business promotion in the fields of financial technology and medical health care. On this basis, alignment and generation of fusion instructions based on visual element information and key text information can solve problems such as audio and video synchronization, and make subsequent processing such as adding visual animation and dubbing more accurate and reliable, thereby improving the overall effect of the generated slideshow explanation video.
[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 This is a flowchart showing the steps of a method for generating a slideshow video according to an embodiment of the present application;
[0026] Figure 2 This is a flowchart showing the steps of a method for generating a slideshow video according to an embodiment of the present application;
[0027] Figure 3 This is a flowchart showing the steps of the model training method provided in one embodiment of the present application;
[0028] Figure 4 This is a structural diagram of a device for generating a slideshow video provided in an embodiment of the present application;
[0029] Figure 5 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application.
[0030] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0032] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0033] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0034] It should be understood that, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the first data and the second data are merely used to distinguish different data and do not limit their order. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences.
[0035] It should be further understood that the term “and / or” used in this specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0036] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.
[0037] With the acceleration of digital transformation, the demand for converting PPT files to videos is growing in the financial technology, healthcare, and elderly care sectors. However, most existing conversion tools rely on static conversions based on fixed templates and rules. This technical implementation has exposed numerous flaws in practical applications, seriously affecting the efficiency and quality of video generation.
[0038] In the field of financial technology, PPTs usually contain complex data charts, professional terms and detailed business processes. Existing tools can only splice PPT pages into videos in sequence, lack the ability to deeply analyze the content, and cannot automatically identify and understand the key information on the page. This requires users to manually add animations, dubbing and transition effects, which is not only inefficient but also prone to errors. In addition, existing technologies fail to effectively combine natural language processing (NLP), speech synthesis (TTS) and visual generation technologies. The rhythm of the voice and picture in the generated video is not synchronized, and there is a lack of logical connection between the text summary and the visual elements, which makes the video presentation effect stiff and difficult to meet the high requirements of professionalism and accuracy in the financial industry. At the same time, existing solutions cannot adaptively adjust the video style according to user needs (such as audience type, application scenario), only support fixed conversion rules, lack dynamic optimization capabilities, and are difficult to adapt to the diverse business needs of the financial industry.
[0039] In the healthcare and elderly care sectors, presentations often contain a rich and diverse range of information, including detailed patient care processes, precise and rigorous medical terminology, and vast amounts of clinical data and imaging. This content requires not only a high degree of professionalism to ensure accurate communication, but also clear logic and comprehensibility in the video presentation, so that different audiences (such as patients, medical staff, and family members) can easily access key information. Regarding multimodal fusion, existing technologies fail to effectively integrate natural language processing (NLP), text-to-speech (TTS), and visual generation technologies. This results in a mismatch between the rhythm of speech and image in the generated videos, a lack of logical connection between text summaries and visual elements, and a stilted video presentation that fails to meet the high standards of professionalism and accuracy required by the healthcare industry. Furthermore, existing solutions are unable to adaptively adjust the video style based on user needs (such as audience type and application scenario). They only support fixed conversion rules and lack dynamic optimization capabilities. This forces medical staff and pharmaceutical and medical device engineers to manually add animations, dubbing, and transition effects when producing videos, which not only consumes a lot of time and effort but is also prone to errors due to improper operation.
[0040] Therefore, existing PPT-to-video tools lack significant capabilities in content understanding and dynamic adaptation, multimodal fusion, and personalization and interactivity. These issues not only reduce the efficiency of video generation but also impact video quality and user experience. Therefore, the industry urgently needs a PPT-to-video technology that can deeply analyze and understand PPT content, dynamically adapt and integrate multimodal data, and support personalization and interactivity to meet the professional needs of the financial technology and healthcare and elderly care sectors.
[0041] To solve the above problems, the present application proposes a method for generating a slideshow video. Figure 1 , Figure 1This is a schematic flow chart of the steps of the method for generating a slideshow video according to an embodiment of the present application. Figure 1 As shown, the slideshow explanation video generation method specifically includes steps S101 to S105.
[0042] S101: Acquire a slide file, process the slide file, and obtain processed file data.
[0043] It should be understood that in the field of FinTech, slide files may include financial product introductions, market analysis reports, investment strategies, etc. For example, slide files can be obtained from information sources such as the internal knowledge base of a financial institution, market research reports, or client education materials.
[0044] In some embodiments, the slide file can be subjected to various processes such as format conversion, information extraction, data analysis, and semantic understanding to obtain processed file data. For example, the slide file can be converted from a common PPT format to a format suitable for further processing, such as HTML or JSON. This facilitates subsequent data extraction and analysis while ensuring data compatibility and scalability.
[0045] In some embodiments, information extraction can be performed on slide files to obtain file data. Specifically, natural language processing (NLP) technology can be used to extract key information from the slides, such as the name of the financial product, risk level, expected return, and investment period. For market analysis reports, information such as market trends, industry dynamics, and key indicators can be extracted.
[0046] In some embodiments, data analysis can be performed on slide files to obtain file data. Specifically, the extracted data can be subjected to in-depth analysis. For example, the return volatility and risk assessment indicators of different financial products can be calculated, or market trend data can be analyzed to identify potential investment opportunities. Furthermore, statistical analysis of the historical performance of financial products can be performed to provide investors with a more comprehensive reference.
[0047] In some embodiments, semantic understanding can be performed on slide files to obtain file data. Specifically, semantic analysis techniques can be used to understand the logical structure of the slide content. For example, key semantic information about product advantages, market trends, and the core points of investment strategies can be identified. This helps generate more accurate and easy-to-understand explanation video content.
[0048] In some embodiments, data association can be performed on the slide file. Specifically, the extracted and analyzed data can be associated with the financial institution's customer database. For example, based on the customer's risk preferences and investment goals, matching financial product information can be screened and highlighted in the presentation script or file data.
[0049] It should be understood that in the field of medical care, health care, and elderly care, slide files may include health science knowledge, disease prevention methods, and introductions to elderly care service programs. Slide files can be obtained through information sources such as health education materials from medical institutions, introductions to activities in elderly care communities, or health science lectures.
[0050] In some embodiments, the slide file can be converted from the common PPT format to a format suitable for further processing, such as Markdown or XML, to facilitate subsequent data extraction and analysis while ensuring data compatibility and scalability.
[0051] In some embodiments, when processing slide files, various technical means can be employed to ensure effective extraction and utilization of information. For example, natural language processing (NLP) technology can be used to analyze the slide content and extract key information, such as disease names, preventive measures, health advice, and features of elderly care services. For health science knowledge, key health indicators and lifestyle recommendations can be further extracted to provide users with more targeted information.
[0052] In some embodiments, semantic analysis techniques can be used to better understand and organize slide content, identifying its logical structure. For example, this could identify key health science topics, key steps in disease prevention, and the core benefits of elderly care services. This approach not only helps generate more precise and accessible video content, but also ensures the accuracy and effectiveness of information conveyed.
[0053] In some embodiments, the extracted and analyzed data can be linked to a medical institution's patient database or to a retirement community's resident information. For example, based on the health status and needs of the elderly, matching health science knowledge and elderly care services can be screened and highlighted in the video explanation. This approach can provide users with personalized content, improving the relevance and practicality of information, and better meeting their needs.
[0054] In some embodiments, the file data may include text data, image data, and attribute data. Specifically, the attribute data may include content type data, page sequence data, or theme style data. Content type data may include content types such as course training, marketing proposals, investor education, and technical analysis introductions. Theme style data may include data such as the selected theme color and the selected theme template. Page sequence data may include paging data, chapter data, etc. This allows for more accurate and efficient generation of explanation script data when the file data is subsequently input into the first multimodal large model, and helps users, developers, and artificial intelligence models compare the explanation script data with the file data to achieve incremental optimization.
[0055] S102: Input the file data into the first multimodal large model to obtain explanation script data.
[0056] The explanation script data can include visual element information and key text information. Key text information can provide important prompts and contextual tools for the explanation video generation process, thereby achieving effects such as detailed and concise explanations and clear logical connections. Furthermore, the generation of key text information makes it possible to generate corresponding visual element information based on text-based timestamps, node associations, or guidance, saving the additional matching computing resources required for unrelated generation, greatly improving the quality and efficiency of subsequent content instructions and explanation video generation processes.
[0057] In some embodiments, the aforementioned text data, image data, and attribute data can be input into the first multimodal macromodel to generate explanation script data. Because detailed data such as text data, image data, and attribute data are extracted from the file data, the first multimodal macromodel can perform classification and refined processing based on the different characteristics of different data, thereby effectively improving the processing efficiency and technical effects of the first multimodal macromodel.
[0058] In some embodiments, the first multimodal macromodel may include an analysis model and a synthesis model. Inputting document data into the first multimodal macromodel may specifically include inputting the document data into the analysis model to obtain keyword data, key sentence data, and contextual relationship data; and then inputting the keyword data, key sentence data, and contextual relationship data into the synthesis model to generate explanation script data.
[0059] It should be understood that extracting keyword data, key sentence data, and contextual relationship data from file data does not conflict with parsing the slide file to obtain text data, image data, and attribute data. These two steps can be performed simultaneously or separately. Specifically, the contextual relationship data can include text-to-text association data and text-to-visual element association data. The text-to-text association data includes causal relationship data, contrast relationship data, or hierarchical relationship data. The text-to-visual element association data includes timestamp information, condition information, or timing information.
[0060] In some embodiments, see Figure 2 , Figure 2 This is a schematic flow chart of the steps of the method for generating a slideshow video according to an embodiment of the present application. Figure 2 As shown, the method for generating a slideshow explanation video may specifically include steps S102a to S102b.
[0061] S102a: Segment the text data to obtain sentence data.
[0062] For example, the text data can be processed by Chinese word segmentation, segmentation, and other sentence processing to obtain sentence data. In the financial and medical fields, from the user's perspective, the granularity of the voice and visual alignment of slide presentations is often at the Chinese sentence level. Too detailed a split will not only cause unnecessary computing power consumption, but also result in poor overall video effects of the slide presentation, misalignment of voice and visual generation, or excessive visual effects generation, giving users a "dazzled" feeling. Too rough a split often results in long-winded voice content and relatively monotonous visuals, which is also not conducive to the slide presentation effect.
[0063] S102b: Input the sentence data into the first pre-training model, perform semantic structure analysis to obtain contextual relationship data, and / or input the sentence data into the first pre-training model, extract key information to obtain keyword data and / or key sentence data.
[0064] In the fintech sector, slide decks often contain complex financial terminology, data charts, investment strategies, and market analysis. The data structure of this content is unique. For example, financial data often exhibits time series characteristics, and investment strategy descriptions may involve causal relationships and conditional logic. Therefore, for parsing slide decks in the fintech sector, a specialized multimodal model architecture, such as the first pre-trained model described above, can be constructed to analyze the semantic structure of sentence data and extract key information.
[0065] In some embodiments, the first pre-trained model may include a contextual relationship extractor that can extract text-to-text association data and text-to-visual element association data from the fused features, or from the pre-fused sentence data. Specifically, this may include text-to-text association data, such as causal relationship data, contrast relationship data, and hierarchical relationship data.
[0066] In some embodiments, the first pre-trained model may also construct a key information extractor for extracting keywords and key sentences from the text content. Specifically, keyword extraction may be based on similarity calculations using TF-IDF or BERT word embeddings, and key sentence extraction may be based on the semantic importance and contextual relevance of sentences. For example, summary paragraphs and data explanation paragraphs may be given higher weights.
[0067] It should be understood that the contextual relationship data may also include text-text association data and text-visual element association data. The text-text association data includes causal relationship data, comparative relationship data or hierarchical relationship data, and the text-visual element association data includes timestamp information, condition information or timing information.
[0068] In some embodiments, some slide data can be manually annotated first, including the contextual relationships in the text (causality, contrast, hierarchy) and the associations between text and visual elements (timestamps, conditions, timing). In the specific training process, the large language model is first pre-trained based on a large-scale medical field corpus, and then fine-tuned on the annotated data to adapt to the slide file parsing task. Furthermore, the large language model can also be used to train text feature extraction, visual feature extraction, and relationship modeling tasks, and the comprehensive performance of the multimodal model can be improved through multi-task learning.
[0069] Furthermore, we can use graph neural networks (GNNs) to model the relationship between text and visual elements. We build a graph structure where nodes represent text fragments and visual elements, and edges represent the relationship between them (such as timestamp association, conditional association, and timing information).
[0070] In some embodiments, a pre-trained large language model can be used to encode the text content in a slide and extract text semantic features, or contextual information and key concepts can be directly extracted from the text. Pre-trained convolutional neural networks can also be used to extract image semantic information from images and charts, and OCR technology can be used to obtain text semantic information within the images and charts. This image semantic information, text semantic information, and the text semantic features extracted from the text content can then be integrated using a multi-head attention mechanism.
[0071] Specifically, the first multimodal large model may include an input processing module, a feature extraction module, a modal fusion module and a result output module. The input processing module may include, the feature extraction module may include a text feature extractor and an image feature extractor, which are respectively used for feature extraction of text semantics and image semantics in the slide file data. The modal fusion module may include an attention mechanism and a relationship modeling mechanism, and the relationship modeling mechanism may include at least one graph structure. The result output module may include a contextual relationship extractor, which may be used to extract text-text association data and text-visual element association data from the fused features, wherein the text-text association data may include causal relationship data, contrast relationship data or hierarchical relationship data, and the text-visual element association data may include timestamp information, condition information, and timing information.
[0072] Furthermore, a knowledge graph can be constructed based on the output results of the first multimodal large model, so as to store and identify the internal logic and contextual relationships of the slide content in the form of structured data, which is more conducive to the R&D team to carry out manual inspection and correction and targeted model iteration in actual deployment and operation.
[0073] S103: Input the explanation script data into the second multimodal large model, align and generate fusion instruction information.
[0074] The fusion instruction information may include visual content instruction information and voice content instruction information. For example, the fusion instruction information may include emotion annotation, content annotation, style annotation, theme annotation, etc. Exemplarily, the second multimodal large model may be used to perform emotion annotation, content annotation, style annotation, or theme annotation on the explanation script data, split the explanation script data according to visual and voice classification, and align them to obtain the fusion instruction information. For example, the second multimodal large model may be used to perform emotion annotation on the explanation script data to obtain the emotion tags in the key text information.
[0075] For example, the second multimodal large model can be used to directly process the explanation script data, align it, and generate fused instruction information. It should be understood that alignment can include aligning key text information with visual element information along multiple dimensions, such as time, condition, and / or order. For example, in explanation A, the specific form of alignment can be timestamps, conditional order, etc. For example, after animation effect A ends, voice B is played. For another example, at time A, voice B is played. For another example, after voice B ends, animation effect C is activated.
[0076] S104: According to the visual content instruction information, the visual material library is called to obtain animation data.
[0077] It should be understood that the visual element information obtained by the first multimodal large model may include icon information, chart information, motion effect information, or visual element position information, and thus the second multimodal large model can obtain visual content instruction information such as icon instructions, chart instructions, motion effect instructions, or visual element position instructions based on the icon information, chart information, motion effect information, or visual element position information. Specifically, the visual content instruction information may also include a method for calling a visual material library. Users can design corresponding data structures adaptively or programmatically based on different visual material library calling methods.
[0078] In some embodiments, the visual content instruction information may include natural language instructions such as "Please generate a bar chart object based on data A, and place the bar chart object in Section E of Chapter B, Section C, Page D of the slideshow video in a fade-in / fade-out animation format." Alternatively, the visual content instruction information may be in a specific data structure. Furthermore, the visual content instruction information may also limit specific parameters such as scale, contrast, object size, position, and rotation for objects such as bar charts, pie charts, and inserted images.
[0079] S105: Input the voice content instruction information into the text-to-speech model to obtain dubbing data.
[0080] It should be understood that the voice content instruction information may also include a call key for the text-to-speech service. Specifically, the text-to-speech model may be called directly through a software interface, or the related service of the text-to-speech model may be called indirectly through a cloud service or network communication.
[0081] It should be further understood that the key text information obtained by the first multimodal large model may include emotional tags, and thus the second multimodal large model can obtain emotional instructions from the voice content instruction information based on the emotional tags. The text-to-speech model can generate dubbing data with a specific style, appropriate details, reasonable speaking speed, and pleasant sound quality based on the voice content instruction information such as emotional instructions.
[0082] S106: synthesize the slideshow explanation video based on the animation data and the dubbing data.
[0083] It should be understood that both the slideshow explanation video and the animation data can be in formats such as MP4 and AVI. In some embodiments, the animation data can include compressed animation data of multiple keyframes. Alternatively, the animation data can be synthesized using a multimodal large language model to obtain intermediate animation data, which can then be synthesized with the dubbing data to obtain the slideshow explanation video. Alternatively, the animation data and dubbing data can be directly synthesized to obtain the slideshow explanation video. The intermediate animation data can be synthesized in conjunction with the dubbing data to improve the semantic consistency between the animation data and the dubbing data of the slideshow explanation video.
[0084] Since in the aforementioned process of generating animation data, compressed animation data that only generates key frames can be used, the computing power consumed by directly generating videos based on semantics can be greatly reduced. In comparison, the computing power required to further generate transition frames between key frames based on key frames will be greatly reduced, and the number of key frames themselves is limited, which is more conducive to alignment and synthesis with dubbing data.
[0085] In some embodiments, see Figure 3 , Figure 3 This is a flowchart illustrating the steps of a model training method provided in one embodiment of the present application. The model training method can be used to implement iterative training of the aforementioned first multimodal large model, thereby improving the interactivity, performance, and user experience of the first multimodal large model. Specifically, the model training method may include steps S201 to S202.
[0086] S201. Obtain a user change instruction, and change the explanation script data according to the user change instruction to obtain script change data.
[0087] It should be understood that an interactive interface can be set up so that both R&D personnel and ordinary users can view, edit or correct at least a part of the explanation script data such as visual element information and key text information and / or fusion instruction information such as visual content instruction information and voice content instruction information generated by the model during a single use of the slideshow explanation video generation service. Thereby significantly improving the professionalism and accuracy of the generated slideshow explanation videos in the field of financial technology or the slideshow explanation videos in the field of medical health and elderly care, and users can also bypass the existing professional video post-processing work with complex operations and high thresholds, which greatly facilitates user control and use. In essence, it is also possible to transform the traditional black box, device-based slideshow explanation video generation system into a slideshow explanation video generation system based on human-computer collaboration, which is controllable and interactive, significantly improving the quality of the generated slideshow explanation videos and the user experience of the generation process.
[0088] In some embodiments, the user change instruction can be used to change explanation script data such as visual element information and key text information, or can be directly used to change fusion instruction information such as visual content instruction information and voice content instruction information.
[0089] S202: Train the first multimodal large model according to the script change data and the user change instruction.
[0090] It should be understood that because the work undertaken by the first multimodal large model is directly related to the instruction information generation task of the second multimodal large model, compared to the complexity of the second multimodal large model, continuous iteration and reinforcement learning of the first multimodal large model is more necessary and more effective. Therefore, the first multimodal large model can be trained based on scripted data changes and user-modified instructions.
[0091] Specifically, techniques such as Dropout and L2 regularization can be used to prevent overfitting. Learning rate decay strategies can also be used to dynamically adjust the learning rate to accelerate model convergence. During training data preprocessing, machine and / or manual screening can also be performed to ensure data quality and the efficiency of tasks such as annotation training during incremental training.
[0092] See also Figure 4 , Figure 4 The present invention provides a schematic diagram of a device for generating a slideshow video according to an embodiment of the present invention, wherein the device is configured to execute the aforementioned method for generating a slideshow video. The device can be configured in a terminal or a server.
[0093] like Figure 4 As shown, the slideshow explanation video generation device 100 includes a file acquisition module 101, a script generation module 102, an instruction generation module 103, a visual effect generation module 104, a sound effect generation module 105 and a video synthesis module 106.
[0094] The file acquisition module 101 is used to acquire a slide file, process the slide file, and obtain processed file data.
[0095] The script generation module 102 is used to input the file data into the first multimodal large model to obtain the explanation script data.
[0096] The instruction generation module 103 is used to input the explanation script data into the second multimodal large model, align and generate fused instruction information, and the fused instruction information includes visual content instruction information and voice content instruction information.
[0097] The visual effect generation module 104 is used to call the visual material library according to the visual content instruction information to obtain animation data.
[0098] The sound effect generation module 105 is used to input the voice content instruction information into the text-to-speech model to obtain dubbing data.
[0099] The video synthesis module 106 is used to synthesize the slideshow explanation video according to the animation data and the dubbing data.
[0100] In some embodiments, the slideshow explanation video generating device 100 may further include a user interaction module, which may be used to obtain a user change instruction, and change the explanation script data according to the user change instruction to obtain script change data.
[0101] In some embodiments, the slideshow explanation video generating device 100 may further include an iterative optimization module, which may be used to train the first multimodal large model according to script change data and user change instructions.
[0102] It should be noted that those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0103] The above-mentioned device can be realized in the form of a computer program. The computer program can be used in Figure 5 Runs on the computer equipment shown.
[0104] See also Figure 5 , Figure 5 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a server. Figure 5 The computer device includes a processor, a memory, and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0105] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause the processor to execute any of the slideshow video generation methods of the present application.
[0106] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0107] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the slideshow video generation methods.
[0108] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0109] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0110] In some embodiments, the processor is configured to execute a computer program stored in the memory to implement the following steps:
[0111] Obtain a slide file, process the slide file to obtain processed file data; input the file data into a first multimodal large model to obtain explanation script data, wherein the explanation script data includes visual element information and key text information; input the explanation script data into a second multimodal large model, align and generate fusion instruction information, wherein the fusion instruction information includes visual content instruction information and voice content instruction information; based on the visual content instruction information, call a visual material library to obtain animation data; input the voice content instruction information into a text-to-speech model to obtain dubbing data; based on the animation data and the dubbing data, synthesize a slide explanation video.
[0112] In some embodiments, the file data includes text data, image data and attribute data. When the processor is used to process the slide file to obtain the processed file data, it is used to implement: parsing the slide file to obtain text data, image data and attribute data, and the attribute data includes content type data, page sequence data or theme style data.
[0113] In some embodiments, the first multimodal large model includes an analysis model and a synthesis model. When the processor is used to input the file data into the first multimodal large model to obtain explanation script data, it is used to implement: inputting the file data into the analysis model to obtain keyword data, key sentence data and context relationship data; inputting the keyword data, the key sentence data and the context relationship data into the synthesis model to generate explanation script data.
[0114] In some embodiments, the contextual relationship data includes text-text association data and text-visual element association data, the text-text association data includes causal relationship data, contrast relationship data or hierarchical relationship data, and the text-visual element association data includes timestamp information, condition information or timing information.
[0115] In some embodiments, when the processor is used to input the text data, the image data and the attribute data into the first multimodal large model to obtain the explanation script data, it is used to implement: sentence processing of the text data to obtain sentence data; inputting the sentence data into the first pre-trained model, semantic structure analysis to obtain contextual relationship data, and / or, inputting the sentence data into the first pre-trained model, extracting key information to obtain keyword data and / or key sentence data.
[0116] In some embodiments, the key text information includes emotional tags, and the voice content instruction information includes emotional instructions; or, the visual element information includes icon information, chart information, motion information or visual element position information, and the visual content instruction information includes icon instructions, chart instructions, motion instructions or visual element position instructions.
[0117] In some embodiments, the processor is further used to implement: obtaining user change instructions, changing the explanation script data according to the user change instructions, and obtaining script change data; and training the first multimodal large model according to the script change data and the user change instructions.
[0118] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement any one of the slideshow explanation video generation methods provided in the embodiments of the present application.
[0119] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.
[0120] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for generating a slideshow video, characterized in that: include: Acquire a slide file, process the slide file, and obtain processed file data; Inputting the file data into the first multimodal large model to obtain explanation script data, wherein the explanation script data includes visual element information and key text information; Inputting the explanation script data into the second multimodal large model, aligning and generating fused instruction information, wherein the fused instruction information includes visual content instruction information and voice content instruction information; According to the visual content instruction information, calling the visual material library to obtain animation data; Inputting the voice content instruction information into a text-to-speech model to obtain dubbing data; A slideshow explanation video is synthesized based on the animation data and the dubbing data.
2. The method according to claim 1, characterized in that The file data includes text data, image data and attribute data; the slide file is processed to obtain the processed file data, including: Parsing the slide file to obtain text data, image data, and attribute data, wherein the attribute data includes content type data, page sequence data, or theme style data; The step of inputting the file data into the first multimodal large model to obtain explanation script data includes: inputting the text data, the image data and the attribute data into the first multimodal large model to obtain explanation script data.
3. The method according to claim 2, characterized in that The step of inputting the text data, the image data, and the attribute data into the first multimodal macro model to obtain explanation script data includes: Sentence processing is performed on the text data to obtain sentence data; The sentence data is input into the first pre-training model, and the semantic structure is analyzed to obtain contextual relationship data, and / or the sentence data is input into the first pre-training model, and key information is extracted to obtain keyword data and / or key sentence data.
4. The method according to claim 1, wherein The first multimodal large model includes an analysis model and a comprehensive model; the step of inputting the file data into the first multimodal large model to obtain the explanation script data includes: Inputting the document data into the analysis model to obtain keyword data, key sentence data and contextual relationship data; The keyword data, the key sentence data and the contextual relationship data are input into the comprehensive model to generate explanation script data.
5. The method according to claim 4, characterized in that The contextual relationship data includes text-text association data and text-visual element association data, the text-text association data includes causal relationship data, contrast relationship data or hierarchical relationship data, and the text-visual element association data includes timestamp information, condition information or timing information.
6. The method according to claim 1, characterized in that The key text information includes an emotion tag, and the voice content instruction information includes an emotion instruction; or The visual element information includes icon information, chart information, motion effect information or visual element position information, and the visual content instruction information includes icon instructions, chart instructions, motion effect instructions or visual element position instructions.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Obtaining a user change instruction, and changing the explanation script data according to the user change instruction to obtain script change data; The first multimodal large model is trained according to the script change data and the user change instruction.
8. A device for generating a slideshow video, characterized in that: include: A file acquisition module, used to acquire a slide file, process the slide file, and obtain processed file data; A script generation module, configured to input the file data into the first multimodal large model to obtain explanation script data; An instruction generation module, configured to input the explanation script data into a second multimodal large model, align and generate fused instruction information, wherein the fused instruction information includes visual content instruction information and voice content instruction information; A visual effect generation module, configured to call a visual material library according to the visual content instruction information to obtain animation data; A sound effect generation module, configured to input the voice content instruction information into a text-to-speech model to obtain dubbing data; The video synthesis module is used to synthesize the slideshow explanation video according to the animation data and the dubbing data.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the method for generating a slideshow explanation video according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the method for generating a slideshow explanation video according to any one of claims 1 to 7.