Presentation automatic generation method and device, electronic equipment and readable storage medium

By analyzing the PPT outline and topic information, and combining the large language model and layout design model, a PPT with a clear logical structure and unique layout design is generated. This solves the problems of loose PPT structure and lack of personalization in existing technologies, and improves the quality and efficiency of PPT generation.

CN119540403BActive Publication Date: 2026-08-25CHENGDU CELIS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411442058.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2026-08-25
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

Existing AI-generated PPTs have loose structures, weak logic, and poor novelty and uniqueness in layout and design, making it difficult to meet users' personalized needs.

Method used

By obtaining the presentation outline or theme information, the large language model interface is called to parse it, generate a sequence of pagination prompts, and combine it with the target theme style to input a pre-trained presentation layout design model to generate a PPT with a clear logical structure and unique layout design.

Benefits of technology

The generated PPT has a compact structure, good content relevance and coherence, which can meet the diverse needs of users and improve the innovation and production efficiency of PPT.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540403B_ABST
    Figure CN119540403B_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence, and provides a presentation automatic generation method and device, electronic equipment and a readable storage medium. The method comprises the following steps: calling a first model interface, inputting presentation outline information or presentation theme information, and inputting layout prompt information to obtain a first return result; generating a pagination prompt content sequence according to the first return result; inputting a target theme style and the pagination prompt content sequence into a pre-trained presentation layout design model to obtain a pagination layout design description sequence, the pagination layout design description sequence comprising a plurality of layout pages, and one layout page corresponding to one layout design description; generating a target presentation based on the pagination layout design description sequence and the target theme style, and outputting. The structure of the PPT generated by the application is more compact, the logic is better, the PPT layout design has better novelty and uniqueness, and the individualized needs of users can be better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly to a method, apparatus, electronic device, and readable storage medium for automatically generating presentations. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence (AI) technology, various AI-powered PPT (Microsoft Office PowerPoint) generators have emerged.

[0003] Existing AI-powered PPT generators lack sufficient depth and accuracy in understanding the content and structure of PPTs, resulting in PPTs with loose structures and weak logic. At the same time, their use of relatively fixed templates or rules leads to PPTs with poor novelty and uniqueness in layout design, making it difficult to meet users' various personalized needs. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, electronic device, and readable storage medium for automatically generating presentations, in order to solve the problems that existing technologies generate PPTs with loose structures, weak logic, and poor novelty and uniqueness in layout design, making it difficult to meet various personalized needs of users.

[0005] A first aspect of this application provides a method for automatically generating presentation documents, including:

[0006] Get presentation outline information or presentation theme information;

[0007] Call the first model interface, pass in the presentation outline information or presentation theme information, and layout tips, and get the first return result;

[0008] Based on the first returned result, a pagination prompt content sequence is generated. The pagination prompt content sequence includes multiple ordered pages. Each ordered page corresponds to one page prompt content. Pages belonging to the same page category have the same page prompt content, while pages belonging to different page categories have different page prompt content.

[0009] Determine the target theme and style;

[0010] Input the target theme style and pagination prompt content sequence into the pre-trained presentation layout design model to obtain the pagination layout design description sequence. The pagination layout design description sequence includes multiple layout pages, and one layout page corresponds to one layout design description.

[0011] Based on the pagination layout design description sequence and target theme style, generate the target presentation and output it.

[0012] A second aspect of this application provides a presentation automatic generation device, comprising:

[0013] The retrieval module is configured to retrieve presentation outline information or presentation theme information;

[0014] The calling module is configured to call the first model interface, pass in the presentation outline information or presentation theme information, as well as layout hints, and get the first return result;

[0015] The first generation module is configured to generate a pagination prompt content sequence based on the first returned result. The pagination prompt content sequence includes multiple ordered pages, with one ordered page corresponding to one page prompt content. Pages belonging to the same page category have the same page prompt content, while pages belonging to different page categories have different page prompt content.

[0016] The module is configured to determine the target theme style;

[0017] The input module is configured to input the target theme style and pagination prompt content sequence into a pre-trained presentation layout design model to obtain a pagination layout design description sequence, which includes multiple layout pages, with one layout page corresponding to one layout design description.

[0018] The second generation module is configured to generate and output the target presentation based on the pagination layout design description sequence and the target theme style.

[0019] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0020] A fourth aspect of this application provides a readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0021] Compared with existing technologies, the beneficial effects of this application embodiment include at least the following: By calling the first model interface to parse the input presentation outline information or presentation theme information, as well as the layout prompt information, it can ensure that the subsequently generated PPT has a clear logical structure and a complete content framework. By generating a sequence of pagination prompts based on the first returned result, the PPT content and structure can be further parsed into expressions similar to natural language, enabling the presentation layout design model to understand the content and structure of the PPT more comprehensively and accurately, thereby making the generated PPT more compact in structure and with better content relevance, coherence, and logic. In addition, the presentation layout design model used in this application embodiment can learn and integrate diverse design styles and elements, generating more novel and unique PPT layout designs, greatly improving the innovation of the generated PPT, and better meeting the various personalized needs of users. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating a method for automatically generating presentation documents provided in an embodiment of this application;

[0024] Figure 2 This is a flowchart illustrating another method for automatically generating presentation documents provided in an embodiment of this application;

[0025] Figure 3 This is a schematic diagram of the structure of an automatic presentation generation device provided in an embodiment of this application;

[0026] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0028] A method and apparatus for automatically generating presentation documents according to embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0029] In existing technologies, PPTs are automatically filled into corresponding positions based on user input using a pre-defined template library. While this method can quickly generate PPTs, it lacks flexibility and innovation, often resulting in generic PPTs that fail to meet personalized needs. Another approach involves extracting key information from user-provided text using predefined rules and formatting it according to fixed layout rules. However, this method is limited by the scope of the rules and struggles to handle complex and varied content structures. Some tools have begun to incorporate AI technology to assist PPT design, such as automatic color adjustment and intelligent chart generation. However, these functions are often isolated and lack a deep understanding and optimization of the overall PPT structure and content.

[0030] In view of this, this application provides a method for automatically generating presentations. By using innovative methods such as deep analysis of PPT elements, content understanding using a large language model, and construction of a specialized PPT design model, this method overcomes some limitations of existing technologies, thereby achieving more intelligent, flexible, and personalized automatic PPT generation, significantly improving the efficiency and quality of PPT production, and providing users with a better experience.

[0031] Furthermore, this application provides an automatic presentation generation method. By calling a first model interface, the method parses the input presentation outline information or presentation theme information, as well as layout prompts, ensuring that the subsequently generated PPT has a clear logical structure and a complete content framework. By generating a sequence of pagination prompts based on the first returned result, the PPT content and structure can be further parsed into expressions similar to natural language. This allows the presentation layout design model to understand the PPT content and structure more comprehensively and accurately, resulting in a more compact structure and better content relevance, coherence, and logic in the generated PPT. In addition, the presentation layout design model used in this application can learn and integrate diverse design styles and elements, generating more novel and unique PPT layout designs, greatly improving the innovation of the generated PPT and better meeting various personalized needs of users.

[0032] Figure 1 This is a flowchart illustrating a method for automatically generating presentation documents provided in an embodiment of this application. Figure 1 The method for automatically generating presentations can be executed by terminal devices or servers. For example... Figure 1 As shown, the method for automatically generating this presentation includes the following steps:

[0033] Step S101: Obtain presentation outline information or presentation theme information.

[0034] Specifically, presentation outline information or presentation theme information can be obtained using a general model (such as the GPT-4o model), or it can be obtained through other means, such as when the user inputs the presentation outline information or presentation theme information through a terminal device.

[0035] As an example, the presentation outline information can be formatted in Markdown format, as follows:

[0036] #Main Title

[0037] ## Level 1 Heading

[0038] ###Secondary Directory

[0039] content.

[0040] The presentation topic information can be determined by the user based on their actual needs.

[0041] Step S102: Call the first model interface, pass in the presentation outline information or presentation theme information, and layout prompts, and get the first return result.

[0042] The first model interface can be a general model interface such as the GPT-4o model.

[0043] Layout tips are instructions from general models such as GPT-4o on how to plan the pagination of a presentation based on the outline or theme information.

[0044] As an example, the layout prompt could be: "Based on the currently provided presentation outline or presentation theme information and production requirements, plan the pagination of the PPT and the content of each page, and output a JSON file in the following format: {"Cover":PPT title,"Table of Contents":[Table of Contents List],"Content Page 1"{'Table of Contents': the first-level heading to which this page belongs, 'Second-level heading' (optional): subtitle, 'content': complete content}}". The part "{"Cover":PPT title,"Table of Contents":[Table of Contents List],"Content Page 1"{'Table of Contents': the first-level heading to which this page belongs, 'Second-level heading' (optional): subtitle, 'content': complete content}}" is the first returned result.

[0045] Step S103: Based on the first returned result, generate a pagination prompt content sequence. The pagination prompt content sequence includes multiple ordered pages. Each ordered page corresponds to one page prompt content. Pages belonging to the same page category have the same page prompt content, while pages belonging to different page categories have different page prompt content.

[0046] An ordered slide refers to slides arranged in sequence, such as slide 1, slide 2, slide 3, ... slide i.

[0047] Page categories include at least one of cover page, table of contents page, content page, or end page. For example, page categories include cover page, table of contents page, content page, and end page. Another example is that page categories include cover page, content page, and end page.

[0048] As an example, a sequence of pagination prompts can be represented as [prompt1, ..., prompt1]. i ], where prompt1 represents the page prompt content corresponding to slide 1. i This indicates the page prompt content corresponding to slide page i.

[0049] Different page categories have different page prompts. For example, for a cover page, the prompt might be: 'Please design a PowerPoint cover page for {Title}.' For a table of contents page, the prompt might be: 'Please design a PowerPoint table of contents page with the following list: {List of Contents}.' For a content page, the prompt might be: 'Please design a PowerPoint content page with the first-level heading {Title}, the second-level heading {Subheading}, and the content {Complete Content}.' For a closing page, the prompt might be: 'Please design a PowerPoint closing page with the title {Title} based on the provided theme.'

[0050] Step S104: Determine the target theme style.

[0051] The themes and styles include, but are not limited to, minimalism, mechanicalism, minimalism, postmodernism, naturalism, Mediterranean style, and neoclassical style.

[0052] The target theme style can be a theme style selected by the user based on their personal preferences, or a theme style randomly selected by the system.

[0053] Step S105: Input the target theme style and pagination prompt content sequence into the pre-trained presentation layout design model to obtain the pagination layout design description sequence. The pagination layout design description sequence includes multiple layout pages, and one layout page corresponds to one layout design description.

[0054] As an example, a pagination layout design description sequence can be represented as [discribe1,…,discribe] i], where discribe1 represents the layout design description corresponding to the first layout page, discribe i This represents the layout design description corresponding to the i-th layout page.

[0055] The layout design description is a description of the page elements related to the layout of the page.

[0056] Step S106: Based on the pagination layout design description sequence and target theme style, generate the target presentation and output it.

[0057] As an example, the decoding result (PPT layout) can be obtained by decoding the pagination layout design description sequence and target theme style, and then the Python-pptx library can be used to render the decoding result as a whole to obtain the target presentation and output it.

[0058] The technical solution provided in this application embodiment, by calling the first model interface, parses the input presentation outline information or presentation theme information, as well as layout prompts, ensuring that the subsequently generated PPT has a clear logical structure and a complete content framework. By generating a sequence of pagination prompts based on the first returned result, the PPT content and structure can be further parsed into expressions similar to natural language, enabling the presentation layout design model to understand the PPT content and structure more comprehensively and accurately. This results in a more compact structure and better relevance, coherence, and logic in the generated PPT. Furthermore, the presentation layout design model used in this application embodiment can learn and integrate diverse design styles and elements, generating more novel and unique PPT layout designs, greatly improving the innovation of the generated PPT and better meeting users' personalized needs.

[0059] In some embodiments, the presentation layout design model is trained using the following training method:

[0060] Obtain multiple training samples, with each training sample representing a complete finished presentation.

[0061] For each training sample, the training sample is parsed once to obtain the first parsed file;

[0062] The first parsed file is parsed a second time to obtain the theme style file, page element file, and text content file;

[0063] Create the first prompt message, call the second model interface, and pass in the first prompt message, theme style file, page element file and text content file to get the output result of the first model;

[0064] Create a second prompt message, call the second model interface, pass in the second prompt message and the output result of the first model, and obtain the output result of the second model;

[0065] The initial model is trained using the outputs of the first and second models until the preset model convergence conditions are met, thus obtaining the presentation layout design model.

[0066] Specifically, you can obtain a finished presentation (finished PPT) through the internet or other means and save it to a folder named "PPT".

[0067] A complete finished presentation is used as a training sample. The python-pptx library can be used to parse each training sample stored in the "PPT" folder, resulting in a first parsed file in JSON format. This first parsed file can be represented as A, with all distances in cm and all angles in °.

[0068] A can be represented as:

[0069] A = {

[0070] “Slide_shape”:[H,W],

[0071] "Theme_color":T_C,

[0072] “Slides”:[S1,S2,…,S i ]

[0073] }

[0074] Where H and W represent the height and width of the PPT page; T_C represents the theme color dictionary; S1 represents the element list in the first slide; S2 represents the element list in the second slide; S i Represents the list of elements in the i-th slide, where i is the total number of slides in this PowerPoint presentation.

[0075] T_C can be represented as:

[0076] T_C={

[0077] Theme color code 1: RGB value of code 1

[0078] Theme color code 2: RGB value of code 2

[0079]

[0080] }

[0081] S i It can be represented as:

[0082] S i =[sp1,sp2,…,sp n [;sp is a dictionary representing page elements in the PowerPoint presentation. sp1 represents the first page element in the i-th slide, sp2 represents the second page element in the i-th slide, and sp...] n This represents the nth page element in the i-th slide. The i-th slide has a total of n page elements, and the earlier the page element appears, the lower its hierarchical position in the slide.

[0083] Page elements include, but are not limited to: auto_shape (built-in PPT shape), line (built-in PPT connector), place_holder, image (image element), freeform (freeform), and textbox (text box). A placeholder is a symbol that reserves a fixed position, waiting for the user to add content to it.

[0084] The six page elements listed above share some common attributes, including: left, top (the distance between the leftmost and topmost edges of the element and the top left corner of the screen), height (the height of the element), width (the width of the element), shadow (the shadow parameter of the element), fill (the fill attribute), line (the edge attribute), and text (paragraphs, text and their attributes). Their saving format can be customized.

[0085] After parsing is complete, the first parsed file can be saved in a separate json folder, which can be named "PPT+.json".

[0086] In some embodiments, the first parsing file is parsed a second time to obtain a theme style file, a page element file, and a text content file, specifically including:

[0087] Extract the theme style parameters and page size parameters of the training samples from the first parsing file and save them as a theme style file;

[0088] Extract the page elements of each page of the training samples from the first parsing file and save them as a page element file;

[0089] Iterate through the page elements of each page in the training samples, extract the text content of each page, and save it as a text content file.

[0090] As an example, the first parsed file stored in the "PPT.json" folder is parsed a second time. The theme style parameters (parameters contained in "Theme_color") and page size parameters (parameters contained in "Slide_shape"), excluding "Slides," are extracted from file A and saved separately as a file named "theme.txt," which is the theme style file. The format of the "theme.txt" file is as follows: 'height:H / width:W / theme color:[RGB value of code 1|RGB value of code 2…]'.

[0091] For the i-th member S of “Slides” in A i Iterate through its sps and save each page element separately as a file named "slide_i.txt", which is the page element file. Multiple sps are concatenated in order, with the following format: {element type / [left,top,height,width] / shadow:[…] / fill[…] / line[…] / text[…]}{…}.

[0092] For the i-th member S of “Slides” in A i Iterate through its sp and record its text content into variable T. i The text content of different sps is separated by line breaks, and page elements are separated by '----------+line break', resulting in T. i = {T1, T2, ...} and save it as a file named "PPT.json+.txt", which is the text content file.

[0093] Next, create the first message. For example, the first message might look like this:

[0094] The above is the content of a multi-page PPT. Please determine if it has a table of contents, then categorize each page into (cover, table of contents, content pages, and end page), and output it in the following format: {'Cover':[{'slide_xxx':Title},...],'Content Pages':['slide_xxx',...],End Page:['slide_xxx',...]}. If a table of contents exists, add {'Table of Contents':[{'slide_xxx':[Table of Contents List]},...]} to the JSON; if no table of contents exists, add {'Table of Contents':[Inferred Table of Contents List]} to the JSON. Please do not output any other content.

[0095] Call the second model interface (which can be a model interface of a general model such as GPT 4o model), pass in the first prompt information, theme style file, page element file and text content file mentioned above, and get the first model output result in JSON format. You can name the first model output result as "PPT+Res1.json".

[0096] Create a second prompt message. For example, the second prompt message might look like this:

[0097] "Filter out text that does not belong to the content page, determine which directory the content belongs to, and then output the complete content of the content page (note line breaks and special character escaping). If there are subdirectories, output the subdirectories. The output format is JSON: {slide_xxx:{'Directory': the first-level heading to which this page belongs, 'Second-level heading' (optional): subheading, 'content': complete content}}".

[0098] Call the second model interface, pass in the second prompt information and the output result of the first model, and get the output result of the second model in JSON format. You can name the output result of the second model as "PPT+Res2.json".

[0099] In some embodiments, the output results of the first model and the output results of the second model are used to train the initial model until a preset model convergence condition is met, thereby obtaining a presentation layout design model, specifically including:

[0100] The outputs of the first and second models are analyzed to obtain the training pagination prompt content sequence;

[0101] Using page element files as training tags, the initial model is trained using training pagination suggestion content sequences to obtain training model parameters;

[0102] If the preset model convergence condition is met, the trained model parameters will be determined as the final model parameters, and the initial model parameters of the initial model will be updated to the final model parameters to obtain the presentation layout design model.

[0103] The initial model can be one of some pre-trained open-source PPT generation algorithm models released on the Internet, or it can be a clone of a copyrighted pre-trained private PPT generation model, or it can be an untrained algorithm model.

[0104] If the number of training samples is large enough, an untrained algorithm model can be used as the initial model. If the number of training samples is small, a pre-trained open-source PPT generation algorithm model or a proprietary PPT generation model can be used as the initial model.

[0105] As an example, the outputs of the first and second models above are analyzed to obtain the training pagination prompt content sequence. The training pagination prompt content sequence can be represented as: [train_prompt1,train_prompt2…,train_prompt…] j [train_prompt1 represents the training prompts for the first slide, train_prompt2 represents the training prompts for the second slide, train_prompt] j This represents the training prompt content corresponding to the j-th slide. j is the slide page index. `train_prompt` is the training prompt content sequence for each pagination. j With pagination prompt content sequence i The structures are the same.

[0106] As an example, using the file "slide_i.txt" as the training label, prepare a training environment for LoRA (Low-Rank Adaptation, a lightweight and efficient method for fine-tuning large language models). Each time, train_prompt is used to train the initial model (either a pre-trained open-source PPT generation algorithm model or a proprietary PPT generation model). j Together with the "theme.txt" file, iterative training is performed using tools such as DeepSpeed ​​to obtain the trained model parameters. DeepSpeed ​​is a distributed training tool provided by Microsoft, designed to support larger-scale models and provide more optimization strategies and tools. For training larger models, DeepSpeed ​​offers more strategies, such as Zero and Offload.

[0107] The preset model convergence condition can be that the number of training epochs reaches a preset threshold (e.g., 5 epochs, 10 epochs, etc.) or the model accuracy reaches a preset accuracy threshold (e.g., 80%, 85%, etc.).

[0108] When training samples are limited, using the LoRA fine-tuning method for model training can improve model training efficiency and model quality.

[0109] In some embodiments, a target presentation is generated based on the pagination layout design description sequence and the target theme style, including:

[0110] The target theme style and pagination layout design description sequence are decoded to obtain a second parsing file, wherein the second parsing file has the same file structure as the first parsing file;

[0111] Generate an initial presentation based on the second parsed file;

[0112] The initial presentation is rendered to obtain the target presentation.

[0113] As an example, code can be written to describe the target theme style (e.g., the theme style file corresponding to training sample 1) and the pagination design sequence [discribe1,…,discribe] i The code decodes the first file to obtain a second parsed file in JSON format. This second parsed file has the same file structure as the first. Then, the Python-pptx library can be used to generate an initial presentation based on the second parsed file. This initial presentation is then rendered to obtain the target presentation, which can be saved as a file named "result.pptx".

[0114] In some embodiments, the method further includes:

[0115] Obtain modification instructions for the target presentation, including modifying page indexes and modifying prompts.

[0116] Based on the modified page index, determine the target page to be modified and its corresponding first page layout design description;

[0117] Based on the modification prompts, the layout design description of the first page was modified to obtain the revised presentation.

[0118] The modification prompts can be determined by the user based on their personal preferences for the layout and design of the PPT. For example, it could be "move the image down a little" or "reduce the overall transparency".

[0119] As an example, suppose the page index is modified to 5. The target page to be modified is the 5th slide of the currently generated presentation. The first page layout description is the page layout description of the 5th slide, i.e., discribe5. If the modification prompt is 'Reduce overall opacity', then reducing the overall opacity of the 5th slide in the target presentation will result in the corrected presentation.

[0120] Using the methods described above, users can modify one or more slides in a target presentation to obtain a revised presentation that better suits their individual needs.

[0121] In some embodiments, the page layout design description is modified according to the modification prompts to obtain a revised presentation, including:

[0122] Input the modified page index and modification prompt into the presentation layout design model to obtain the second page layout design description corresponding to the target modified page;

[0123] Replace the layout design description on the first page with the layout design description on the second page to obtain the revised presentation.

[0124] As an example, suppose the page indices to be modified are 1, 4, 6, and 7. The target modified pages include slides 1, 4, 6, and 7 of the currently generated target presentation. The modification prompt is "Move the image down a little." Therefore, the page indices 1, 4, 6, and 7, along with the prompt "Move the image down a little," can be input into the presentation layout design model to obtain the second page layout design descriptions, namely divide'1, divide'4, divide'6, and divide'7. Replace the first page layout design descriptions divide'1, divide'4, divide'6, and divide'7 of slides 1, 4, 6, and 7 in the target presentation with the corresponding second page layout design descriptions divide'1, divide'4, divide'6, and divide'7 to obtain the revised presentation.

[0125] Using the methods described above, users can modify one or more slides in a target presentation to obtain a revised presentation that better suits their individual needs.

[0126] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0127] Figure 2 This is a flowchart illustrating another method for automatically generating presentation documents provided in this application embodiment.

[0128] Please see Figure 2 This method can include four steps: the first is the data construction process, the second is the training process, the third is the PPT layout generation process, and the fourth is the PPT layout modification process.

[0129] The data construction process mainly includes the following steps: 1. Create and collect finished PPTs. Specifically, complete finished PPTs can be created or collected from the internet or other channels as training samples. 2. Fully parse the elements within the PPT to generate JSON. Specifically, for each training sample, the elements within its PPT are fully parsed to generate a first parsed file in JSON format. 3. Encode the JSON into natural language description. Specifically, the first parsed file is encoded to obtain a theme style file and a page element file. 4. Extract all text content. Specifically, all text content in the first parsed file is extracted to obtain a text content file.

[0130] The training process mainly includes the following steps: 1. Training the initial model. Specifically, the theme style files, page element files, and text content files obtained during the data construction process can be used to train the initial model (such as the LLAMA3 large language model), and the training model parameters can be output. 2. Adjusting model parameters. Specifically, the training model parameters can be fine-tuned based on the LoRA+ optimization algorithm. 3. Saving the model. Specifically, when the preset model convergence conditions are met, the training process ends, the presentation layout design model is obtained, and it is saved.

[0131] The PPT layout generation process mainly includes the following steps: 1. Obtain the PPT outline. 2. Paginate the PPT outline. Specifically, the large model interface can be called to plan the pagination of the PPT outline based on language-based theme hints, resulting in a sequence of pagination hint content. 3. Input the pagination hint content sequence and the selected target theme style into the saved presentation layout design model. 4. Obtain the second parsing file. Specifically, the content output by the presentation layout design model can be parsed into a PPT representation based on natural language description, and then the natural language can be decoded into JSON representing PPT elements, thus obtaining the second parsing file. 5. Generate the PPT file. Specifically, the PPT file is rendered and generated based on the second parsing file.

[0132] The main steps in modifying a PPT layout are as follows: 1. Obtain the modification prompts and input them into the saved presentation layout design model. 2. Select the page to be modified and input it into the saved presentation layout design model.

[0133] The technical solutions provided in this application have at least the following beneficial effects:

[0134] 1) By parsing various elements in the PPT into expressions similar to natural language, the Large Language Model (LLM) can understand the content and structure of the PPT more comprehensively and accurately. This deep understanding makes the generated PPT more consistent with the intent and logic of the original content, significantly improving the relevance and coherence of the content.

[0135] 2) By using the large model interface to parse the saved text content into pages and generate a PPT outline, the generated PPT can be guaranteed to have a clear logical structure and a complete content framework. This not only improves the quality of the PPT, but also helps viewers better understand and remember the presentation content.

[0136] 3) Through the above training methods, the initial model can learn and integrate diverse design styles and elements. The resulting presentation layout design model can generate more novel and unique PPT designs, which greatly improves the innovation of generated PPTs and can better meet the various personalized needs of users.

[0137] 4) The presentation layout design model trained by this application has powerful language understanding and generation capabilities, which can adapt to the needs of PPT in various fields and styles. At the same time, through continuous learning and updating, the model can continuously improve and expand its capabilities, maintaining its long-term advanced status.

[0138] 5) By breaking down the PPT design and generation process into steps such as analysis, outline generation, and design, the system's decision-making process becomes more transparent and explainable, helping users understand and trust the results generated by AI.

[0139] 6) The embodiments of this application can realize automated PPT generation, which greatly reduces the time for manual design and editing. Especially for enterprises and individuals who need to frequently create PPTs, it can significantly reduce time and labor costs.

[0140] 7) This application combines deep content understanding and intelligent design generation. Users only need to provide a basic outline or idea to obtain a high-quality PPT, which greatly simplifies the PPT production process and improves the user experience.

[0141] In summary, this application, through deep integration of natural language processing, large language models, and design intelligence, improves the quality, efficiency, and personalization of PPT generation while also enhancing the system's adaptability, interpretability, and user interactivity. These advantages make this application significantly innovative and practical in the field of AI-assisted PPT generation, contributing to technological advancements and application development in this area.

[0142] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0143] Figure 3 This is a schematic diagram of the structure of an automatic presentation generation device provided in an embodiment of this application. Figure 3 As shown, the presentation automatic generation device includes:

[0144] The acquisition module 301 is configured to acquire presentation outline information or presentation theme information;

[0145] Module 302 is configured to call the first model interface, passing in presentation outline information or presentation theme information, as well as layout hints, and obtain the first return result;

[0146] The first generation module 303 is configured to generate a pagination prompt content sequence based on the first return result. The pagination prompt content sequence includes multiple ordered pages, with one ordered page corresponding to one page prompt content. Pages belonging to the same page category have the same page prompt content, while pages belonging to different page categories have different page prompt content.

[0147] The first determination module 304 is configured to determine the target theme style;

[0148] Input module 305 is configured to input the target theme style and pagination prompt content sequence into a pre-trained presentation layout design model to obtain a pagination layout design description sequence, which includes multiple layout pages, with one layout page corresponding to one layout design description.

[0149] The second generation module 306 is configured to generate and output the target presentation based on the pagination layout design description sequence and the target theme style.

[0150] In some embodiments, the presentation layout design model is trained using the following training method:

[0151] Obtain multiple training samples, with each training sample representing a complete finished presentation.

[0152] For each training sample, the training sample is parsed once to obtain the first parsed file;

[0153] The first parsed file is parsed a second time to obtain the theme style file, page element file, and text content file;

[0154] Create the first prompt message, call the second model interface, and pass in the first prompt message, theme style file, page element file and text content file to get the output result of the first model;

[0155] Create a second prompt message, call the second model interface, pass in the second prompt message and the output result of the first model, and obtain the output result of the second model;

[0156] The initial model is trained using the outputs of the first and second models until the preset model convergence conditions are met, thus obtaining the presentation layout design model.

[0157] In some embodiments, the first parsed file is parsed a second time to obtain a theme style file, a page element file, and a text content file, including:

[0158] Extract the theme style parameters and page size parameters of the training samples from the first parsing file and save them as a theme style file;

[0159] Extract the page elements of each page of the training samples from the first parsing file and save them as a page element file;

[0160] Iterate through the page elements of each page in the training samples, extract the text content of each page, and save it as a text content file.

[0161] In some embodiments, the output results of the first model and the output results of the second model are used to train the initial model until a preset model convergence condition is met, thereby obtaining a presentation layout design model, including:

[0162] The outputs of the first and second models are analyzed to obtain the training pagination prompt content sequence;

[0163] Using page element files as training tags, the initial model is trained using training pagination suggestion content sequences to obtain training model parameters;

[0164] If the preset model convergence condition is met, the trained model parameters will be determined as the final model parameters, and the initial model parameters of the initial model will be updated to the final model parameters to obtain the presentation layout design model.

[0165] In some embodiments, the second generation module 306 described above includes:

[0166] The decoding unit is configured to decode the target theme style and pagination layout design description sequence to obtain a second parsing file, wherein the second parsing file has the same file structure as the first parsing file;

[0167] The generation unit is configured to generate an initial presentation based on the second parsed file;

[0168] The rendering unit is configured to render the initial presentation to obtain the target presentation.

[0169] In some embodiments, the above-described apparatus further includes:

[0170] The instruction acquisition module is configured to acquire modification instructions for the target presentation, including modifying page indexes and modifying prompt content;

[0171] The second determination module is configured to determine the target page to be modified and its corresponding first page layout design description based on the modified page index.

[0172] The modification module is configured to modify the layout design description of the first page based on the modification prompts, and obtain a revised presentation.

[0173] In some implementations, the modified module described above includes:

[0174] The input unit is configured to input the modified page index and modification prompt content into the presentation layout design model to obtain the second page layout design description corresponding to the target modified page.

[0175] The replacement unit is configured to replace the first page layout design description with the second page layout design description to obtain a revised presentation.

[0176] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0177] Figure 4 This is a schematic diagram of the electronic device 4 provided in an embodiment of this application. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, it implements the steps in the various method embodiments described above. Alternatively, when the processor 401 executes the computer program 403, it implements the functions of each module / unit in the various device embodiments described above.

[0178] Electronic device 4 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 4 may include, but is not limited to, processor 401 and memory 402. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 4 and does not constitute a limitation on electronic device 4. It may include more or fewer components than shown, or different components.

[0179] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0180] The memory 402 can be an internal storage unit of the electronic device 4, such as a hard disk or RAM of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 4. The memory 402 can also include both internal and external storage units of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0181] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0182] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0183] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for automatically generating presentations, characterized in that, include: Get presentation outline information or presentation theme information; Call the first model interface, pass in the presentation outline information or presentation theme information, and layout prompts, and get the first return result; Based on the first returned result, a pagination prompt content sequence is generated. The pagination prompt content sequence includes multiple ordered pages, with one ordered page corresponding to one page prompt content. Pages belonging to the same page category have the same page prompt content, while pages belonging to different page categories have different page prompt content. Determine the target theme and style; The target theme style and the pagination prompt content sequence are input into a pre-trained presentation layout design model to obtain a pagination layout design description sequence, which includes multiple layout pages, with one layout page corresponding to one layout design description. Based on the pagination layout design description sequence and the target theme style, generate the target presentation and output it.

2. The method according to claim 1, characterized in that, The presentation layout design model was trained using the following method: Obtain multiple training samples, where each training sample is a complete finished presentation. For each training sample, the training sample is parsed once to obtain the first parsed file; The first parsed file is parsed a second time to obtain the theme style file, page element file, and text content file; Create a first prompt message, call the second model interface, and pass in the first prompt message, theme style file, page element file and text content file to obtain the output result of the first model; Create a second prompt message, call the second model interface, pass in the second prompt message and the output result of the first model, and obtain the output result of the second model; Using the output results of the first model and the second model, the initial model is trained until the preset model convergence condition is met, thus obtaining the presentation layout design model.

3. The method according to claim 2, characterized in that, The first parsed file is parsed a second time to obtain a theme style file, a page element file, and a text content file, including: Extract the theme style parameters and page size parameters of the training samples from the first parsed file and save them as a theme style file; Extract the page elements of each page of the training sample from the first parsed file and save them as a page element file; Iterate through the page elements of each page of the training samples, extract the text content of each page, and save it as a text content file.

4. The method according to claim 2, characterized in that, Using the output results of the first model and the second model, the initial model is trained until the preset model convergence condition is met, resulting in a presentation layout design model, including: The output results of the first model and the second model are analyzed to obtain the training pagination prompt content sequence; Using the page element file as training tags, the initial model is trained using the training pagination prompt content sequence to obtain the training model parameters; If the preset model convergence condition is met, the training model parameters are determined as the final model parameters, and the initial model parameters of the initial model are updated to the final model parameters to obtain the presentation layout design model.

5. The method according to claim 2, characterized in that, Based on the pagination layout design description sequence and the target theme style, a target presentation is generated, including: The target theme style and pagination layout design description sequence are decoded to obtain a second parsing file, wherein the second parsing file has the same file structure as the first parsing file; Generate an initial presentation based on the second parsed file; The initial presentation is rendered to obtain the target presentation.

6. The method according to claim 1, characterized in that, The method further includes: Obtain modification instructions for the target presentation, the modification instructions including modifying the page index and modifying the prompt content; Based on the modified page index, determine the target modified page and its corresponding first page layout design description; Based on the modification prompts, the layout design description of the first page is modified to obtain a revised presentation.

7. The method according to claim 6, characterized in that, Based on the modification prompts, the page layout design description is modified to obtain a revised presentation, including: Input the modified page index and modification prompt content into the presentation layout design model to obtain the second page layout design description corresponding to the target modified page; Replace the first page layout design description with the second page layout design description to obtain the revised presentation.

8. A presentation automatic generation device, characterized in that, include: The retrieval module is configured to retrieve presentation outline information or presentation theme information; The calling module is configured to call the first model interface, pass in the presentation outline information or presentation theme information, and layout prompt information, and get the first return result; The first generation module is configured to generate a pagination prompt content sequence based on the first returned result. The pagination prompt content sequence includes multiple ordered pages, with one ordered page corresponding to one page prompt content. Pages belonging to the same page category have the same page prompt content, while pages belonging to different page categories have different page prompt content. The module is configured to determine the target theme style; The input module is configured to input the target theme style and the pagination prompt content sequence into a pre-trained presentation layout design model to obtain a pagination layout design description sequence, wherein the pagination layout design description sequence includes multiple layout pages, and one layout page corresponds to one layout design description; The second generation module is configured to generate and output the target presentation based on the pagination layout design description sequence and the target theme style.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for realizing document content generation based on text keywords

    CN117952067A

  • Transforming Content Across Visual Mediums Using Artificial Intelligence and User Generated Media

    US20240257420A1