Presentation generation method and generation apparatus
By generating presentations through large language models and database retrieval, the problem of low generation efficiency in existing tools is solved, enabling fast, efficient, and user-relevant automatic generation of presentations, thus improving the relevance and professionalism of the content.
Patent Information
- Application Number
- CN202411827195.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2026-06-19
AI Technical Summary
Existing presentation generation tools require users to manually adjust content and format, resulting in low generation efficiency and an inability to dynamically adjust according to specific needs, affecting visual effects and professionalism.
By leveraging large language models and database retrieval, an initial heading list matching the theme is generated. Based on the template page's section layout information, charts or text information are automatically populated. Combined with filtering and matching rules, presentation pages are generated, reducing manual intervention and formatting time.
It improves the efficiency and quality of presentation generation, ensures content consistency and meets user needs, reduces manual modification workload, and enhances the relevance and professionalism of the content.
Smart Images

Figure CN122242466A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for generating presentations. Background Technology
[0002] Presentations are an essential tool for information delivery, required in many scenarios to aid in expressing viewpoints. Presentations can visualize complex information, enabling audiences to understand and remember key information more intuitively. Through charts and text, presentations effectively capture audience attention and enhance information delivery. A well-structured and content-rich presentation not only helps speakers better organize their thoughts but also provides strong support during the presentation, ensuring the coherence and logic of information delivery. In summary, presentations not only improve communication efficiency and enhance information retention but also, to some extent, reflect the speaker's presentation skills.
[0003] Current technologies for presentation generation primarily rely on manual operation and auxiliary tools. Traditional presentation creation methods include using software such as Microsoft PowerPoint and Google Slides, which provide templates and editing functions, allowing users to select templates and edit content page by page according to their needs. However, manually editing the content and formatting of each page is time-consuming, especially for presentations that require frequent updates and modifications. Inconsistencies in content and formatting between different pages can easily occur, affecting the overall visual appeal and professionalism. Furthermore, tools and plugins have been developed to assist users in creating presentations more efficiently, such as through predefined templates and automated layout functions to improve production efficiency.
[0004] However, these tools typically still require users to manually adjust the content, failing to dynamically adapt to specific user needs. Users still need to invest significant time and effort in editing content, adjusting formatting, and inserting charts. This not only increases workload but also leads to low generation efficiency. Therefore, there is an urgent need for a presentation generation method to address the technical problem of low presentation generation efficiency in existing technologies. Summary of the Invention
[0005] This application provides a presentation generation method and apparatus to solve the technical problem of low presentation generation efficiency.
[0006] Firstly, this application provides a method for generating presentations, including:
[0007] Based on the theme information of the presentation to be generated, one or more historical title information matching the theme information are retrieved through a pre-set presentation database. Based on the one or more historical title information and the theme information, a large language model is used to obtain the initial title table of contents of the presentation to be generated.
[0008] Based on the initial title directory, the tagging information of the initial title directory, and the theme information, a title directory to be generated is obtained using a preset page title rewriting rule; wherein, the title directory includes a sequence of page titles;
[0009] Select any page title in the title directory to be generated, obtain the template page corresponding to the page title, and determine the fill type of each section in the template page according to the section layout information of the template page;
[0010] Based on the fill type of each section, a preset fill rule is used to obtain the information to be filled in the template page; wherein, the information to be filled includes chart information to be filled and / or text information to be filled.
[0011] Based on the section layout information, the text information to be filled and / or the chart information to be filled are filled into the template page to obtain the first presentation page corresponding to the page title, and according to the presentation database, a second presentation page matching the page title is obtained by using preset filtering and matching rules.
[0012] Traverse the sequence of page titles, sequentially obtain the first presentation page and the second presentation page corresponding to each page title, and output the first presentation page and the second presentation page corresponding to all page titles to complete the generation of the presentation.
[0013] Optionally, in the method described above, obtaining the information to be filled in the template page according to the fill type of each section and using a preset fill rule includes:
[0014] If the current section's fill type is text, then based on a preset text database, one or more similar text segments matching the page title are retrieved, and the text format of the current section is read. Based on the one or more text segments and the page title, a large language model is used to obtain the text information to be filled that matches the text format; or,
[0015] If the current section's fill type is a chart and the chart belongs to a preset chart type, then based on the chart's title information, according to the theme information and the page title, using preset query rules, the chart data is retrieved from a preset table database, and based on the chart data, the chart information to be filled that matches the chart type of the chart is retrieved.
[0016] Based on all sections of the template page, obtain the text information or chart information to be filled for each section, so as to obtain the information to be filled for the template page;
[0017] The step of filling the template page with the text information to be filled and / or the chart information to be filled based on the section layout information, and obtaining the first presentation page corresponding to the page title, includes:
[0018] Read and analyze the section layout information of the template page to determine the position, size, and type to be filled for each section;
[0019] If the current section's fill type is text, then the text information to be filled will be inserted into the current section according to its position and size; or,
[0020] If the current section's fill type is a chart and the chart belongs to a preset chart type, then the chart information to be filled will be inserted into the current section according to the current section's position and size;
[0021] Traverse all sections of the template page to obtain the first presentation page corresponding to the page title.
[0022] Optionally, in the method described above, obtaining the title directory to be generated based on the initial title directory, the tagging information of the initial title directory, and the topic information, using a preset page title rewriting rule, includes:
[0023] The tagging information of the initial title directory is obtained and analyzed. Based on the tagging information, one or more initial titles in the initial title directory are modified to obtain the first title directory of the presentation to be generated. The tagging information is the modified content marked by the user based on the theme information.
[0024] Based on the first title sequence in the first title directory, obtain any first title, and rewrite the first title using a trained title rewriting model according to the first title, the context information of the first title, and the topic information to obtain the page title;
[0025] Iterate through the first title sequence, obtain the page title after each rewritten first title, and obtain the title directory to be generated based on all page titles.
[0026] Optionally, in the method described above, before selecting any page title from the title directory to be generated and obtaining the template page corresponding to the page title, the method further includes:
[0027] Analyze the demand tags in the page title to obtain demand information for the page title; wherein, the demand information is the demand content marked by the user based on the page title;
[0028] Based on the historical page database, and according to the required information, multiple candidate template pages are retrieved and sent to the display terminal for the user to select and mark the template page with the page title;
[0029] The historical page database includes at least one of user historical data and interaction historical data;
[0030] After determining that the page title is marked as a template page, a template page tag is created for the page title;
[0031] Then, select any page title from the title directory to be generated, and obtain the template page corresponding to the page title, including:
[0032] If the page title carries a template page tag, then based on the historical page database, the template page corresponding to the page title is obtained by reading the template page tag; or...
[0033] If the page title does not carry a template page tag, then any presentation page is selected and output as the template page for the page title based on the page template database.
[0034] Optionally, in the method described above, if the current section's fill type is a chart and the chart belongs to a preset chart type, then based on the chart's title information, according to the theme information and the page title, using preset query rules, querying and obtaining chart data in a preset table database, and based on the chart data, obtaining the chart information to be filled that matches the chart type, including:
[0035] If the current section's fill type is a chart, then the chart type is obtained based on the trained image classification model;
[0036] After determining that the chart type belongs to a preset chart type, a preset image recognition model is used to obtain the chart title information, and based on the title information, according to the theme information and the page title, a large language model is used to construct and obtain the question statement;
[0037] Based on the question statement and the data architecture set of the table database, a natural language to SQL model with preset prompt words is used to obtain an SQL query statement, and the SQL query statement is executed in the table database to obtain the chart data corresponding to the question statement;
[0038] Based on the chart type, the chart data is converted into a corresponding chart data structure to obtain chart information to be filled that matches the chart type.
[0039] Optionally, in the method described above, the step of retrieving one or more historical title information matching the theme information of the presentation to be generated from a pre-set presentation database, and obtaining the initial title table of contents of the presentation to be generated using a large language model based on the one or more historical title information and the theme information, includes:
[0040] Based on the presentation to be generated, the system receives topic information input by the user, quantifies the topic information, and obtains quantified topic information.
[0041] Based on the presentation title database, retrieve one or more historical title information that matches the quantified topic information; wherein, the presentation database includes the presentation title database;
[0042] Based on the one or more historical title information and the topic information, a large language model with preset prompt words is used to generate and obtain the initial title directory of the presentation to be generated.
[0043] Optionally, in the method described above, determining the fill type of each section in the template page based on the section layout information of the template page includes:
[0044] If the template page is obtained from the page template database, then the section layout information of the template page is obtained from the page template database, and the fill type of each section in the template page is determined by analyzing the section layout information; or,
[0045] If the template page is obtained based on the template page tags, then the document parsing tool is used to obtain the section layout information of the template page, and by parsing the section layout information, the position and size of each section are determined. Based on the position and size of the sections, an image recognition algorithm is used to determine the fill type of each section in the template page.
[0046] Optionally, in the method described above, obtaining the second presentation page that matches the page title based on the presentation database using preset filtering and matching rules includes:
[0047] Based on the presentation title database, multiple similar title information is retrieved according to the page title vector of the page title; wherein, the presentation database includes a presentation title database and a presentation content database;
[0048] Based on the presentation content database and the multiple similar title information, the presentation content corresponding to each similar title information is determined, and the structural information of each presentation content is extracted; wherein, the structural information includes at least one of the following: presentation name, table of contents, and text content per page;
[0049] Based on at least one of the following: all presentation titles, table of contents, and text content per page, obtain summary information for all similar title information to obtain a summary set;
[0050] Based on the summary set, a rearranged vector set of the summary is obtained through rearranged vectorization; and based on the page title, a rearranged vector set of the page tag is obtained through rearranged vectorization.
[0051] Based on the summary and rearranged vector set, cosine similarity is used to obtain the summary information most similar to the page tag rearranged vector, so as to obtain the original presentation page corresponding to the summary information, and complete the acquisition of the second presentation page matching the page title.
[0052] Optionally, according to the method described above, the step of traversing the page title sequence, sequentially obtaining the first presentation page and the second presentation page corresponding to each page title, and outputting the first presentation page and the second presentation page corresponding to all page titles to complete the generation of the presentation includes:
[0053] Traverse the sequence of page titles and sequentially obtain the first presentation page and the second presentation page corresponding to each page title;
[0054] The first presentation page and the second presentation page are combined, and the combined presentation page is recorded as the output presentation page with the page title.
[0055] Merge all the output presentation pages corresponding to the page titles into a single presentation file, and output the presentation file to complete the presentation generation.
[0056] Secondly, this application provides a presentation generation apparatus, comprising:
[0057] The table of contents generation module is used to retrieve one or more historical title information that match the theme information of the presentation to be generated through a preset presentation database, and to obtain the initial title table of contents of the presentation to be generated using a large language model based on the one or more historical title information and the theme information.
[0058] The directory modification module is used to obtain the directory to be generated based on the initial title directory, the tag information of the initial title directory, and the theme information, using preset page title rewriting rules; wherein, the title directory includes a sequence of page titles;
[0059] The page analysis module is used to select any page title in the title directory to be generated, obtain the template page corresponding to the page title, and determine the fill type of each section in the template page according to the section layout information of the template page.
[0060] The content filling module is used to obtain the information to be filled in the template page according to the filling type of each section and using preset filling rules; wherein, the information to be filled includes chart information to be filled and / or text information to be filled;
[0061] The presentation page acquisition module is used to fill the template page with the text information to be filled and / or the chart information to be filled based on the section layout information, acquire the first presentation page corresponding to the page title, and acquire the second presentation page matching the page title according to the presentation database and using preset filtering and matching rules.
[0062] The presentation output module is used to traverse the page title sequence, sequentially obtain the first presentation page and the second presentation page corresponding to each page title, and output the first presentation page and the second presentation page corresponding to all page titles to complete the generation of the presentation.
[0063] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0064] The memory stores computer-executed instructions;
[0065] The processor executes computer execution instructions stored in the memory to implement the above method.
[0066] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the above-described method.
[0067] The presentation generation method and apparatus provided in this application improve the efficiency and accuracy of presentation production by combining a large language model, template page generation rules, and database retrieval. First, by matching topic information with historical titles, an initial title table of contents related to the topic can be quickly generated. Furthermore, by using page title rewriting rules, the generated title table of contents ensures that it meets user needs and specifications, thereby improving the relevance and applicability of the content. Second, based on the page titles and the section layout information of the template page, the fill type of each section can be determined, and the charts or text information to be filled can be automatically generated according to preset rules. This automated filling process increases the speed of presentation production and reduces manual intervention and formatting time. Simultaneously, by using a filtering mechanism that matches the second presentation page based on the page title, while ensuring the consistency of the final generated presentation content, an alternative presentation option is provided to the user, further optimizing the content presentation effect. Finally, by traversing the title sequence, the first and second presentation pages corresponding to each page title are automatically output and merged into a complete presentation file, completing the presentation generation. This process not only reduces the workload of users manually modifying and piecing together content, but also improves the quality and operability of presentations through data matching, achieving a fast, efficient, and user-friendly automatic presentation generation effect. Attached Figure Description
[0068] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0069] Figure 1 A flowchart illustrating the presentation document generation method provided in this application embodiment;
[0070] Figure 2 A schematic flowchart illustrating a method for obtaining information to be filled in a template page, as provided in an embodiment of this application.
[0071] Figure 3 This is a schematic flowchart illustrating a method for obtaining a first presentation page provided in an embodiment of this application.
[0072] Figure 4 A schematic flowchart illustrating a method for obtaining chart information to be filled that matches the chart type, as provided in an embodiment of this application.
[0073] Figure 5 This is a flowchart illustrating a method for obtaining a title table of contents to be generated, as provided in an embodiment of this application.
[0074] Figure 6 This is a schematic flowchart illustrating the method for marking a page title template page as provided in an embodiment of this application.
[0075] Figure 7 This is a flowchart illustrating a method for obtaining a template page corresponding to a page title, as provided in an embodiment of this application.
[0076] Figure 8 This is a schematic diagram of the method for determining the fill type of a plate according to an embodiment of this application;
[0077] Figure 9 This is a schematic flowchart illustrating a method for obtaining an initial title table of contents as provided in an embodiment of this application.
[0078] Figure 10 This is a schematic flowchart illustrating a method for obtaining a second presentation page provided in an embodiment of this application.
[0079] Figure 11 This is a schematic flowchart illustrating the method for generating a presentation document provided in an embodiment of this application.
[0080] Figure 12 This is a schematic diagram of a method for establishing a database provided in an embodiment of this application;
[0081] Figure 13 This is a schematic diagram of the structure of the presentation document generation device provided in the embodiments of this application;
[0082] Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0083] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0084] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0085] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0086] While existing technologies offer various presentation templates and design tools, most provide fixed templates, lacking flexible content generation and automated processing capabilities. This forces users to spend considerable time adjusting and setting up content. Furthermore, most existing systems rely on preset templates, generating presentations through simple text replacement. While this method can improve generation speed to some extent, the resulting presentations lack personalization and fail to fully utilize user-provided or historical data, making it difficult to meet the user's specific needs.
[0087] Based on the aforementioned technical problems and needs, the inventive concept of this application lies in solving the cumbersome operations and inefficiencies in the traditional presentation generation process by automating functions such as title generation, content filling, template selection, and presentation matching. Specifically, firstly, through a pre-set presentation database, the system can automatically retrieve historical title information matching the user-input topic and, combined with a large language model, generate an initial title directory related to the topic information. This process reduces the steps of manual input and title design, improves efficiency, and ensures the relevance and accuracy of the title content.
[0088] Secondly, this application further optimizes the generation process of the table of contents. Based on the initial table of contents and its tagging information, the title content is dynamically adjusted and optimized through preset page title rewriting rules. This allows for adjustments and customization of titles according to theme information and user needs, further enhancing the personalization and professionalism of the presentation. Regarding content filling, this application analyzes and determines the layout of the template page's sections, automatically identifying the fill type (such as text or chart) for each section. Based on this fill type, preset rules are used to retrieve relevant content from the chart or text database and fill it into the template, achieving automatic content filling. This method not only reduces user operation time but also improves content matching and accuracy. By merging the first and second presentation pages, it ensures that the content of each page remains consistent with the overall structure and theme of the presentation. By traversing the page title sequence, the final output presentation page more closely matches user needs, reducing the workload of manual adjustments. Therefore, this application can automatically generate a compliant presentation based on the user's theme requirements, presentation template, and content type, improving the efficiency and quality of presentation production.
[0089] This application is applicable to scenarios requiring efficient and personalized content generation. For example, in the corporate office sector, by automatically generating titles, tables of contents, and content filler, users only need to provide brief topic information to generate complete and high-quality presentations, saving significant time and effort and improving work efficiency. In common business scenarios such as report presentations, sales demonstrations, and project summaries, it can quickly generate relevant content based on different topics, ensuring the professionalism and efficiency of the presentations. Secondly, in the research field, such as academic paper reports, research results presentations, and academic conferences, researchers and scholars often need to analyze large amounts of research data and transform it into easily understandable presentation content when preparing reports. This application can automatically generate presentations based on topic information and improve the readability and professionalism of research reports through accurate content filler and chart generation. In summary, the application scenarios of this application cover multiple fields such as corporate offices, education and training, research reports, and marketing, and have broad practical application value. This application not only improves work efficiency and saves significant time and labor costs but also ensures the accuracy and professionalism of the content.
[0090] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0091] Figure 1 This is a flowchart illustrating the presentation document generation method provided in an embodiment of this application. Figure 1 As shown, it includes:
[0092] S11. Based on the theme information of the presentation to be generated, retrieve one or more historical title information that match the theme information through a pre-set presentation database, and obtain the initial title table of contents of the presentation to be generated using a large language model based on the one or more historical title information and the theme information.
[0093] In this embodiment, analyzing the topic information of the presentation to be generated effectively provides direction and a foundation for content generation. For example, the topic information of a presentation typically includes the topic name, keywords, and content outline. This information provides clear guidance for subsequent retrieval and generation. Based on the topic information of the presentation to be generated, a search operation is performed using a pre-built presentation database to obtain historical title information related to that topic. The presentation database typically contains a large number of generated presentation titles, which encompass various forms and expressions related to different topics, industries, and related fields.
[0094] Next, the retrieved historical title information is compared and analyzed with the current topic information. Historical title information typically contains title structures and keyword combinations that have proven highly relevant and acceptable in presentation generation. Therefore, an initial title list is generated using a large language model (such as a pre-trained large Transformer model). Large language models can handle large amounts of input data and generate title structures that match the current topic by understanding the relationship between topic information and historical title information. Large language models can not only infer topic-related titles based on semantic understanding but also automatically perform natural language generation, adjusting the expression of the titles to ensure they are logical, highly readable, and attractive.
[0095] This method generates an initial table of contents that, while maintaining relevance to the theme, offers excellent flexibility and scalability, meeting the personalized requirements of different presentation needs. Furthermore, by referencing historical title information, the generated titles conform to industry conventions and language habits, contributing to the professionalism and standardization of the presentation and providing a solid foundation for subsequent presentation generation. Clearly, this not only improves the efficiency of title generation but also ensures the relevance and accuracy of the generated content.
[0096] S12, based on the initial title directory, the tag information of the initial title directory, and the theme information, the title directory to be generated is obtained by adopting the preset page title rewriting rules; wherein, the title directory includes the page title sequence.
[0097] In this embodiment, a new title directory is generated based on the initial title directory, its tagging information, and theme information, using preset page title rewriting rules. This process involves optimizing and adjusting the initial title directory to ensure that the final generated title directory better meets the theme requirements in terms of semantics, structure, and expression, while more accurately reflecting the core content and logical framework of the presentation. During this process, the initial title directory is first modified based on the tagging information. For example, the tagging information can be added by the user during the initial title generation and includes the user's specific modification requirements for certain title content, such as adding, deleting, or modifying certain keywords, adjusting the tone or style of the title, etc. This tagging information provides detailed instructions for title rewriting, ensuring that the rewritten titles better meet the user's specific needs and expectations. By analyzing this tagging information, it is possible to identify which titles need adjustment, which titles need to be retained, and which titles need further detail.
[0098] After acquiring the tagging information, the initial titles are rewritten using preset page title rewriting rules, combined with the initial heading directory and topic information. These rules typically incorporate various natural language processing techniques and grammatical transformation rules. For example, synonym replacement, phrase rearrangement, and grammatical optimization may be used to optimize and rewrite the titles. Alternatively, the rules can be tailored to different scenarios and needs. These rules cover multiple levels, including title length, language style, and structural form, aiming to ensure that each page title accurately conveys the core information of the page content while conforming to the overall style and format requirements of the presentation. Furthermore, the order of the titles is optimized during the application of the rewriting rules to ensure that the generated sequence of page titles clearly reflects the overall logical structure of the presentation. Each page title is adjusted according to preset rules, ensuring that all titles are interconnected and logically clear. For example, the order of page titles may be rearranged according to the hierarchical structure of the topics, allowing the presentation content to present a logical narrative or flow.
[0099] Ultimately, after processing with the title rewriting rules, the generated table of contents will better suit the needs of the presentation. At this point, the table of contents already contains optimized page titles. This process yields a sequence of titles that accurately reflects the presentation content and meets the actual requirements, laying a solid foundation for subsequent presentation generation. In summary, by combining the initial table of contents, tag information, theme information, and preset page title rewriting rules, the generation of the table of contents is ensured, enhancing the presentation's logic and readability, and providing users with a more tailored presentation framework.
[0100] S13, select any page title in the title directory to be generated, obtain the template page corresponding to the page title, and determine the fill type of each section in the template page according to the section layout information of the template page.
[0101] In this embodiment, a template page corresponding to a page title is obtained based on the page titles in the title directory to be generated. A template page is a basic template used to present page content in a presentation, typically composed of multiple sections, each with a specific display function or containing a specific type of content. For example, a template page may include text sections, chart sections, and image sections. After obtaining the template page, its section layout information is analyzed. For example, the section layout information of a template page typically includes the position, size, style, and function of each section. The position and size of each section determine its display area on the page, and the size and proportion of the sections also determine the display space for different content. The layout information also includes the arrangement order of the sections; for example, a text section may be located at the top or left of the page, while a chart section may be located in the middle or right of the page.
[0102] Based on the established layout, the fill type for each section is further determined. For example, the fill type might be "Chart," while for sections displaying explanatory text, the fill type would be "Text." Furthermore, some sections in the template page may require both text and charts. Once the fill type for each section is determined, it becomes clear which sections need text, which need charts, and which may require images or multimedia content. This classification of fill types helps ensure that the generated presentation not only meets content requirements but also achieves a good visual presentation. Overall, identifying the fill type in the template page ensures consistency in content display and visual effects, laying the foundation for subsequent text and chart content filling. This not only improves the efficiency of presentation generation but also enhances the professionalism and visual appeal of the presentation.
[0103] S14. Based on the fill type of each section, use preset fill rules to obtain the information to be filled in the template page; wherein, the information to be filled includes chart information to be filled and / or text information to be filled.
[0104] In this embodiment, different filling rules are applied to obtain the information to be filled based on the determined section fill type (text or chart). For example, for text fill type, a search is performed using a preset text database to obtain text information related to the current page title and section requirements. On the other hand, when the section fill type is chart, data matching the required chart type is queried according to preset filling rules, and relevant chart data is obtained based on the chart title information and page title. The chart fill acquisition process typically includes querying a tabular database and matching chart types. Specifically, relevant data is searched based on the chart title and content requirements. This data can be historical data, real-time collected data, or user-input data.
[0105] Furthermore, the acquisition of information to be filled is not limited to text and chart data; it may also involve the filling of images, videos, or other multimedia content. For example, some sections may need to display photos or other forms of image data. In this case, images related to the page title and theme are selected from relevant resource libraries through a preset image database or image acquisition rules. This process involves the selection, editing, and formatting of images to ensure their coordinated display with other content on the template page. In other words, by using preset filling rules based on the filling type of each section, it is possible to retrieve and obtain text information, chart information, and other types of information to be filled from different databases.
[0106] S15, based on the section layout information, fill the template page with the text information and / or chart information to be filled, obtain the first presentation page corresponding to the page title, and obtain the second presentation page matching the page title according to the preset presentation database and the preset filtering and matching rules.
[0107] In this embodiment, the position, size, and fill type of each section are obtained by reading the section layout information of the template page, ensuring that the fill information matches the template requirements. Then, the corresponding text or chart data to be filled is filled into the corresponding section according to the instructions in the template layout information. For example, if the current section's fill type is text, the text information to be filled is inserted into the specified position according to the template page's layout parameters. This includes adjusting the text's font, size, line spacing, and other layout parameters according to text format requirements to ensure the text content is consistent with other page elements. For sections filled with charts, their position and size are first determined according to the chart type, and then the chart information to be filled is inserted into the corresponding chart area in the template. Adjustments are also needed based on elements such as the chart's display scale to ensure the chart effectively conveys the data. Once the information to be filled is in the template page, the first presentation page can be generated.
[0108] To enhance the diversity and richness of presentations, a pre-defined presentation database is used, employing preset filtering and matching rules, to find a second presentation page that matches the current page title. Matching this second presentation page involves simply reusing existing presentation content, ensuring that the selected content aligns with the requirements of the current page title. This not only leverages historical presentation data but also increases user choice. For example, if the current page involves performance analysis of a product, relevant historical presentations related to that product might be extracted from the database, including market performance and competitive analysis of the same product, enhancing the page's depth and breadth. In this way, presentation generation better meets the needs of different themes, ensuring that each page fully displays relevant information, resulting in an efficient and engaging presentation.
[0109] S16, traverse the page title sequence, obtain the first presentation page and the second presentation page corresponding to each page title in turn, output the first presentation page and the second presentation page corresponding to all page titles, and complete the generation of the presentation.
[0110] In this embodiment, each page title is processed sequentially according to a preset page title sequence. The first and second presentation pages for each page are acquired and merged one by one. Once the first and second presentation pages are merged, each pair of pages is used as an output page, and the specific position and order of each page are recorded. All pages are arranged in the order of the title sequence to form complete presentation content. Finally, all pages are merged into a single complete presentation file. This traversal output method not only improves the efficiency of presentation generation but also ensures the comprehensiveness and consistency of the presentation content.
[0111] This application improves the efficiency and accuracy of presentation creation by combining large language models, template page generation rules, and database retrieval. First, by matching topic information with historical titles, an initial title table of contents related to the topic can be quickly generated. Furthermore, page title rewriting rules ensure that the generated title table of contents meets user needs and specifications, thereby improving content relevance and applicability. Second, based on page titles and template page layout information, the fill type for each section can be determined, and charts or text information to be filled can be automatically generated according to preset rules. This automated filling process increases the speed of presentation creation and reduces manual intervention and formatting time. Simultaneously, by using a filtering mechanism that matches a second presentation page to the page title, the consistency of the final generated presentation content is ensured while providing users with another presentation option, further optimizing the content presentation. Finally, by traversing the title sequence, the first and second presentation pages corresponding to each page title are automatically output and merged into a complete presentation file, completing the presentation generation. This process not only reduces the workload of users manually modifying and piecing together content, but also improves the quality and operability of presentations through data matching, achieving a fast, efficient, and user-friendly automatic presentation generation effect.
[0112] Figure 2 This is a flowchart illustrating a method for obtaining the information to be filled in a template page, as provided in an embodiment of this application. It specifically describes one implementation of step S14 above, which involves obtaining the information to be filled in a template page. Based on the above embodiment, as... Figure 2 As shown, it includes:
[0113] S21. If the current section's fill type is text, then according to the preset text database, retrieve one or more similar text segments that match the page title, and read the text format of the current section. Based on one or more text segments and the page title, use a large language model to obtain the text information to be filled that matches the text format.
[0114] In this embodiment, when the current section's fill type is text, a pre-set text database is used to retrieve one or more text segments related to the current page title. These text segments are typically pre-organized and stored in other relevant content libraries. By matching keywords related to the topic, text related to the page title can be found. The text database can include text content from various sources, such as text from historical presentations, technical documents, white papers, industry reports, etc. Therefore, based on the keywords or topic information extracted from the page title, matching text segments are filtered to ensure that this text content effectively supports the topic. After retrieving one or more similar text segments, the text formatting requirements of the current section are analyzed. These formatting requirements include, but are not limited to, visual and typographical requirements such as font size, line spacing, paragraph style, and text alignment. Different sections may require different types of text formats. For example, some sections may require concise key points, while other sections may require detailed explanations or case studies. Therefore, based on the requirements of the section layout, the required format type for the current section is read and identified.
[0115] Next, the retrieved text segments are combined with the page title information and further processed using a large language model. The large language model generates text to be filled that matches the current section format based on the provided text segments and page title information. This process involves not only generating text content but also the tone, style, and structure of the generated text, ensuring accuracy, readability, and coherence with other parts of the page. For example, if a page title is "2024 Enterprise Market Trend Analysis," relevant paragraphs about "2024 market trends" will be retrieved, and text will be generated based on these paragraphs and the title information. If the section requires a concise, list-style presentation of key information, the text format will be automatically adjusted to conform to the preset list style, ensuring each point is clear and concise. If the section requires more explanatory text, more detailed paragraphs will be generated, covering the background of the trends, market drivers, and forecasts. Through this text generation process, thematic information is effectively combined with historical text content to generate text content that matches the section layout and format. This not only saves time for manual editing but also ensures the accuracy and professionalism of the text content, meeting the requirements for presentation generation.
[0116] S22. If the current section's fill type is a chart and the chart belongs to a preset chart type, then based on the chart's title information, according to the theme information and page title, using preset query rules, query and obtain chart data in a preset table database, and based on the chart data, obtain the chart information to be filled that matches the chart type.
[0117] In this embodiment, when the current section's fill type is a chart, and the chart belongs to a preset chart type, the chart title information corresponding to the current section is obtained from the template page. The chart title typically describes the chart's theme or the data content it displays; therefore, its information is crucial for generating correct chart data. The chart title information is extracted, and the specific content of the chart is determined based on its relevance to the page title and theme information. For example, assuming the page title is "2024 Quarterly Sales Data Analysis," the chart title might be "2024 First Quarter Sales" or "2024 Second Quarter Sales Comparison," and this information will influence the selection of chart content. Next, based on the chart title information, a preset query rule is used to query a pre-defined tabular database. The tabular database typically contains a large amount of structured data, which may come from different sources, such as a company's historical sales records, financial statements, market research data, etc. Through the set query rules, relevant chart data is filtered from the tabular database based on the chart title information (such as the specific data dimensions and time intervals displayed by the chart).
[0118] After acquiring the relevant chart data, the system formats the data according to the preset chart type (such as bar chart, line chart, pie chart, etc.) and converts it into a data structure that matches the chart type. For example, if the required chart is a bar chart, the data (such as sales revenue of different product categories) will be organized into a format suitable for bar chart display, including the sales data and time points corresponding to each bar. If the required chart is a pie chart, the data will be organized according to the data category (such as sales revenue percentage) to meet the display requirements of a pie chart. This process typically involves data normalization, classification, and sorting steps to ensure that the final chart clearly and intuitively expresses the required information. In summary, this automated chart generation method not only saves users time in manually creating charts but also ensures the relevance and quality of the chart content, helping presentations to present more professional and accurate data analysis and visualization effects.
[0119] S23. Based on all sections of the template page, obtain the text information or chart information to be filled for each section, so as to obtain the information to be filled for the template page.
[0120] In this embodiment, based on all sections of the template page, the specific information required to fill in each section is obtained through analysis, thereby completing the template page filling. The specific filling process has been described in the above steps and will not be repeated here.
[0121] Furthermore, Figure 3This is a schematic flowchart illustrating a method for obtaining a first presentation page according to an embodiment of this application. Based on the above embodiment, further explanation is provided regarding obtaining the first presentation page. Figure 3 As shown, it includes:
[0122] S31, Read and analyze the layout information of the template page to determine the position, size and type to be filled for each section;
[0123] S32, If the current block's fill type is text, then insert the text information to be filled into the current block according to the current block's position and size;
[0124] S33, If the current panel's fill type is a chart and the chart belongs to a preset chart type, then insert the chart information to be filled into the current panel according to the current panel's position and size;
[0125] S34, iterate through all sections of the template page and get the first presentation page corresponding to the page title.
[0126] In this embodiment, by reading and analyzing the layout information of the template page, the position and size of each section can be identified, thereby effectively inserting the content to be filled (text or charts) into the corresponding positions and reducing the occurrence of content overlap or misalignment. First, the layout information of the template page includes the coordinates, size, style, and layout method of each section. Furthermore, the layout information not only covers the physical location of the section (such as the coordinates of the upper left corner), but also the size, proportion, and whether additional layout adjustments are needed. By analyzing this layout data, the requirements of each section and its relative position in the template page can be fully understood, providing accurate spatial guidance for subsequent content filling. Once the layout information of each section is determined, the content to be filled is processed according to the filling type of each section. If the current section's filling type is text, the text information to be filled will be inserted into the section according to its specific position and size. Based on the section's size, the text content will be automatically wrapped, segmented, and its alignment adjusted to ensure that the text is clear, neat, and conforms to specifications within the section.
[0127] If the current section's fill type is a chart, and the chart is a preset chart type, the chart's size and position will be adjusted according to the section's layout information. Charts typically need to be scaled to fit the section's display area. Simultaneously, the chart's data will be formatted according to its type (e.g., bar chart, pie chart, line chart, etc.) to ensure a visually intuitive presentation. Inserting a chart requires not only processing the chart's data itself but also adjusting the position and size of its various elements (e.g., data labels, titles, axes, etc.). The process iterates through all sections of the template page, filling the content of each section sequentially. Finally, the template page, after filling all sections, will become the first presentation page corresponding to the page title.
[0128] In one specific embodiment, Figure 4 This is a flowchart illustrating a method for obtaining chart information matching the chart type, as provided in an embodiment of this application. It describes one implementation of step S22 above. Based on the above embodiment, as... Figure 4 As shown, it includes:
[0129] S41, If the current panel's fill type is a chart, then obtain the chart type based on the trained image classification model;
[0130] S42, after determining that the chart type belongs to the preset chart type, the preset image recognition model is used to obtain the chart title information, and based on the title information, according to the theme information and page title, the large language model is used to construct and obtain the question statement;
[0131] S43. Based on the data architecture set of the question statement and the table database, a natural language to SQL model with preset prompt words is used to obtain the SQL query statement, and the SQL query statement is executed in the table database to obtain the chart data corresponding to the question statement;
[0132] S44. Based on the chart type, convert the chart data into the corresponding chart data structure to obtain the chart information to be filled that matches the chart type.
[0133] In this embodiment, after determining that the current section's fill type is a chart, a trained image classification model analyzes the chart in the current section to determine its chart type. The image classification model uses deep learning algorithms, such as Convolutional Neural Networks (CNNs), to learn and recognize the visual features of the chart. These features may include elements such as the chart's shape, color, axes, and labels. Through training, the image classification model can accurately identify common chart types, such as bar charts, line charts, pie charts, and scatter plots. Once the chart type is identified and determined, subsequent operations are performed based on the chart type. If the chart type belongs to a preset chart type, processing continues; if it does not belong to a preset type, the chart is skipped. When the chart type is determined to be a preset type, a preset image recognition model is used to obtain the chart's title information. The image recognition model typically uses Optical Character Recognition (OCR) technology to extract text information from the chart image, recognizing the chart's title and other relevant text. This text information is the chart's metadata, used to determine the chart's theme and content.
[0134] Based on chart title information, theme information, and page title, a large language model is used to construct and obtain the question statement. This question statement is typically a query expression that specifies the chart data to be retrieved from the database. For example, for a line chart, the question statement might be: "Retrieve sales data for each month in the first quarter of 2024." After the question statement is constructed, based on the question statement and the structure of the tabular database, a pre-defined natural language to SQL model is used to transform the question statement into an SQL query statement. Natural language to SQL models typically learn how to map natural language queries (such as question statements) to corresponding SQL queries through large-scale training data. This process utilizes Natural Language Processing (NLP) technology and SQL generation algorithms. In this way, the natural language requirements can be automatically transformed into query commands that the database can understand and execute. Based on the data architecture of the tabular database (such as table structure, field names, data types, etc.), an accurate SQL query statement is generated. Then, the generated SQL query statement is executed, and the relevant chart data is retrieved from the tabular database. The database contains all the original data corresponding to the chart titles. By executing an SQL query, the required data is returned. This data may include fields such as date, number, and category. This data is used to generate charts.
[0135] After acquiring the chart data, it is converted into a data structure that matches the chart type. Different types of charts (such as bar charts, line charts, pie charts, etc.) have different requirements for data structures. For example, for bar charts, it may be necessary to group the data by category and use the value of each category as the height of the bar; for line charts, it may be necessary to sort the data in chronological order and use the values as the coordinates of the line. Based on the chart type, the raw data returned by the database is converted into a data structure suitable for that chart type to facilitate subsequent rendering and display. The converted chart data is then populated into the template page to generate the final chart content. At this point, the chart data not only conforms to the user-defined chart type but also displays the corresponding data content. During this process, the chart controls in the template page are automatically updated to display the correct chart style, title, and data content. Through the above process, relevant data can be automatically queried and processed based on the chart type, title information, and question statement, and a matching data structure can be generated according to different chart types, ultimately populating the presentation template page. This entire process not only improves the level of automation but also ensures the accuracy and visualization of the chart content, improving the efficiency and quality of presentation generation.
[0136] In one embodiment, Figure 5 This is a flowchart illustrating a method for obtaining a title table of contents to be generated, provided in an embodiment of this application. It describes one implementation of step S12 above, which involves obtaining the title table of contents. Based on the above embodiment, as... Figure 5 As shown, it includes:
[0137] S51, Obtain and analyze the tagging information of the initial title directory, and modify one or more initial titles in the initial title directory according to the tagging information to obtain the first title directory of the presentation to be generated; wherein, the tagging information is the modified content marked by the user according to the theme information;
[0138] S52, based on the first title sequence in the first title directory, obtain any first title, and rewrite the first title using a trained title rewriting model according to the first title, the context information of the first title, and the topic information, to obtain the page title;
[0139] S53, traverse the first title sequence, obtain the page title after each first title is rewritten, and obtain the title directory to be generated based on all page titles.
[0140] In this embodiment, the initial heading list first needs to be acquired and analyzed. This information is typically created by the user based on actual needs and topic information, reflecting their requirements for headline modifications. By analyzing these markers, the initial heading list is adjusted. Modifying the initial heading list generates a first heading list that meets the user's needs, making it clearer, more concise, and more logical. After modifying and generating the first heading list, each first heading is sequentially acquired based on the first heading sequence. Each heading is then rewritten using a pre-trained heading rewriting model, based on the heading, its context, and topic information. The rewriting model, trained on a large amount of data, can analyze the context and content requirements of the headings, thereby generating more logical and topic-appropriate page headings. The rewritten page headings will better reflect the topic. For example, if a heading is too general or insufficiently specific, the rewritten heading will more accurately express that content, ensuring that the page headings clearly convey the topic and key information of the page. By traversing the first heading sequence, the rewritten page headings for each first heading are sequentially acquired, and all rewritten headings are integrated to finally generate the heading list for the presentation. This process not only improves the logic and structure of the presentation but also allows for flexible adjustments to the style and expression of headings based on different content needs, ensuring that each page clearly and accurately presents the relevant content. In summary, these steps optimize the structure of the table of contents and ensure consistency and efficiency in both visual presentation and content expression. This automated heading rewriting method improves the efficiency of presentation generation while ensuring the professionalism and thematic relevance of the presentation.
[0141] In one embodiment, Figure 6 This is a schematic flowchart illustrating a method for marking a page title template page as provided in an embodiment of this application. It describes one implementation of marking a page title template page before obtaining the template page corresponding to the page title in step S13. Based on the above embodiment, as... Figure 6 As shown, it includes:
[0142] S61, Analyze the demand tags in the page title to obtain the demand information of the page title; where demand information is the demand content marked by the user based on the page title;
[0143] S62, based on the historical page database, according to the demand information, retrieves multiple candidate template pages, and sends the multiple candidate template pages to the display terminal for the user to select and mark the template page title; wherein, the historical page database includes at least one of user historical data and interaction historical data;
[0144] S63, after determining that the page title is marked with a template page, create the template page tag for the page title.
[0145] In this embodiment, the requirement tags for page titles are first analyzed to obtain specific requirement information. These requirement tags typically contain specific requirements for the page title, such as indications of content type, style, format, and structure. Users usually mark relevant requirements based on the title's theme and purpose. Natural Language Processing (NLP) technology and semantic analysis algorithms can accurately identify and extract this requirement information. For example, some titles may require a combination of text and images, while others may prefer a concise text layout. The requirement information includes requirements for page structure, the data type to be displayed (such as charts or text), and specific visual presentation requirements (such as color, font, and layout). By analyzing the requirement tags, accurate understanding and extraction of user needs are achieved, reducing human intervention and improving the efficiency and accuracy of data processing.
[0146] Next, based on the acquired requirements information, the historical page database is accessed to retrieve multiple candidate template pages that meet the requirements. The historical page database typically stores templates of various types and styles, including different layouts, color schemes, and content structures. These templates are categorized and preset according to different page requirements, facilitating quick matching and selection. For example, the historical page database contains user history data and interaction history data. This data records the presentation pages that the user has previously generated and used, along with related information. The interaction history data records the user's behavior and preferences during the template selection process. During the retrieval process, based on the title's requirements information, such as the type of charts or text styles to be presented, and the requirements for the page structure, multiple candidate template pages that meet the criteria are selected, reducing the time spent manually filtering templates. For instance, by utilizing machine learning algorithms and similarity matching technology, multiple candidate template pages that highly match the current requirements information are filtered from the database.
[0147] Subsequently, these candidate template pages are sent to the display terminal for users to select and mark. Users can see previews of different templates on the display terminal and choose the most suitable template based on their needs for page content and presentation objectives. This process, by combining historical data and real-time user interaction, ensures the personalization and efficiency of template page selection, improving user experience and the accuracy of template matching. After the user completes the template selection, a template page tag for the page title is created based on the selected template. This tag allows a specific template to be associated with the page title, enabling the selected template page to be automatically called and applied during subsequent presentation generation. In summary, through the analysis of page title requirement tags, template page retrieval and candidate generation based on historical databases, and the creation of template page tags, not only is the automation and efficiency of presentation generation improved, but also the diverse and personalized needs of users are met through requirement matching and personalized template selection.
[0148] Next, in one embodiment, Figure 7 This is a flowchart illustrating a method for obtaining a template page corresponding to a page title, provided in an embodiment of this application. Based on the above embodiments, as follows... Figure 7 As shown, it includes:
[0149] S71, if the page title carries a template page tag, then based on the historical page database, retrieve the template page corresponding to the page title by reading the template page tag;
[0150] S72, if the page title does not carry a template page tag, then select and output any presentation page as the template page for the page title based on the page template database.
[0151] In this embodiment, when a page title carries a template page tag, the template page tag is read, and the template page information corresponding to the page title is retrieved from the historical page database. This allows for direct retrieval of a template that matches the page title, ensuring that the page's visual style and content structure meet the user's needs. For example, the template page tag can point to a specific template and also carry detailed information about the template, such as its layout, fill requirements, and design style. Not all page titles carry template page tags. In some cases, the user may not have explicitly marked a template. In this case, any presentation page is selected as the template page from the page template database. The page template database contains a variety of templates to choose from. These templates are categorized according to different functions, content types, and design styles to facilitate automatic matching and selection based on the characteristics of the page title. For example, for any presentation page as the template page, a matching template can be selected based on the page title. For instance, if the title involves displaying chart data, a chart template will be selected; if the title needs to display long text or summary content, a text template will be selected. This automatic matching process allows for the selection of the most suitable template from the page template database based on the content and requirements of the page title, even without template page tags. In summary, the combination of template page tags and the page template database enables flexible selection and matching of template pages, ensuring that the content and requirements of each page title are processed accurately and efficiently. Whether through tag reading or automatic database matching, this mechanism improves the automation level, accuracy, and user satisfaction of presentation generation.
[0152] In one embodiment, Figure 8 This is a schematic flowchart illustrating the method for determining the fill type of a plate according to an embodiment of this application. It explains the implementation of step S13 above, specifically the method for determining the fill type. Based on the above embodiment, as... Figure 8 As shown, it includes:
[0153] S81, If the template page is obtained from the page template database, then obtain the section layout information of the template page from the page template database, and determine the fill type of each section in the template page by analyzing the section layout information;
[0154] S82, if the template page is obtained based on the template page tags, then the document parsing tool is used to obtain the section layout information of the template page, and the position and size of each section are determined by parsing the section layout information. Based on the position and size of the sections, an image recognition algorithm is used to determine the fill type of each section in the template page.
[0155] In this embodiment, when the template page is obtained through a page template database, the database is queried to retrieve template pages related to the current presentation theme and page title. These template pages are typically predefined and stored in the database, possessing a certain degree of universality. After retrieving the template page from the database, its section layout information is further extracted. By parsing the section layout information, the basic structure, position, size, and other information of each section are obtained, and the fill type of each section is analyzed. Fill types typically include text, charts, images, videos, etc. By analyzing the structure of the template page, the specific content type that each section needs to be filled with is determined. When the template page is obtained through template page tags, after obtaining the template page tags, the detailed structural information of the page is read through a document parsing tool, and its section layout is further analyzed. This process includes identifying the position and size of each section on the page, and using an image recognition algorithm to identify the specific type of the section. The image recognition algorithm determines the fill type of each section by analyzing the graphic features of different sections on the page, such as borders, positional relationships, and size ratios. Through image recognition technology, text boxes, chart areas, image areas, etc., can be automatically distinguished, thereby efficiently determining the content type that each section needs to be filled with. Whether retrieving template pages from a page template database or through template page tags, the ultimate goal is to ensure that the layout information of the template page is extracted and parsed, and to determine the fill type for each section based on this information. This process lays the foundation for subsequent content filling, ensuring that the structure and content of the presentation match the user's needs.
[0156] In one embodiment, Figure 9 This is a flowchart illustrating a method for obtaining an initial title table of contents as provided in an embodiment of this application. It describes one implementation of step S11 described above. Based on the above embodiment, as... Figure 9 As shown, it includes:
[0157] S91: Based on the presentation to be generated, receive the topic information input by the user, quantify the topic information, and obtain quantified topic information;
[0158] S92, based on the presentation title database, retrieve one or more historical title information that matches the quantified topic information; wherein, the presentation database includes the presentation title database;
[0159] S93, based on one or more historical title information and theme information, uses a large language model with preset prompt words to generate and obtain the initial title table of contents for the presentation to be generated.
[0160] In this embodiment, user-inputted topic information is received. This topic information is typically descriptive, relating to the core content, purpose, and context of the presentation. It may include multiple aspects, such as business reports, product introductions, academic research, and technical analysis. Next, the user-input topic information is quantified. Quantification aims to transform concepts from natural language into structured data, facilitating subsequent matching and processing. After quantifying the topic information, these quantified features are used as query conditions to retrieve matching historical title information from the presentation title database. For example, quantified keywords can be directly matched with titles in the historical title database. If there is a high degree of overlap between keywords, the historical title is considered relevant to the current topic. Alternatively, a semantic embedding model can be used to calculate the semantic similarity between the quantified topic information and historical title information. For example, cosine similarity and Euclidean distance can be used to select the historical title most relevant to the current topic. The historical title information is a pre-collected set of presentation titles related to different topics.
[0161] After obtaining historical title information that matches the quantified topic information, an initial title directory is generated using a large language model. This involves analyzing one or more historical titles and combining them with current topic information to automatically construct a hierarchical title directory. The large language model can understand the logical structure of historical titles based on contextual information, thereby generating a matching directory hierarchy. After title generation, a complete initial title directory is output. This directory includes the titles of all pages and ensures a hierarchical structure and logical relationship between the titles. In summary, an initial title directory is generated through retrieval and a large language model. This title directory not only meets user needs but can also be dynamically adjusted based on context, historical data, and topic information, ensuring that the generated presentation accurately and effectively conveys the user's core content.
[0162] In one embodiment, Figure 10 This is a schematic flowchart illustrating a method for obtaining a second presentation page according to an embodiment of this application. It provides a detailed explanation of step S15 above, specifically the process of obtaining the second presentation page. Based on the above embodiment, as... Figure 10 As shown, it includes:
[0163] S101, based on the presentation title database, retrieve multiple similar title information according to the page title vector of the page title; wherein, the presentation database includes a presentation title database and a presentation content database;
[0164] S102, based on the presentation content database and multiple similar title information, determine the presentation content corresponding to each similar title information, and extract the structural information of each presentation content; wherein, the structural information includes at least one of the following: presentation name, table of contents, and text content of each page;
[0165] S103, based on at least one of all presentation titles, headings, and text content of each page, obtain summary information of all similar heading information through summary summarization to obtain a summary set;
[0166] S104. Based on the summary set, obtain the summary rearranged vector set through rearranged vectorization, and based on the page title, obtain the page tag rearranged vector through rearranged vectorization.
[0167] S105: Based on the summary, the vector set is reordered. Cosine similarity is used to obtain the summary information that is most similar to the reordered vector of the page tag, so as to obtain the original presentation page corresponding to the summary information, and complete the acquisition of the second presentation page that matches the page title.
[0168] In this embodiment, based on the page title vector, multiple title information similar to the given title is obtained by searching the presentation database. The presentation database consists of two parts: a presentation title database and a presentation content database. The title database stores various titles of the presentation and is organized according to a certain classification system for easy retrieval. The content database contains the complete presentation content, including text, charts, and other types of information for each page. Through retrieval, multiple historical titles that are most similar to the given title in semantics and structure can be found and filtered based on the vector characteristics of the page title. This process utilizes a vector space model and a text similarity measurement method. After obtaining multiple similar title information, the presentation content database is further accessed to analyze the specific content corresponding to these titles. Each similar title has its corresponding presentation page content in the database, from which its structural information is extracted. The structural information includes the presentation name, table of contents, and text content of each page, including at least one of these. By analyzing the structural information, the layout and organization of the presentation can be obtained, providing a reference for subsequent page content generation. It will also extract specific text content, chart data, or other key information from each page as needed to enhance the relevance and targeting of the content.
[0169] To further improve matching accuracy, a summary processing step is performed based on information from all acquired presentation titles, headings, and page text content. This step uses natural language processing (NLP) technology to extract key information from the presentation content and generate a condensed set of text data through the summary. Next, a rearrangement vectorization technique is used to vectorize the summary set and page titles. This process converts both the summary set and page titles into vector representations for easy similarity calculation. Based on this, cosine similarity calculation is used to compare the similarity between the page tag rearrangement vector and the summary rearrangement vector. Cosine similarity is a standard method for measuring the directional similarity between two vectors; it can determine the semantic relevance between the title vector and the summary vector, thus identifying the most matching summary information. After completing the cosine similarity calculation, the summary information most similar to the page tag rearrangement vector is selected, and this information is used to backtrack to the original presentation page. This process ensures the accurate acquisition of the second presentation page that best matches the page title, improving the automation and accuracy of the presentation generation process.
[0170] In one embodiment, Figure 11 This is a schematic flowchart illustrating the method for generating a presentation document according to an embodiment of this application. It further explains the presentation document generation process in step S16 above. Based on the above embodiment, as... Figure 11 As shown, it includes:
[0171] S111, traverse the page title sequence and obtain the first presentation page and the second presentation page corresponding to each page title in turn;
[0172] S112, combine the first presentation page and the second presentation page, and record the combined presentation page as the output presentation page with the page title;
[0173] S113, merge all the output presentation pages corresponding to the page titles into a presentation file, and output the presentation file to complete the generation of the presentation.
[0174] In this embodiment, the page title sequence is first traversed, and each page title is processed sequentially. For each page title, the corresponding first presentation page and second presentation page are obtained. The first presentation page is typically generated by filling data into a template page and contains preliminary content corresponding to the page title. The second presentation page is obtained from the presentation database by matching it with historical presentation content and combining relevant information from the page title. Next, the first and second presentation pages are concatenated. For example, to distinguish them, the title of the first presentation page is changed to: Page Title + Generate; the title of the second presentation page is changed to: Original Presentation Title + Historical Data + Source; for example: the title of the first presentation page is: Application of Classification Algorithms + Generate; the title of the second presentation page is: Application Scenarios of Classification Models + Historical Data + "Classification Basics.pptx". After concatenation, the concatenated presentation page corresponding to each page title is recorded as the output presentation page for that page title.
[0175] Once all the output presentation pages with page titles have been assembled and recorded, all these pages are merged into a single, complete presentation file. During the merging process, the output presentation pages are arranged according to the order of their page titles to ensure that the logical structure of the final document meets the user's requirements. Finally, the merged presentation file is output, completing the presentation generation process. The output presentation file can be saved to a specified storage location or directly sent to the user for further modification. This automated and efficient process reduces the need for manual intervention while also improving the accuracy and consistency of the generation process, making the final presentation more in line with the user's requirements.
[0176] Here, we will provide further explanation regarding the establishment of the database. Figure 12 This is a schematic flowchart illustrating a method for establishing a database according to an embodiment of this application. Different document types are extracted, processed, vectorized, and stored to form different databases, facilitating subsequent data retrieval and management. Document types include presentations (PowerPoint, PPT), text files (such as Word and TXT formats), and spreadsheet files (i.e., Excel format).
[0177] For the PPT document, the entire document is uploaded and stored in an object storage service, with Minio chosen as the platform. Minio can efficiently store large-scale files and supports integration with other services. After the PPT file is uploaded to the object storage, it is parsed and split into pages. The text content of each page is extracted, and the main title of each page is extracted separately. For title processing, based on contextual information (i.e., the titles of the previous and next pages) and the overall title of the PPT document, a large language model or other model is used to rewrite the titles. This makes the rewritten titles more semantically deep and context-relevant, covering more information and ensuring greater accuracy in subsequent retrieval. The title and text content of each page are vectorized and stored separately in the presentation database, with Elasticsearch used as the vector database. Each vector item is indexed and labeled as a title vector and text content vector, respectively, and stored in the presentation content database (i.e., database A in the diagram), along with additional storage of PPT page numbers, PPT metadata information, etc.
[0178] Simultaneously, all page titles in the entire PPT are combined and concatenated into a single, complete text. This combined text is then vectorized and stored in an Elasticsearch vector database. Based on this, the concatenated title text is indexed, and additional fields are created to store the organizational structure of each title (e.g., in JSON format) for feature processing. This title information is stored in the presentation title database (i.e., database B in the diagram) to facilitate faster retrieval of topic-related title content during subsequent searches. It should be noted that the presentation title database and the presentation content database are used to distinguish the stored content; they can be separate databases. This is merely an illustrative example and is not intended to limit the scope of this application.
[0179] For Word and TXT documents, the content, including text and table data, is first extracted. Then, the text content is preprocessed and vectorized to form corresponding text vectors. Table data is also extracted and processed row by row. Both text content and table data are stored in Elasticsearch's vector database. For these text documents, they are stored in a text database (i.e., database C in the diagram), ensuring that the document content can be effectively combined and retrieved with data from other document types.
[0180] Finally, for Excel documents, we process historical data tables. First, the Excel table headers are identified and used as the database's schema data. This schema information is stored in a dedicated set called the data schema set. Then, each row of data from the Excel file is entered into the database, forming structured data, which is stored in a relational database, specifically a tabular database (database D in the diagram). The data structure in the tabular database is defined according to the Excel headers to ensure data standardization and consistency, facilitating subsequent queries and analysis.
[0181] Figure 13 This is a schematic diagram of the presentation document generation device provided in an embodiment of this application. Figure 13 As shown, the generating apparatus 13 includes:
[0182] The table of contents generation module 131 is used to retrieve one or more historical title information that match the theme information of the presentation to be generated through a pre-set presentation database, and to obtain the initial title table of contents of the presentation to be generated using a large language model based on the one or more historical title information and the theme information.
[0183] The table of contents modification module 132 is used to obtain the table of contents to be generated based on the initial table of contents, the tag information of the initial table of contents, and the theme information, using preset page title rewriting rules; wherein, the table of contents includes a sequence of page titles;
[0184] Page analysis module 133 is used to select any page title in the title directory to be generated, obtain the template page corresponding to the page title, and determine the fill type of each section in the template page according to the section layout information of the template page.
[0185] The content filling module 134 is used to obtain the information to be filled in the template page according to the filling type of each section and using preset filling rules; wherein, the information to be filled includes chart information to be filled and / or text information to be filled.
[0186] The presentation page acquisition module 135 is used to fill the template page with text information and / or chart information to be filled based on the section layout information, acquire the first presentation page corresponding to the page title, and acquire the second presentation page matching the page title according to the presentation database and using preset filtering and matching rules.
[0187] The presentation output module 136 is used to traverse the page title sequence, sequentially obtain the first presentation page and the second presentation page corresponding to each page title, and output the first presentation page and the second presentation page corresponding to all page titles to complete the generation of the presentation.
[0188] The generating apparatus provided in this embodiment can execute the generating method of the above embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0189] Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 14 As shown, the electronic device 14 includes at least one processor 141 and a memory 142. The electronic device 14 also includes a communication component 143. The processor 141, memory 142, and communication component 143 are connected via a bus 144.
[0190] In the specific implementation process, at least one processor 141 executes computer execution instructions stored in memory 142, causing at least one processor 141 to execute the presentation generation method executed on the electronic device side as described above.
[0191] The specific implementation process of processor 141 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0192] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0193] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage.
[0194] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0195] The above description of the functions implemented by electronic devices and main control devices has introduced the solutions provided by the embodiments of the present invention. It is understood that, in order to implement the above functions, the electronic device or main control device includes hardware structures and / or software modules corresponding to the execution of each function. By combining the units and algorithm steps of the various examples described in the embodiments of the present invention, the embodiments of the present invention can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of the present invention.
[0196] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above method.
[0197] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0198] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in an electronic device or a host device.
[0199] This application also provides a computer program product, comprising: a computer program stored in a readable storage medium, wherein at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to perform the scheme provided in any of the above embodiments.
[0200] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0201] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0202] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0203] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0204] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0205] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0206] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0207] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0208] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
Claims
1. A method for generating presentation slides, characterized in that, include: Based on the theme information of the presentation to be generated, one or more historical title information matching the theme information are retrieved through a pre-set presentation database. Based on the one or more historical title information and the theme information, a large language model is used to obtain the initial title table of contents of the presentation to be generated. Based on the initial title directory, the tagging information of the initial title directory, and the theme information, a title directory to be generated is obtained using a preset page title rewriting rule; wherein, the title directory includes a sequence of page titles; Select any page title in the title directory to be generated, obtain the template page corresponding to the page title, and determine the fill type of each section in the template page according to the section layout information of the template page; Based on the fill type of each section, a preset fill rule is used to obtain the information to be filled in the template page; wherein, the information to be filled includes chart information to be filled and / or text information to be filled. Based on the section layout information, the text information to be filled and / or the chart information to be filled are filled into the template page to obtain the first presentation page corresponding to the page title, and according to the presentation database, a second presentation page matching the page title is obtained by using preset filtering and matching rules. Traverse the sequence of page titles, sequentially obtain the first presentation page and the second presentation page corresponding to each page title, and output the first presentation page and the second presentation page corresponding to all page titles to complete the generation of the presentation.
2. The method according to claim 1, characterized in that, The step of obtaining the information to be filled in the template page according to the fill type of each section and using preset fill rules includes: If the current section's fill type is text, then based on a preset text database, one or more similar text segments matching the page title are retrieved, and the text format of the current section is read. Based on the one or more text segments and the page title, a large language model is used to obtain the text information to be filled that matches the text format; or, If the current section's fill type is a chart and the chart belongs to a preset chart type, then based on the chart's title information, according to the theme information and the page title, using preset query rules, the chart data is retrieved from a preset table database, and based on the chart data, the chart information to be filled that matches the chart type of the chart is retrieved. Based on all sections of the template page, obtain the text information or chart information to be filled for each section, so as to obtain the information to be filled for the template page; Based on the section layout information, the text information to be filled and / or the chart information to be filled are filled into the template page, and the first presentation page corresponding to the page title is obtained, including: Read and analyze the section layout information of the template page to determine the position, size, and type to be filled for each section; If the current section's fill type is text, then the text information to be filled will be inserted into the current section according to its position and size; or, If the current section's fill type is a chart and the chart belongs to a preset chart type, then the chart information to be filled will be inserted into the current section according to the current section's position and size; Traverse all sections of the template page to obtain the first presentation page corresponding to the page title.
3. The method according to claim 1, characterized in that, The step of obtaining the title directory to be generated based on the initial title directory, the tagging information of the initial title directory, and the theme information, using preset page title rewriting rules, includes: The tagging information of the initial title directory is obtained and analyzed. Based on the tagging information, one or more initial titles in the initial title directory are modified to obtain the first title directory of the presentation to be generated. The tagging information is the modified content marked by the user based on the theme information. Based on the first title sequence in the first title directory, obtain any first title, and rewrite the first title using a trained title rewriting model according to the first title, the context information of the first title, and the topic information to obtain the page title; Iterate through the first title sequence, obtain the page title after each rewritten first title, and obtain the title directory to be generated based on all page titles.
4. The method according to claim 1, characterized in that, Before selecting any page title from the title directory to be generated and obtaining the template page corresponding to the page title, the method further includes: Analyze the demand tags in the page title to obtain demand information for the page title; wherein, the demand information is the demand content marked by the user based on the page title; Based on the historical page database, and according to the required information, multiple candidate template pages are retrieved and sent to the display terminal for the user to select and mark the template page with the page title; The historical page database includes at least one of user historical data and interaction historical data; After determining that the page title is marked as a template page, a template page tag is created for the page title; Then, select any page title from the title directory to be generated, and obtain the template page corresponding to the page title, including: If the page title carries a template page tag, then based on the historical page database, the template page corresponding to the page title is obtained by reading the template page tag; or... If the page title does not carry a template page tag, then any presentation page is selected and output as the template page for the page title based on the page template database.
5. The method according to claim 2, characterized in that, If the current section's fill type is a chart and the chart belongs to a preset chart type, then based on the chart's title information, according to the theme information and the page title, using preset query rules, chart data is retrieved from a preset table database, and based on the chart data, the chart information to be filled that matches the chart type is retrieved, including: If the current section's fill type is a chart, then the chart type is obtained based on the trained image classification model; After determining that the chart type belongs to a preset chart type, a preset image recognition model is used to obtain the chart title information, and based on the title information, according to the theme information and the page title, a large language model is used to construct and obtain the question statement; Based on the question statement and the data architecture set of the table database, a natural language to SQL model with preset prompt words is used to obtain an SQL query statement, and the SQL query statement is executed in the table database to obtain the chart data corresponding to the question statement; Based on the chart type, the chart data is converted into a corresponding chart data structure to obtain chart information to be filled that matches the chart type.
6. The method according to claim 1, characterized in that, The step involves retrieving one or more historical title entries matching the theme information of the presentation to be generated from a pre-set presentation database, and then using a large language model to obtain the initial title list of the presentation based on the one or more historical title entries and the theme information, including: Based on the presentation to be generated, the system receives topic information input by the user, quantifies the topic information, and obtains quantified topic information. Based on the presentation title database, retrieve one or more historical title information that matches the quantified topic information; wherein, the presentation database includes the presentation title database; Based on the one or more historical title information and the topic information, a large language model with preset prompt words is used to generate and obtain the initial title directory of the presentation to be generated.
7. The method according to claim 4, characterized in that, The step of determining the fill type of each section in the template page based on the section layout information of the template page includes: If the template page is obtained from the page template database, then the section layout information of the template page is obtained from the page template database, and the fill type of each section in the template page is determined by analyzing the section layout information; or, If the template page is obtained based on the template page tags, then the document parsing tool is used to obtain the section layout information of the template page, and by parsing the section layout information, the position and size of each section are determined. Based on the position and size of the sections, an image recognition algorithm is used to determine the fill type of each section in the template page.
8. The method according to claim 1, characterized in that, The step of obtaining a second presentation page that matches the page title based on the presentation database and using preset filtering and matching rules includes: Based on the presentation title database, multiple similar title information is retrieved according to the page title vector of the page title; wherein, the presentation database includes a presentation title database and a presentation content database; Based on the presentation content database and the multiple similar title information, the presentation content corresponding to each similar title information is determined, and the structural information of each presentation content is extracted; wherein, the structural information includes at least one of the following: presentation name, table of contents, and text content per page; Based on at least one of the following: all presentation titles, table of contents, and text content per page, obtain summary information for all similar title information to obtain a summary set; Based on the summary set, a rearranged vector set of the summary is obtained through rearranged vectorization; and based on the page title, a rearranged vector set of the page tag is obtained through rearranged vectorization. Based on the summary and rearranged vector set, cosine similarity is used to obtain the summary information most similar to the page tag rearranged vector, so as to obtain the original presentation page corresponding to the summary information, and complete the acquisition of the second presentation page matching the page title.
9. The method according to claim 1, characterized in that, The process of traversing the page title sequence, sequentially obtaining the first presentation page and the second presentation page corresponding to each page title, and outputting the first presentation page and the second presentation page corresponding to all page titles completes the generation of the presentation, including: Traverse the sequence of page titles and sequentially obtain the first presentation page and the second presentation page corresponding to each page title; The first presentation page and the second presentation page are combined, and the combined presentation page is recorded as the output presentation page with the page title. Merge all the output presentation pages corresponding to the page titles into a single presentation file, and output the presentation file to complete the presentation generation.
10. A presentation generation device, characterized in that, include: The table of contents generation module is used to retrieve one or more historical title information that match the theme information of the presentation to be generated through a preset presentation database, and to obtain the initial title table of contents of the presentation to be generated using a large language model based on the one or more historical title information and the theme information. The directory modification module is used to obtain the directory to be generated based on the initial title directory, the tag information of the initial title directory, and the theme information, using preset page title rewriting rules; wherein, the title directory includes a sequence of page titles; The page analysis module is used to select any page title in the title directory to be generated, obtain the template page corresponding to the page title, and determine the fill type of each section in the template page according to the section layout information of the template page. The content filling module is used to obtain the information to be filled in the template page according to the filling type of each section and using preset filling rules; wherein, the information to be filled includes chart information to be filled and / or text information to be filled; The presentation page acquisition module is used to fill the template page with the text information to be filled and / or the chart information to be filled based on the section layout information, acquire the first presentation page corresponding to the page title, and acquire the second presentation page matching the page title according to the presentation database and using preset filtering and matching rules. The presentation output module is used to traverse the page title sequence, sequentially obtain the first presentation page and the second presentation page corresponding to each page title, and output the first presentation page and the second presentation page corresponding to all page titles to complete the generation of the presentation.