Presentation document generation method, apparatus, and medium

By combining a large language model and a graph-based text multimodal model, and using a local database to generate presentation documents, the problem of insufficient depth and personalization in generated content in existing technologies is solved, achieving efficient and accurate presentation document generation and improving user experience.

CN119830884BActive Publication Date: 2025-11-04CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411864869.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-11-04
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing technologies cannot fully utilize the potential of large language models when generating presentation documents, resulting in content that lacks depth and personalization, and is time-consuming and prone to errors.

Method used

By comparing user input with preset topics in the local database, the desired topics and subtopics are determined. The title and content of the presentation document are generated using a large language model and a graph-based text multimodal model. Combined with professional content data and pre-trained models in the local database, efficient and accurate presentation document generation is achieved.

Benefits of technology

It enables the efficient and accurate generation of presentation documents that meet user needs, improving generation efficiency and content quality while reducing the time and workload of manual editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830884B_ABST
    Figure CN119830884B_ABST
Patent Text Reader

Abstract

The present disclosure provides a presentation document generation method and device and medium, relating to the technical field of artificial intelligence, the method comprising: comparing user-initiated input information with a preset theme in a local database to determine a user-desired theme for generating a presentation document; obtaining a preset sub-theme under the user-desired theme in the local database, generating a user question according to the sub-theme, and determining several user-desired sub-themes for generating the presentation document through user question and user dialogue; generating a main title and several sub-titles of the presentation document according to the user-desired theme and the several user-desired sub-themes; obtaining content data corresponding to each user-desired sub-theme in the local database, generating sub-content of each sub-title according to the corresponding content data, and generating main content of the main title according to all the sub-content; and combining the main title and the main content and the sub-titles and the sub-content to generate the presentation document. The present disclosure combines multi-level classification of the database to generate a comprehensive and professional presentation document.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a presentation document generation method, a presentation document generation device and a computer readable storage medium. BACKGROUND

[0002] In the field of artificial intelligence, existing technologies for automatically generating presentation documents often rely on manual writing or simple template matching, which cannot fully exploit the potential of large language models. Manual writing of presentation documents is flexible, but time-consuming and prone to errors; while simple template matching may result in a lack of depth and personalization in the generated content. Therefore, how to generate efficient, accurate and demand-oriented presentation documents is still an important topic in current technical research. SUMMARY

[0003] The technical problem to be solved by the present disclosure is to provide a presentation document generation method, a presentation document generation device and a computer readable storage medium to solve the problem of how to generate efficient, accurate and demand-oriented presentation documents.

[0004] In a first aspect, the present disclosure provides a presentation document generation method, comprising:

[0005] determining a user desired theme for generating a presentation document by comparing user-initiated input information with a preset theme in a local database;

[0006] obtaining preset sub-themes under the user desired theme in the local database, generating user questions according to the sub-themes, and determining several user desired sub-themes for generating a presentation document through user questions and user dialogue;

[0007] generating a main title and several sub-titles of the presentation document according to the user desired theme and the several user desired sub-themes;

[0008] obtaining content data corresponding to each user desired sub-theme in the local database, generating sub-content for each sub-title according to the corresponding content data, and generating main content for the main title according to all the sub-content;

[0009] combining the main title and the main content and the sub-titles and the sub-content to generate a presentation document.

[0010] Further, the method further comprises establishing a local database, specifically comprising:

[0011] collecting professional content data including professional pictures and professional documents;

[0012] classifying the professional content data according to preset main categories, and setting themes for each main category;

[0013] adding a first keyword to a professional picture, extracting a professional chart in a professional document and generating a second keyword of the professional chart in combination with a context in the professional document, extracting a professional document paragraph and a third keyword thereof according to a structure of the professional document;

[0014] calculating a first similarity of the first keyword, the second keyword and the third keyword, classifying the professional picture, the professional chart and the professional document paragraph into a same sub-category if a first similarity is greater than a first threshold, and generating a sub-topic of each sub-category according to the first keyword, the second keyword and the third keyword;

[0015] storing the professional picture, the professional chart and the professional document paragraph in the local database in correspondence with the sub-topic and the topic.

[0016] Further, the method further comprises pre-training a model, specifically comprising:

[0017] pre-training a picture description capability of a graph-to-text model using at least part of data stored in the local database, including pre-training a capability of the graph-to-text model to generate a picture description text according to the professional picture and the professional chart through artificial annotation;

[0018] pre-training a content generation capability of a large language model using at least part of data stored in the local database, including a capability of generating a user question according to the sub-topic, a capability of generating a main title and a sub-title according to the topic and the sub-topic, a capability of generating an expanded document paragraph according to the professional document paragraph and the picture description text, and a capability of generating an abstract according to a plurality of expanded document paragraphs.

[0019] Further, comparing the information actively input by the user with the preset topic in the local database to determine a user desired topic of the presentation document, specifically comprising:

[0020] identifying a professional field and a topic type indicated by the user for generating the presentation document in the information actively input by the user;

[0021] calculating a second similarity of the professional field and the topic type with the preset topic in the local database;

[0022] determining the topic with a second similarity greater than a second threshold as the user desired topic for generating the presentation document.

[0023] Further, obtaining a preset sub-topic under the user desired topic in the local database, generating a user question according to the sub-topic, and determining a plurality of user desired sub-topics for generating the presentation document through the user question and a user dialogue, specifically comprising:

[0024] obtaining a preset presentation document layout template and obtaining all preset sub-topics under the user desired topic in the local database;

[0025] generating a user question including asking the user to select and modify a presentation document layout template, the total number of pages of the presentation document expected to be generated by the user, selected subtopics, and the number of subpages corresponding to each subtopic;

[0026] obtaining, through the user question and a user dialogue, a first presentation document layout template specified by the user, the total number of pages of the presentation document expected to be generated by the user, a number of user-expected subtopics selected from the subtopics, and the number of subpages corresponding to each user-expected subtopic.

[0027] Further, generating a main title and a number of sub-titles of the presentation document according to the user-expected topic and the number of user-expected subtopics, specifically including:

[0028] filtering content data in the local database according to the user-expected subtopics and the corresponding number of subpages, the content data being professional pictures, professional charts, and professional document paragraphs;

[0029] generating a sub-title of the presentation document according to a first keyword, a second keyword, and a third keyword of the filtered content data in combination with the user-expected subtopic;

[0030] generating a main title of the presentation document according to the number of sub-titles in combination with the user-expected topic.

[0031] Further, obtaining content data corresponding to each user-expected subtopic in the local database, generating sub-content of each sub-title according to the corresponding content data, and generating main content of the main title according to all sub-contents, specifically including:

[0032] obtaining professional pictures, professional charts, and professional document paragraphs filtered from the local database;

[0033] inputting the filtered professional pictures and professional charts and their first keywords and second keywords into a picture generation model to generate picture description texts;

[0034] inputting the generated picture description texts and the filtered professional document paragraphs into a large language model to generate expanded document paragraphs;

[0035] generating sub-content of each sub-title according to the expanded document paragraphs and the professional pictures and professional charts;

[0036] inputting the expanded document paragraphs of all sub-contents into a large language model to generate an abstract to generate main content of the main title.

[0037] Further, combining the main title with the main content and the sub-titles with the sub-contents to generate the presentation document, specifically including:

[0038] filling the main title with the main content and the sub-titles with the sub-contents into the first presentation document layout template and showing the user, receiving confirmation or modification instructions of the user on the displayed content to generate the presentation document.

[0039] In a second aspect, the present disclosure provides a presentation document generation apparatus, the apparatus comprising:

[0040] a theme determination unit configured to determine a user desired theme for generating the presentation document by comparing the user active input information with preset themes in a local database;

[0041] a sub-theme determination unit connected with the theme determination unit, configured to obtain preset sub-themes under the user desired theme in the local database, generate user questions according to the sub-themes, and determine several user desired sub-themes for generating the presentation document by dialoging with the user through the user questions;

[0042] a title generation unit connected with the sub-theme determination unit, configured to generate a main title and several sub-titles of the presentation document according to the user desired theme and the several user desired sub-themes;

[0043] a content generation unit connected with the title generation unit, configured to obtain content data corresponding to each user desired sub-theme in the local database, generate sub-content of each sub-title according to the corresponding content data, and generate main content of the main title according to all the sub-content;

[0044] a document generation unit connected with the content generation unit, configured to combine the main title and the main content and the sub-titles and the sub-content to generate the presentation document.

[0045] In a third aspect, the present disclosure provides a computer readable storage medium, the computer readable storage medium storing a computer program, when the computer program is run by a processor, the presentation document generation method as described above is implemented.

[0046] The present disclosure provides a presentation document generation method, a presentation document generation apparatus and a computer readable storage medium, by interacting with the user and combining with preset multi-level classification content in the local database, the theme and sub-theme of the presentation document are determined step by step, and each level of title and content corresponding to the title are generated, finally an accurate and user demand meeting presentation document is efficiently generated. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 is a flow chart of a presentation document generation method according to an embodiment of the present disclosure;

[0048] Figure 2 is a structural schematic diagram of a presentation document generation apparatus according to an embodiment of the present disclosure;

[0049] Figure 3 is a structural schematic diagram of a presentation document generation system according to an embodiment of the present disclosure;

[0050] Figure 4is a schematic diagram of a local database forming method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] In order to enable a person skilled in the art to better understand the technical solutions of the present disclosure, the embodiments of the present disclosure will be further described in detail below with reference to the drawings.

[0052] It can be understood that the specific embodiments and drawings described herein are only used to explain the present disclosure, but not to limit the present disclosure.

[0053] It can be understood that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0054] It can be understood that, for the convenience of description, only parts related to the present disclosure are shown in the drawings of the present disclosure, and parts unrelated to the present disclosure are not shown in the drawings.

[0055] It can be understood that each module and unit involved in the embodiments of the present disclosure can only correspond to one entity structure, or can be composed of multiple entity structures, or multiple modules and units can be integrated into one entity structure.

[0056] It can be understood that the functions and steps marked in the flowcharts and block diagrams of the present disclosure can occur in an order different from that marked in the drawings without conflict.

[0057] It can be understood that in the flowcharts and block diagrams of the present disclosure, the architecture, functions and operations of possible implementations of the systems, devices, apparatuses and methods according to the embodiments of the present disclosure are shown. Each block in the flowchart or block diagram can represent a module, unit, program segment, code, which contains executable instructions for realizing the specified functions. Moreover, each block or combination of blocks in the block diagram and flowchart can be realized by a hardware-based device for realizing the specified functions, or by a combination of hardware and computer instructions.

[0058] It can be understood that the modules and units involved in the embodiments of the present disclosure can be realized in the form of software or hardware, for example, the modules and units can be located in a processor.

[0059] Embodiment 1:

[0060] As shown in Figure 1 The present disclosure provides a presentation document generation method, which comprises:

[0061] S1, comparing the user-initiated input information with the preset theme in the local database to determine the user's expected theme for generating the presentation document;

[0062] S2, obtaining preset sub-topics under the user desired topic in the local database, generating user questions according to the sub-topics, and determining several user desired sub-topics for generating the presentation document by dialoguing with the user through the user questions;

[0063] S3, generating a main title and several sub-titles of the presentation document according to the user desired topic and the several user desired sub-topics;

[0064] S4, obtaining content data corresponding to each user desired sub-topic in the local database, generating sub-content of each sub-title according to the corresponding content data, and generating main content of the main title according to all the sub-content;

[0065] S5, combining the main title and the main content and the sub-titles and the sub-content to generate the presentation document.

[0066] In this embodiment, by interacting with the user and combining the multi-level classified content preset in the local database, the topic, sub-topic and each level title and the content corresponding to the title of the presentation document are gradually determined, and finally an accurate and user demand meeting presentation document is efficiently generated. The present embodiment can provide an apparatus corresponding to the method, i.e., the apparatus as shown in Figure 2 The apparatus includes a topic determination unit 1 for performing step S1, a sub-topic determination unit 2 for performing step S2, a title generation unit 3 for performing step S3, a content generation unit 4 for performing step S4, and a document generation unit 5 for performing step S5.

[0067] More specifically, the present embodiment provides a presentation document generation method and system based on knowledge base retrieval and large language model, aiming to provide a more efficient and reliable way for users to create various types of presentation documents. The method and system can meet the different needs of users in various application scenarios, thereby significantly improving the overall work efficiency and the quality of the slides. The local knowledge base can not only be updated and expanded at any time, but also effectively ensure the timeliness and accuracy of the generated content. Users can rely on this knowledge base to obtain the latest information and data, ensuring that the generated presentation document can reflect the most accurate and most relevant content at the moment.

[0068] The system (i.e., a form of presentation document generation apparatus) can include an information collection module, a model fine-tuning module, a data search module, and a human-computer interaction module, as shown in Figure 3 and 4

[0069] ​The information collection module is used to obtain the construction of the local document picture knowledge base, the large language model fine-tuning data, and the multi-modal model fine-tuning data; the model fine-tuning module fine-tunes the content generation model by using the collected document data, so that the content generation model has the ability to generate relevant content according to the user input prompt; the picture data is used to fine-tune the picture-to-text multi-modal model, so that the picture-to-text multi-modal model has the ability to automatically generate detailed descriptions of pictures; the data search module classifies the information in the user input, and retrieves relevant knowledge information in the corresponding category knowledge base for supplementation; the human-computer interaction module accepts the theme, keyword and other prompt information input by the user, and combines the supplementary information with the original information by using a prompt engineering to input the large language model, inputs the final prompt information into the large language model, generates a presentation document outline, splits the presentation document outline into multiple titles, determines the target text information corresponding to each title from the search results, then combines the target text information and the prompt information template to input the large language model, and generates a paragraph of text according to the target text information; according to the presentation document title and the module content, the database is queried and matched to find similar picture descriptions, so as to obtain the corresponding pictures; finally, the high-quality text and professional pictures output by the model are filled into the prepared presentation document template and returned to the user, and the reference document and the picture are returned to the user.

[0070] The large language model, the RAG (Retrieval-augmented Generation) retrieval enhancement technology, and the multi-modal picture-to-text large model are combined to fully exert the advantages of various technologies. The large language model has strong language understanding and generation capabilities, and can generate high-quality text content according to the input instructions. The RAG retrieval enhancement technology retrieves relevant information from an external knowledge base, enhancing the accuracy and relevance of the generated content. The multi-modal large model can process various types of data such as images and tables, providing more abundant materials and forms for the generation of presentation documents. By building and utilizing the document and picture knowledge base, not only the problem of low content quality and insufficient credibility of the large model is solved, but also the characteristics of efficient, easy-to-maintain, and good traceability are utilized, showing a broad application prospect in the field of presentation document generation and auxiliary writing. This method not only improves the accuracy and reliability of content generation, but also provides a powerful and efficient tool for users to meet their diverse needs.

[0071] In an embodiment, the method further comprises establishing a local database, specifically comprising:

[0072] Collecting professional content data including professional pictures and professional documents;

[0073] Classifying the professional content data according to preset main categories, and setting the theme of each main category;

[0074] The first keyword is added to the professional picture, the second keyword of the professional chart is generated by extracting the professional chart in the professional document and combining the context in the professional document, and the third keyword of the professional document paragraph is extracted according to the structure of the professional document;

[0075] The first similarity of the first keyword, the second keyword and the third keyword is calculated, the professional picture, the professional chart and the professional document paragraph with the first similarity greater than the first threshold value are classified into the same subcategory, and the subtopic of each subcategory is generated according to the first keyword, the second keyword and the third keyword;

[0076] The professional picture, the professional chart and the professional document paragraph are stored in the corresponding subtopic and the topic in the local database.

[0077] In the embodiment, in order to realize the retrieval based on the knowledge base, the local knowledge base needs to be established in advance, different content data is stored in multiple levels in the local knowledge base, first, multiple professional fields and main categories (topic types) can be preset, the professional field refers to an industry category such as communication, the topic type includes market analysis, project planning, technical specification and the like, these are taken as the preset main categories, the corresponding professional content data is stored under the main categories, and the corresponding topic is set, at this time, the set topic can be a selectable combination of multiple topic words, the content data under the main categories is classified again, the pictures, tables and article paragraphs with high correlation are stored under the same subcategory through similarity analysis, and the subtopic words including multiple selectable combinations are also set.

[0078] More specifically, to establish an efficient and comprehensive local knowledge base, various formats of document data are systematically collected, processed and stored. Text data can be collected from various reliable sources, and the text data can be in various formats of documents, such as Word documents, wps documents, and PDF documents, each of which contains rich information and professional knowledge suitable for building a knowledge base. The OCR (Optical Character Recognition) technology is used to recognize the text of the PDF (Portable Document Format). The text data is processed to remove redundant information, irrelevant content, and repetitive information, to ensure the neatness and effectiveness of the data. At the same time, relevant image data is collected, which can be in JPEG, PNG, TIFF, etc. The images can be illustrations and charts from the documents, or independent professional images, covering industry-related images, data visualization charts, etc. The document structure is identified using layout analysis technology, and paragraphs, tables and images are identified respectively. Based on the natural language processing capabilities of large models and multi-modal visual analysis capabilities, table information extraction and image text generation are performed, and the segmented data is vectorized and stored. At the same time, document images and descriptive text are extracted, including image titles, legends and text content related to images, or reliable image data is collected and manually annotated to describe the main content, characteristics and classification information of the images, to build a local document and image knowledge base. The extracted image description text includes image titles, legends and text content related to images, so as to facilitate future retrieval and use. Independent professional images are manually annotated to enter keywords describing the main content, characteristics and classification information of the images. The processed text and image data are stored in the local database to build a knowledge base (local database).

[0079] Based on the knowledge base constructed above, when generating a presentation document, the user input is first classified, relevant information is retrieved from the local knowledge base and integrated, and the large model is used to generate an outline based on the retrieval results and prompt word questions. After splitting the title, the target text information is determined from the retrieval results, combined with the prompt information template, and input into the large model again to generate detailed text descriptions. According to the document title and text, the corresponding images are found by searching the text, and finally the final presentation document is generated by integration. The entire process revolves around knowledge base retrieval, multiple input and output operations of the large model, and image query matching, emphasizing the use of knowledge base information to continuously improve the generation of presentation documents. It should be understood that the "local" in the local database described in the present application refers to the professional nature of the database, and does not limit its location.

[0080] In an embodiment, the method further comprises pre-training a model, specifically comprising:

[0081] Pre-training the picture description capability of the image-to-text model using at least part of the data stored in the local database, including the capability of the image-to-text model pre-trained by artificial annotation to generate picture description text according to professional pictures and professional charts;

[0082] Pre-training the content generation capability of the large language model using at least part of the data stored in the local database, including the capability of generating user questions according to sub-topics, the capability of generating main titles and sub-titles according to topics and sub-topics, the capability of generating expanded document paragraphs according to professional document paragraphs and picture description text, and the capability of generating summaries according to several expanded document paragraphs.

[0083] In this embodiment, the model includes an image-to-text model and a large language model, which can be collectively referred to as a multi-modal large model. The multi-modal large model needs to be pre-trained based on the collected professional content data, and the parameters of the large model need to be fine-tuned to adapt to the analysis, recognition and generation of new text content of professional data.

[0084] More specifically, the large language model (LLM) and the multi-modal image-to-text model represent the forefront of language understanding and generation technology.

[0085] The large language model is based on deep learning technology and is trained on a large amount of text data, with strong natural language processing capabilities. It can understand and generate complex language content, support multiple tasks such as dialogue systems, text generation, translation and summary, and through learning the grammar, semantics and context relationships of language, it realizes efficient processing of natural language.

[0086] The multi-modal image-to-text model integrates visual and language information and is committed to processing the complex relationship between images and text. It can not only generate detailed text descriptions from pictures (image-to-text), but also generate relevant images based on text (text-to-image). By combining image recognition technology with natural language processing technology, it promotes the progress of cross-modal information processing, enabling artificial intelligence to convert between vision and language more accurately and naturally. The combination of the two not only enhances the application potential of artificial intelligence in multiple fields, but also provides intelligent systems with richer and more comprehensive information processing capabilities.

[0087] In the process of giving a speech, report, etc., the process of manually editing the presentation document is usually time-consuming and laborious, especially when dealing with a large amount of text and images, the user needs to adjust the position, format and content of each element one by one, which reduces the overall writing efficiency. For similar or identical presentation documents, users often need to repeat similar editing operations, increasing the workload and time consumption. Users need to rely entirely on manual operation and cannot enjoy the convenience and efficiency brought by intelligentization. The efficiency and quality of presentation document production are limited, and the user's workload is increased, so it is necessary to improve these shortcomings through intelligent and automated technical means to improve the writing efficiency and document quality.

[0088] Intelligent presentation document generation using natural language processing (NLP) and machine learning (ML) techniques has significant advantages, which can greatly improve production efficiency and reduce labor costs. The present embodiment uses artificial intelligence to automatically generate and optimize the content and layout of the presentation document, not only saving a lot of time and effort, but also effectively shortening the production cycle to meet the needs of rapid iteration and instant update, providing an intelligent document generation method, deeply analyzing the user's input requirements, automatically extracting and organizing relevant information, and generating high-quality presentation documents, combining human expertise and proofreading to ensure that the final generated document meets the actual needs and has high quality and high reliability.

[0089] Fine-tune the large language model using collected data to enable content generation; fine-tune the image-to-text multi-modal model using image data to enable automatic generation of detailed image descriptions. Specifically, use the image data in the local database, use the image-to-text multi-modal model for image understanding, obtain the image title and image content description, and pre-train to improve the ability to automatically generate detailed image descriptions; pre-train the large language generation model using collected corpus to enable the model to capture deep semantics and contextual relationships in natural language, and enhance the professional content generation capability; and use the collected document data and its corresponding categories to train the BERT (general semantic representation model) classification model to enable document classification.

[0090] In an embodiment, S1, compare the user's active input information with the pre-set topics in the local database to determine the user's desired topic for generating the presentation document, specifically including:

[0091] Identify the professional field and topic type indicated by the user's active input information for generating the presentation document;

[0092] Calculate the second similarity between the professional field and topic type and the pre-set topics in the local database;

[0093] The theme with the second similarity greater than the second threshold value is determined as a theme expected by the user for generating the presentation document.

[0094] In this embodiment, a special presentation document generation tool can be provided, which can prompt the user to input the professional field and theme type for generating the presentation document, such as the user inputting "generate a market analysis report of a certain operator this year", and the tool then enters the flow of automatic generation of the presentation document, and acquires the corresponding theme in the local database according to the positioning of the operator to the communication industry and the market analysis report.

[0095] More specifically, the information input by the user can include various forms of text, such as questions, requests, descriptions, or instructions. The input is analyzed by using a large model to determine the category or theme to which it belongs, which is defined in advance in the local database, and the text and pictures are stored in the local database according to the categories or themes; the categories or themes can be classified based on document types, theme fields, content types, etc., for example, the input can be classified into different categories such as "market analysis report", "project planning suggestion", "technical specification description", etc.

[0096] In an embodiment, S2, the preset sub-themes under the user expected theme in the local database are acquired, and the user questions are generated according to the sub-themes, and the user questions are used for dialogue with the user to determine several user expected sub-themes for generating the presentation document, which specifically includes:

[0097] The preset presentation document layout template is acquired, and all the preset sub-themes under the user expected theme in the local database are acquired.

[0098] The user questions including inquiring the user to select and modify the presentation document layout template, the total number of pages of the presentation document expected by the user to be generated, the selected sub-themes, and the number of sub-pages corresponding to each sub-theme are generated.

[0099] The user specified first presentation document layout template, the total number of pages of the presentation document expected by the user to be generated, the several user expected sub-themes selected from the sub-themes, and the number of sub-pages corresponding to each user expected sub-theme are acquired through the dialogue between the user questions and the user.

[0100] In this embodiment, after the user expected theme is determined, the user expected sub-theme is further determined, and through further interaction with the user, the user expected presentation document layout, length, and other information are also determined, which are used as the basis for designing and generating the title and content of the presentation document.

[0101] More specifically, the user input is assigned to a preset category by a large model, and then relevant information is retrieved from the local knowledge base according to the classification result. The data in the knowledge base includes text, images and other related resources, which can provide detailed information matching the user input. The relevant information is retrieved from the corresponding local category knowledge base, and finally the user input and the supplemented information are integrated by the prompt engineering to obtain the specific requirements of the user for generating the presentation document.

[0102] In an embodiment, S3 generates the main title and several sub-titles of the presentation document according to the user desired theme and several user desired sub-themes, specifically including:

[0103] According to the user desired sub-themes and the corresponding number of sub-pages, the content data in the local database is screened, and the content data is professional pictures, professional charts and professional document paragraphs;

[0104] According to the first keyword, the second keyword and the third keyword of the screened content data, the sub-title of the presentation document is generated in combination with the user desired sub-theme;

[0105] According to the several sub-titles in combination with the user desired theme, the main title of the presentation document is generated.

[0106] In this embodiment, the efficiency of generating the sub-title of the presentation document is improved based on the pre-arranged keywords and other settings of the local database. When the user's question is output, the user's preferences can also be collected, such as asking the user's favorite chart format, keywords, etc. through the question, so as to better screen the content data meeting the user's demand. Since the content under the sub-theme is screened at this time, the sub-title is generated according to the screened content. For the generation of the main title, the user's initial input information can be mainly used to determine the main title, and the selected sub-theme can also be combined to optimize the main title, for example, to increase the sub-title. The main title in this paper refers to the title of the first page of the slide, which can include the main title and the sub-title in the traditional sense.

[0107] More specifically, according to the final prompt information input into the large language model, an outline of the presentation document is generated, and the predetermined prompt information template can be a piece of natural language, for example: "According to the following document information, generate a presentation document outline containing main titles and sub-titles", "Ensure that the outline covers all the key information provided and maintains a clear logical structure", etc., to clearly define the task and requirements of the large language model, which ensures the effectiveness of the large model output based on pre-training. The prompt information template can indicate how the large language model generates an accurate presentation document outline based on the given document information, and requires the generated presentation document outline to conform to the information provided in the dialogue. According to the final prompt information input into the large language model, relevant documents are retrieved from the filtered knowledge base, and embedding and rerank models are used to obtain search results, improving the accuracy and precision of the search. The large model first generates a presentation document outline based on the prompt word template, the question, and the search results, and then splits the presentation document outline into multiple titles.

[0108] In an embodiment, S4, obtain the content data corresponding to each user desired sub-topic in the local database, generate sub-content for each sub-title based on the corresponding content data, and generate main content for the main title based on all sub-contents, specifically including:

[0109] Obtain professional pictures, professional charts, and professional document paragraphs filtered from the local database;

[0110] Input the filtered professional pictures and professional charts and their first and second keywords into the picture-to-text model to generate picture description text;

[0111] Input the generated picture description text and the filtered professional document paragraphs into the large language model to generate expanded document paragraphs;

[0112] Generate sub-content for each sub-title based on the expanded document paragraphs and professional pictures and professional charts;

[0113] Input the expanded document paragraphs of all sub-contents into the large language model to generate an abstract, and generate the main content of the main title based on the abstract.

[0114] In this embodiment, for the generation of the presentation document content, the pre-trained picture-to-text model and the large language model are mainly relied on to achieve the generation capability of the corresponding multi-modal large model, and the order and mutual relationship of the calling of different generation capabilities of the large language model are determined when used. The calling of different generation capabilities of the large language model can be achieved by setting different prompt word templates.

[0115] More specifically, from the initial search results, determine the target text information corresponding to each title, then combine these target text information with the preset prompt information template to form a complete input format, this combined information is input again into the large language model, the large model will generate detailed text description for each title based on the target text information, until the entire document is generated; according to the title of each page of the demonstration document and the generated text, use ES (ElasticSearch) to realize query matching to find similar picture descriptions, so as to obtain the corresponding pictures; integrate the document information and the generated information of the pictures to generate the final version of the demonstration document, return the final high-quality demonstration document to the user, and output the reference data document and the pictures to the user.

[0116] Exemplarily, the outline of the demonstration document usually contains multiple titles, each title corresponds to a specific content module. When making the demonstration document, it is necessary to determine the relevant text, pictures and other specific content for each title. The specific operation steps are as follows:

[0117] 1) Content determination: For each title, determine its core information that needs to be presented, which may include a detailed text explanation, related images, charts, etc. When determining these contents, the large language model can be used to generate extended text related to the title, and the search picture function can be used to find suitable picture resources.

[0118] 2) Text generation: Use the large language model to expand each title to generate detailed text. Input the title into the model, expand the theme covered by the title through the language generation capability of the model, and form clear and coherent text paragraphs. The generated text should be able to elaborate the content involved in the title in depth and be consistent with the theme of the overall demonstration document. The RAG retrieval enhancement technology is used to retrieve related information from the local knowledge base, which enhances the accuracy and relevance of the generated content.

[0119] 3) Picture search: Use the text search function to find related pictures from the picture database according to the theme of each title and the generated text. Through database query matching, find suitable pictures or charts, and form graphic description information and add it to the text. Pictures will be used to enhance the visual effect and information transmission effect of the demonstration document.

[0120] 4) Content integration: Combine the generated text and the searched pictures to form a complete content module under each title. According to the preset format and design principles, integrate these content modules into the demonstration document.

[0121] In an embodiment, the main title and the main content and the sub-title and the sub-content are combined to generate a demonstration document, which specifically includes:

[0122] The main title and main content and sub-title and sub-content are filled into the first presentation document layout template and shown to the user, and the user's confirmation or modification instruction for the displayed content is received to generate the presentation document.

[0123] In this embodiment, the presentation effect of the presentation document is generated according to the template selected by the user, and the user's modification is received to form the final presentation document. The final presentation document, including the information modified by the user, can also be added to the local database as available data in the future.

[0124] More specifically, the system is pre-configured with multiple candidate templates, each defining the layout and design specifications of the presentation document. These candidate templates provide multiple layout options for different pages of the presentation document, including title position, font style, font size, color, image position, image size, etc. The user can choose from these pre-set templates to customize the appearance and format of the presentation document as needed. The following is a detailed operation process and template configuration process:

[0125] 1) Template configuration: Each candidate template defines the page layout in detail, including the position of the title (such as top, middle, bottom), the style of the font (such as bold, italic), the size of the font, the color scheme, etc. Text and image settings: The template also specifies the maximum amount of text allowed to be displayed on the page, the position of the image (such as left, right, center), and the size of the image (such as width and height limits).

[0126] 2) Template selection: The user browses multiple candidate templates in the template library provided by the system and selects a template that best suits their needs and the content of the presentation document. Once the user has determined the template, the system will apply the template as the target template to the design of the presentation document.

[0127] 3) Determine the layout: Based on the target template selected by the user, the system automatically applies the layout rules defined in the template, and the elements of the page such as the title, text, and images will be arranged according to the position and format requirements in the template. The text and images to be displayed are arranged according to the settings in the target template. For example, the position of the title is set according to the position of the title in the template, the style of the text is adjusted according to the font and color requirements, and the placement and scaling of the image are adjusted according to the position and size requirements of the image.

[0128] The following examples are provided: Figure 3 and 4 Two more examples are provided:

[0129] Example one: a large language model demonstration document generation system based on knowledge base retrieval, the system is composed of information collection module, model fine-tuning module, data search module and human-computer interaction module; the information collection module includes knowledge mining function, collects reliable data and picture set, mines data information, and constructs local knowledge base; the model fine-tuning module processes text and provides it to the large language model for fine-tuning, so that it has content generation capability; the constructed picture data is used to fine-tune the picture-to-text multi-modal model, so that it has the ability to automatically generate detailed description of pictures; the data search module classifies the input slide title and keyword prompt text by using the large model, then retrieves the associated document information through the classified local knowledge base, and integrates the two parts of content by using the prompt engineering; the human-computer interaction module returns the generated outline to the user according to the integrated input information by using the large language model, the large language model splits the demonstration document outline into multiple titles, determines the target text information corresponding to each title from the search results, combines the target text information and the prompt information template, and inputs them into the large language model to generate a paragraph of text; at the same time, according to the demonstration document title and the generated content, database query matching is performed to find similar picture descriptions, so as to obtain the corresponding pictures; and finally a high-quality demonstration document is generated according to the predetermined demonstration document template.

[0130] Example two: a demonstration document generation process, including: the information collection module collects reliable source text and picture data, and constructs a local knowledge base; the model fine-tuning module fine-tunes the large language model by using the collected text data after preprocessing (excluding redundant information and format standardization), so that it has content generation capability, and fine-tunes the multi-modal large model by using the collected professional pictures, so that it has the ability to describe picture content; the human-computer interaction module obtains the theme, keywords and other prompt information input by the user; the data search module classifies the user input information by using the large model, and supplements the relevant information from the corresponding category knowledge base; the human-computer interaction module integrates the original input and the supplemented information by using the prompt engineering, and transmits them to the large language model to generate a demonstration document outline; according to the final prompt information, the large language model is input to generate a demonstration document outline, the demonstration document outline is split into multiple titles, the target text information corresponding to each title is determined from the search results, the target text information and the prompt information template are combined and input into the large language model to generate a paragraph of text; according to the demonstration document title and the generated content, database query matching is performed to find similar picture descriptions, so as to obtain the corresponding pictures; finally, the high-quality text and pictures output by the model are filled into the predetermined demonstration document template and returned to the user.

[0131] The above embodiment builds a comprehensive local knowledge base by collecting data from reliable sources and high-quality professional pictures. The system can identify key information in the user's short text input, use the local knowledge base for information retrieval and supplementation, retrieve documents related to the topic, generate ppt content based on the documents, retrieve relevant pictures based on the topic, and supplement the pictures to the predetermined ppt template. This solves the problem of inaccurate and poor quality generated documents due to insufficient user input. At the same time, a professional and reliable picture database is established by using a multi-modal large language model to intelligently describe the picture library. According to the user's specific needs, the system can accurately match appropriate pictures, receive user prompt text, analyze and retrieve relevant document and picture information from the knowledge base, and combine with the preset candidate template to finally generate a complete structure and rich content presentation document. In summary, the construction of the local knowledge base greatly enhances the efficiency and accuracy of data management, and also provides reliable support for content generation, making information maintenance and traceability more efficient and convenient, and ensuring system performance and data security.

[0132] Embodiment 2:

[0133] As shown in Figure 2 The present disclosure provides a presentation document generation device, which comprises:

[0134] A theme determination unit 1 is configured to compare the user's active input information with the preset theme in the local database to determine the user's desired theme for generating the presentation document.

[0135] A sub-theme determination unit 2 is connected to the theme determination unit 1 and is configured to obtain the preset sub-theme under the user's desired theme in the local database, generate user questions according to the sub-theme, and determine several user desired sub-themes for generating the presentation document through user questions and user dialogue.

[0136] A title generation unit 3 is connected to the sub-theme determination unit 2 and is configured to generate the main title and several sub-titles of the presentation document according to the user's desired theme and the several user desired sub-themes.

[0137] A content generation unit 4 is connected to the title generation unit 2 and is configured to obtain the content data corresponding to each user desired sub-theme in the local database, generate the sub-content of each sub-title according to the corresponding content data, and generate the main content of the main title according to all the sub-contents.

[0138] A document generation unit 5 is connected to the content generation unit 4 and is configured to combine the main title and the main content and the sub-titles and the sub-contents to generate the presentation document.

[0139] In an embodiment, the device further comprises a database building unit configured to build the local database, specifically comprising:

[0140] a collection subunit configured to collect professional content data including professional pictures and professional documents;

[0141] a theme preset subunit connected to the collection subunit and configured to classify the professional content data according to preset main categories and set a theme for each main category;

[0142] a keyword subunit connected to the theme preset subunit and configured to add a first keyword to the professional pictures, extract a professional chart from the professional documents and generate a second keyword for the professional chart in combination with a context in the professional documents, and extract a third keyword according to a structure of the professional documents;

[0143] a sub-theme preset subunit connected to the keyword subunit and configured to calculate a first similarity of the first keyword, the second keyword and the third keyword, classify the professional pictures, the professional chart and the professional document paragraphs into a same sub-category if the first similarity is greater than a first threshold, and generate a sub-theme for each sub-category according to the first keyword, the second keyword and the third keyword;

[0144] a hierarchical classification storage subunit connected to the sub-theme preset subunit and configured to store the professional pictures, the professional chart and the professional document paragraphs in the local database according to the corresponding sub-theme and theme.

[0145] In an embodiment, the apparatus further comprises a pre-training unit configured to pre-train a model, specifically comprising:

[0146] a picture generation text pre-training subunit configured to pre-train a picture description capability of a picture generation text model using at least part of the data stored in the local database, including pre-training the capability of the picture generation text model to generate picture description texts according to the professional pictures and the professional chart through manual annotation;

[0147] a large model pre-training subunit connected to the picture generation text pre-training subunit and configured to pre-train a content generation capability of a large language model using at least part of the data stored in the local database, including the capability to generate user questions according to the sub-theme, the capability to generate main titles and sub-titles according to the theme and the sub-theme, the capability to generate expanded document paragraphs according to the professional document paragraphs and the picture description texts, and the capability to generate an abstract according to a plurality of expanded document paragraphs.

[0148] In an embodiment, the theme determination unit 1 specifically comprises:

[0149] an identification subunit configured to identify a professional field and a theme type indicated by the user for generating the presentation document in the active input information;

[0150] a calculation subunit connected to the identification subunit and configured to calculate a second similarity of the professional field and the theme type with preset themes in the local database.

[0151] The determining subunit is connected with the calculating subunit and is configured to determine the topic with the second similarity greater than the second threshold as the user expected topic for generating the presentation document.

[0152] In an embodiment, the subtopic determining unit 2 specifically comprises:

[0153] The preset obtaining subunit is configured to obtain a preset presentation document layout template and all preset subtopics of the user expected topic in a local database.

[0154] The question generating subunit is connected with the preset obtaining subunit and is configured to generate a user question including inquiring the user to select and modify the presentation document layout template, the total number of pages of the presentation document expected to be generated by the user, the selected subtopic and the number of subpages corresponding to the subtopic.

[0155] The dialogue analysis subunit is connected with the question generating subunit and is configured to obtain, through the user question and the dialogue with the user, the first presentation document layout template specified by the user, the total number of pages of the presentation document expected to be generated by the user, the number of subtopics selected by the user from the subtopics and the number of subpages corresponding to each user expected subtopic.

[0156] In an embodiment, the title generating unit 3 specifically comprises:

[0157] The screening subunit is configured to screen the content data in the local database according to the user expected subtopic and the corresponding number of subpages, wherein the content data is professional pictures, professional charts and professional document paragraphs.

[0158] The subheading generating subunit is connected with the screening subunit and is configured to generate the subheading of the presentation document according to the first keyword, the second keyword and the third keyword of the screened content data in combination with the user expected subtopic.

[0159] The main heading generating subunit is connected with the subheading generating subunit and is configured to generate the main heading of the presentation document according to the number of subheadings in combination with the user expected topic.

[0160] In an embodiment, the content generating unit 4 specifically comprises:

[0161] The content obtaining subunit is configured to obtain the professional pictures, the professional charts and the professional document paragraphs screened from the local database.

[0162] The figure generating text subunit is connected with the content obtaining subunit and is configured to input the screened professional pictures and professional charts and the first keyword and the second keyword thereof into a figure generating text model to generate picture description text.

[0163] The expansion subunit is connected with the picture generation subunit, and is configured to input the generated picture description text and the screened professional document paragraph into a large language model to generate an expanded document paragraph;

[0164] The sub-content generation subunit is connected with the expansion subunit, and is configured to generate sub-content of each sub-title according to the expanded document paragraph and the professional picture and the professional chart;

[0165] The main content generation subunit is connected with the sub-content generation subunit, and is configured to input the expanded document paragraph of all the sub-contents into the large language model to generate a summary to generate main content of the summary main title.

[0166] In an embodiment, the document generation unit 5 specifically includes:

[0167] The filling subunit is configured to fill the main title and the main content and the sub-title and the sub-content into the first presentation document layout template and show the user;

[0168] The receiving instruction subunit is connected with the filling subunit, and is configured to receive a confirmation or modification instruction of the user on the displayed content to generate the presentation document.

[0169] Embodiment 3

[0170] Embodiment 3 of the present disclosure provides a computer readable storage medium, the computer readable storage medium stores a computer program, when the computer program is run by a processor, the presentation document generation method as described in embodiment 1 is implemented, or the presentation document generation device as described in embodiment 2 is implemented.

[0171] The computer readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information, such as computer readable instructions, data structures, computer program units or other data. The computer readable storage medium includes but is not limited to RAM (Random Access Memory, Random Access Memory), ROM (Read-Only Memory, Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory, Electrically Erasable Programmable Read-Only Memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory, Compact Disc Read-Only Memory), digital versatile disc (DVD) or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other medium that can be used to store desired information and can be accessed by a computer.

[0172] In addition, the present disclosure can also provide a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program stored in the memory, the processor executes the presentation document generation method according to Embodiment 1. The computer device can be the presentation document generation device according to Embodiment 2.

[0173] The memory is connected with the processor, and the memory can be a flash memory or a read-only memory or other memories, and the processor can be a central processing unit or a single-chip microcomputer.

[0174] Embodiments 1-3 of the present disclosure provide a presentation document generation method, a presentation document generation device and a computer readable storage medium, which gradually determine the theme, sub-theme of the presentation document and generate the titles and the content corresponding to the titles of each level by interacting with the user and combining the multi-level classified content preset by the local database, and finally efficiently generate the accurate and user-demand-compliant presentation document.

[0175] It can be understood that the above embodiments are only exemplary embodiments adopted for illustrating the principles of the present disclosure, and the present disclosure is not limited thereto. Various modifications and improvements can be made by those skilled in the art without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also considered as the protection scope of the present disclosure.

Claims

1. A method for generating a presentation document, characterized in that, The method includes: Compare the user-inputted information with the preset themes in the local database to determine the user's desired theme for generating the presentation document; Retrieve the preset subtopics under the user's expected topic in the local database, generate user questions based on the subtopics, and engage in dialogue with the user through the user questions to determine several user-expected subtopics for generating the presentation document; Generate the main title and several subtitles of the presentation document based on the user's desired theme and several user-desired sub-themes; Retrieve content data corresponding to each user's desired subtopic from the local database, generate sub-content for each sub-heading based on the corresponding content data, and generate the main content of the main title based on all sub-content. Specifically, this includes: Retrieves professional images, charts, and document paragraphs filtered from a local database, along with the first keyword added to the professional images and the second keyword extracted from the professional charts in the documents and generated based on the context within the documents. Input the selected professional images and charts, along with their primary and secondary keywords, into the image-to-text model to generate descriptive text for the images. The generated image description text and the selected professional document paragraphs are input into the large language model to generate expanded document paragraphs. Based on the expanded document paragraphs and professional images and charts, generate sub-content for each subheading. The expanded document paragraphs of all sub-contents are input into the large language model to generate a summary, which in turn generates the main content of the main title; Combine the main title and main content, and subtitles and subcontent to generate a presentation document.

2. The method according to claim 1, characterized in that, The method also includes establishing a local database, specifically including: Collect professional content data, including professional images and documents; Classify professional content data according to preset main categories, and set the theme for each main category; Add a first keyword to professional images, extract professional charts from professional documents and generate second keywords for professional charts by combining them with the context of the professional documents, and extract paragraphs and their third keywords from the professional documents according to their structure. Calculate the first similarity of the first keyword, the second keyword, and the third keyword; classify professional images, professional charts, and professional document paragraphs with a first similarity greater than the first threshold into the same sub-category; and generate sub-topics for each sub-category based on the first keyword, the second keyword, and the third keyword. Professional images, charts, and document paragraphs are categorized into their corresponding subtopics and themes and stored in a local database.

3. The method according to claim 2, characterized in that, The method also includes a pre-trained model, specifically comprising: The image description capability of the image-to-text model is pre-trained using at least a portion of the data stored in the local database, including the ability of the pre-trained image-to-text model to generate image description text based on professional images and charts through manual annotation; The content generation capability of a large language model pre-trained using at least a portion of the data stored in a local database includes: the ability to generate user questions based on subtopics; the ability to generate main titles and subtitles based on topics and subtopics; the ability to generate expanded document paragraphs based on professional document paragraphs and image description text; and the ability to generate summaries based on several expanded document paragraphs.

4. The method according to claim 2 or 3, characterized in that, By comparing user-inputted information with preset themes in the local database, the desired theme for generating the presentation document is determined, including: Identify the professional field and topic type specified in the user's actively input information for generating the presentation document; Calculate the second similarity between the professional field and topic type and the preset topics in the local database; Topics with a second similarity greater than the second threshold are identified as the user-expected topics for generating the demonstration document.

5. The method according to claim 2 or 3, characterized in that, Retrieve preset subtopics under the user's expected topic from the local database, generate user questions based on these subtopics, and engage in dialogue with the user through these questions to determine several user-expected subtopics for generating the presentation document. These subtopics include: Retrieve the preset presentation document layout template and retrieve all preset sub-themes under the user's desired theme from the local database; Generate user questions that include asking the user to select and modify the presentation document layout template, the total number of pages the user expects in the generated presentation document, the selected sub-themes, and the number of sub-pages for each sub-theme; By engaging with users through questions and conversations, we can obtain the user-specified first presentation document layout template, the total number of pages the user expects to generate in the presentation document, several user-expected subtopics selected from the subtopics, and the number of subpages corresponding to each user-expected subtopic.

6. The method according to claim 5, characterized in that, Based on the user's desired theme and several desired sub-themes, generate the main title and several sub-titles of the presentation document, specifically including: Filter content data in the local database based on the user's desired sub-topic and corresponding subpage number. The content data consists of professional images, professional charts, and professional document paragraphs. Based on the first, second, and third keywords of the selected content data, combined with the user's desired sub-topics, sub-titles of the presentation document are generated; Generate the main title of the presentation document based on several subheadings and the user's desired theme.

7. The method according to claim 5, characterized in that, Combine the main title and main content, and subtitles and subcontent to generate a presentation document, specifically including: Fill the main title and main content, and subtitles and subcontents into the first presentation document layout template and display it to the user. Receive user confirmation or modification instructions for the displayed content to generate the presentation document.

8. A presentation document generation device, characterized in that, The device includes: The topic determination unit is used to compare the information actively entered by the user with the preset topics in the local database to determine the user's expected topic for generating the presentation document; The subtopic determination unit, connected to the topic determination unit, is used to obtain preset subtopics under the user's expected topic in the local database, generate user questions based on the subtopics, and engage in dialogue with the user through the user questions to determine several user-expected subtopics for generating the presentation document. The title generation unit, connected to the subtopic determination unit, is used to generate the main title and several subtitles of the presentation document based on the user's desired theme and several user-desired subtopics. The content generation unit, connected to the title generation unit, is used to obtain content data corresponding to each user's desired sub-topic from the local database, generate sub-content for each sub-title based on the corresponding content data, and generate the main content of the main title based on all sub-content. Specifically, this includes: The content acquisition sub-unit is used to retrieve professional images, charts, and document paragraphs filtered from the local database, as well as the first keyword added to the professional images and the second keyword extracted from the professional charts in the professional documents and generated by combining the context of the professional documents. The image-to-text subunit, connected to the content acquisition subunit, is used to input selected professional images and charts, along with their first and second keywords, into the image-to-text model to generate image description text. The expanded text subunit, connected to the image-to-text subunit, is used to input the generated image description text and filtered professional document paragraphs into the large language model to generate expanded document paragraphs. The sub-content generation sub-unit, connected to the expansion sub-unit, is used to generate sub-content for each subheading based on the expanded document paragraphs, professional images, and professional charts. The main content generation subunit, connected to the sub-content generation subunit, is used to input the expanded document paragraphs of all sub-contents into the large language model to generate a summary, in order to generate the main content of the main title; The document generation unit, connected to the content generation unit, is used to combine the main title and main content, and the subtitles and subcontent, to generate a presentation document.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the demonstration document generation method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Presentation file generation method and device, equipment and medium

    CN111753108A

  • Generation method, device and equipment of presentation file and medium

    CN118657856A