Presentation file generation method and device based on large model, electronic equipment and medium
Through a large-model-based method, combining the topics entered by users and the nested relationship tree of presentation templates, the large language model is used to generate text content, which solves the problem that automatic generation of presentations in the existing technology is difficult to adapt to complex layout structures, and achieves higher layout diversity and professionalism.
Patent Information
- Application Number
- CN202510121781.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
AI Technical Summary
When the prior art automatically generates presentations, it is difficult to adapt to templates with complex layout structures, resulting in insufficient diversity in the layout of the generated presentations.
Using a large model-based method, by receiving the topics input by users, selecting matching presentation templates, obtaining the nested relationship tree of the template, and using the large language model to generate text content that matches the topic, fill it in the text box in the template, and generate a presentation.
It realizes automatic generation of presentation templates suitable for complex layout structures, improving the layout diversity and professionalism of presentations.
Smart Images

Figure CN120031015A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the field of smart office and large model technology. Background Art
[0002] Presentations (PowerPoint, PPT) can dynamically display static content. Therefore, using presentations can make complex issues easier to understand, and the content display form is more vivid, leaving a deeper impression on the viewers. Summary of the invention
[0003] The present disclosure provides a method, device, electronic device and medium for generating a presentation based on a large model.
[0004] A first aspect of the embodiments of the present disclosure provides a method for generating a presentation based on a large model, comprising:
[0005] Receive the topic input by the user;
[0006] Select a target presentation template matching the theme from among a plurality of presentation templates;
[0007] Acquire a nested relationship tree of the target presentation template, the nested relationship tree comprising a plurality of nodes and logical nested relationships between the nodes, each node representing a text box included in the target presentation template;
[0008] Based on the subject and the nested relationship tree, using a large language model to generate text content that conforms to the subject for the text box included in the target presentation template;
[0009] Fill the text content into the text box included in the target presentation template to obtain a presentation.
[0010] A second aspect of the embodiments of the present disclosure provides a large model-based presentation document generation device, including:
[0011] A receiving module, used for receiving a topic input by a user;
[0012] A selection module, configured to select a target presentation template matching the theme received by the receiving module from a plurality of presentation templates;
[0013] An acquisition module, used for acquiring a nested relationship tree of the target presentation template selected by the selection module, wherein the nested relationship tree includes a plurality of nodes and logical nested relationships between the nodes, and each node represents a text box included in the target presentation template;
[0014] A generating module, configured to generate text content that conforms to the subject for the text box included in the target presentation template using a large language model based on the subject received by the receiving module and the nested relationship tree acquired by the acquiring module;
[0015] A construction module is used to fill the text content generated by the generation module into the text box included in the target presentation template selected by the selection module to obtain a presentation.
[0016] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method according to the first aspect.
[0020] According to a fourth aspect of an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute a method according to any one of the first aspects.
[0021] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method according to any one of the first aspects is implemented.
[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0024] Figure 1 It is a flow chart of a method for generating a presentation based on a large model provided in an embodiment of the present disclosure;
[0025] Figure 2 is an exemplary schematic diagram of a presentation provided by an embodiment of the present disclosure;
[0026] Figure 3 is an exemplary schematic diagram of another presentation provided by an embodiment of the present disclosure;
[0027] Figure 4is a flow chart of a method for generating a nested relationship tree provided by an embodiment of the present disclosure;
[0028] Figure 5 is an exemplary schematic diagram of a nested relationship tree provided by an embodiment of the present disclosure;
[0029] Figure 6 is a flow chart of a method for selecting a presentation template provided by an embodiment of the present disclosure;
[0030] Figure 7 is a flow chart of another method for generating a presentation based on a large model provided by an embodiment of the present disclosure;
[0031] Figure 8 It is a structural schematic diagram of a large model-based presentation document generation device provided in an embodiment of the present disclosure;
[0032] Fig. 9 It is a block diagram of an electronic device used to implement the large model-based presentation generation method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0034] At present, in the traditional method of automatically generating a presentation, after obtaining the subject input by the user, a text matching the subject is generated, and then a template matching the text is selected and the text is filled into the template to generate a presentation.
[0035] However, since the template content is not considered when generating text, this method can only be adapted to templates with simple text box layout structures, and thus generate presentations with simple layout structures.
[0036] In order to improve the diversity of the layout structure of the automatically generated presentation, the embodiment of the present disclosure provides a presentation generation method based on a large model, which is applied to electronic devices, such as servers, desktop computers, laptop computers, tablet computers, etc., which have presentation processing capabilities. Figure 1 As shown, the large model-based presentation generation method provided by the embodiment of the present disclosure includes the following steps:
[0037] S101: Receive a topic input by a user.
[0038] The subject input by the user is a piece of text, which is used to express the central idea, main point or content outline of the presentation that the user wants to generate. For example, the subject is: the team building plan for the first quarter of a certain company.
[0039] It should be noted that when the electronic device is a server, a user terminal such as a mobile phone or a tablet computer can receive a topic input by a user, and then the server can receive the topic sent by the user terminal. When the electronic device is a user terminal, the terminal can directly receive the topic input by the user.
[0040] S102. Select a target presentation template that matches the theme from multiple presentation templates.
[0041] Each presentation template is a template with a relatively complex layout structure that is pre-designed for one or more scenarios. A target presentation template that matches the theme can be determined based on the scenarios that each presentation template is adapted to.
[0042] S103: Obtain a nested relationship tree of the target presentation template.
[0043] The nested relationship tree includes multiple nodes and logical nested relationships between the nodes, and each node represents a text box included in the target presentation template.
[0044] Logical nesting relationships include: position nesting and meaning nesting. Among them, two text boxes have a position nesting relationship, which means that one text box is inside another text box. Two text boxes have a meaning nesting relationship, which means that the content meaning of one text box includes the content meaning of another text box.
[0045] For example, see Figure 2 , Figure 2 For a presentation template, Figure 2 Each dotted box represents a text box, and each text box includes a template text, which is used to prompt the text content recommended to be filled in the text box. Figure 2 The template text included in text box 1 is "main title", and the template text included in text box 2 is "subtitle". Text box 2 is in text box 1, so there is a position nesting relationship between text box 2 and text box 1, that is, text box 1 includes text box 2.
[0046] For example, see Figure 3 , Figure 3 For a presentation template, Figure 3 Each dotted box represents a text box, and each text box includes a section of template text. Figure 3The template text included in text box 3 is "project / activity planning", the template text included in text box 4 is "project / activity goals", the template text included in text box 5 is "task and time management", and the template text included in text box 6 is "resource and budget management". The meanings of the template texts included in text boxes 4 to 6 all belong to the meanings of the template text included in text box 3, so there is a nested relationship between the meanings of text boxes 4 to 6 and text box 3, that is, text box 3 includes text boxes 4 to 6.
[0047] S104. Based on the topic and the nested relationship tree, the large language model is used to generate text content that matches the topic for the text box included in the target presentation template.
[0048] The Large Language Model (LLM) is a model obtained by training a deep learning model using a large amount of text data. For example, the deep learning model can be a transformer.
[0049] S105. Fill the text content into the text box included in the target presentation template to obtain a presentation.
[0050] In addition to the text content generated for each text box, the output results of the large language model may also include text box information corresponding to each text content, such as the text box information being a text box identifier or a text box position, etc., to distinguish text boxes corresponding to different text contents, so that the electronic device can fill each text content into its corresponding text box.
[0051] For example, the output result of the large language model is: "001: The goal of this activity is to enhance team cohesion; 002: Activities include mountain climbing and boating; 003: The activity time is from 9 am to 5 pm." Among them, 001, 002 and 003 are respectively marked with a text box, and the text after ":" represents the generated text content.
[0052] It should be noted that after S105, when the electronic device is a server, the server may also send the presentation obtained in S105 to the user terminal so that the user terminal displays the presentation. When the electronic device is a user terminal, the user terminal may directly display the presentation.
[0053] The disclosed embodiment can select a target presentation template that matches the subject input by the user, and then based on the nested relationship tree between the subject and the target presentation template, use the large language model to generate text content that matches the subject for the text box included in the target presentation template, and then fill the text content into the text box included in the target presentation template to obtain a presentation. Since the nested relationship tree includes multiple nodes and logical nested relationships between the nodes, and each node represents a text box included in the target presentation template, the large language model can combine the nested relationships between the text boxes to generate more accurate text content that matches the presentation structure and the subject input by the user for each text box. Therefore, the disclosed embodiment can be applied to templates with more complex text box layout structures, thereby generating presentations with more diverse text layout structures.
[0054] The following is a detailed description of the large model-based presentation generation method provided in an embodiment of the present disclosure.
[0055] Before automatically generating a presentation, the embodiment of the present disclosure can pre-generate a nested relationship tree of the presentation template. Figure 4 As shown, the method of generating a nested relationship tree includes the following steps:
[0056] S401. For each presentation template, extract the property information of each text box from the code of the presentation template.
[0057] In the disclosed embodiment, the attribute information includes at least one of the following: a text box identifier (ID), a size, a position in a presentation template, and a template text in the text box.
[0058] The size of the text box can be expressed as: length × width, with the unit being pixel (px).
[0059] The position of the text box in the presentation template can be expressed as: the coordinates of at least one key point of the text box in the presentation template, in px. For example, the key point is the center point, the upper left corner point or the lower right corner point of the text box.
[0060] The template text in the text box means: the text content recommended to be filled in the text box.
[0061] In addition, the attribute information may also include other parameters, such as border color, background color, and template text font, etc., which are not specifically limited in the embodiments of the present disclosure.
[0062] The disclosed embodiment can obtain multi-dimensional attribute information such as text box identification, size, position in the presentation template, and template text in the text box, based on which the accuracy of determining the nested relationship tree can be improved.
[0063] Since in the code of the presentation template, all information of the text box included in the presentation template is in the structure of JavaScript Object Notation (JSON), that is, all information is recorded in the form of key-value pairs, the electronic device can obtain the value corresponding to the preset key from the code of the presentation template as the attribute information of the text box.
[0064] S402: Based on the attribute information of each text box and the image of the presentation template, a nested relationship tree of the presentation template is generated using a multimodal large language model.
[0065] The image of the presentation template can be generated by converting the presentation template into an image format, or by taking a screenshot of the presentation template, or by other means. The electronic device can generate the image of the presentation template; or the electronic device can receive the image of the presentation template input by other devices or staff.
[0066] Among them, the Multimodal Large Language Model (MLLM) is a model based on a large language model that has natural language processing and computer vision processing capabilities. It can understand and analyze data in various forms such as text and pictures.
[0067] Exemplarily, the nested relationship tree structure is:
[0068] {
[0069] text:aaa,
[0070] id:001,
[0071] box:(x,y),aa×bb,
[0072] next:[{secondary node 1},{secondary node 2},…,{secondary node n}]
[0073] }
[0074] Among them, text, id, box and next are the information of the first-level node. "text:aaa" means that the template text of the text box corresponding to the first-level node is aaa. "id:001" means that the identifier of the text box corresponding to the first-level node is 001. "box:(x,y),aa×bb" means that the center point coordinates of the text box corresponding to the first-level node are (x,y), and the size of the text box is aa×bb. Next means the second-level node included in the first-level node. The structure of the second-level node is the same as that of the first-level node, which will not be repeated here.
[0075] See also Figure 5 , the logical structure of the nested relationship tree is as follows Figure 5 As shown, a first-level node may include one or more second-level nodes, each second-level node may not include or include at least one third-level node, each third-level node may not include or include at least one fourth-level node, and so on. Moreover, each node may include attribute information such as text, id, and box.
[0076] The parallel or progressive logical relationship between text boxes can be clearly reflected by using the nested relationship tree structure.
[0077] Since the attribute information of the text box can represent multi-dimensional information such as the content and size of the text box, and the picture of the presentation can reflect the layout of each text box, a more accurate nested relationship tree can be generated by combining the attribute information of each text box and the picture of the presentation template. Moreover, compared with the method of manually constructing a nested relationship tree, the embodiment of the present disclosure can automatically construct a nested relationship tree, thereby improving the efficiency of constructing the nested relationship tree.
[0078] In some embodiments of the present disclosure, the above S402 generates a nested relationship tree of the presentation template using a multimodal large language model based on the attribute information of each text box and the image of the presentation template, including the following steps:
[0079] Step 1: Based on the attribute information of each text box and the image of the presentation template, fill in the first prompt template to obtain the first prompt information. The first prompt template may be a prompt template, and the first prompt information may be called prompt.
[0080] The first prompt information is used to indicate that the logical nested relationship between the text boxes is identified based on the attribute information and the picture, and a nested relationship tree is generated.
[0081] For example, the first prompt template is:
[0082] “Task description: Based on the attribute information of each text box included in the PPT and the PPT picture, generate a tree representing the nested relationship between the text boxes.
[0083] Attribute information of each text box:
[0084] PPT pictures:
[0085] ”
[0086] After filling in the first prompt template, the first prompt information obtained is:
[0087] “Task description: Based on the attribute information of each text box included in the PPT and the PPT picture, generate a tree representing the nested relationship between the text boxes.
[0088] Attribute information of each text box: the logo, size, position in the PPT and template text in the text box of text box 1; the logo, size, position in the PPT and template text in the text box of text box 2.
[0089] PPT picture: [Picture]
[0090] ”
[0091] Here, “[Picture]” indicates the picture of the presentation template, which is a real picture in the actual application scenario.
[0092] Step 2: Input the first prompt information into the multimodal large language model to obtain a nested relationship tree of the presentation template output by the multimodal large language model.
[0093] Since the multimodal large language model has powerful reasoning and generalization capabilities, the multimodal large language model can be used to more accurately understand the nested relationship between text boxes, thereby improving the accuracy of the nested relationship tree obtained in the embodiment of the present disclosure.
[0094] After generating the nested relationship tree of each presentation template, you can use Figure 1 The process shown automatically generates the presentation.
[0095] See also Figure 6 , the above Figure 1 The method of S102 for selecting a target presentation template matching a theme from a plurality of presentation templates includes the following steps:
[0096] S601. For each presentation template, based on the theme and the template text in the text box included in the target presentation template, determine the matching degree between the theme and the presentation template.
[0097] Since the template text in the text box included in the target presentation template can represent the recommended text content to be filled in the text box, thereby reflecting the scenario to which the presentation template is adapted, based on the theme and the template text in the text box included in the target presentation template, it is possible to determine whether the theme matches the scenario to which the presentation template is adapted, thereby obtaining the degree of matching.
[0098] S602: Select the presentation template with the highest matching degree as the target presentation template.
[0099] The template text in the disclosed embodiment can reflect the scenario to which the presentation template is adapted, so based on the template text in the text box included in the subject and the target presentation template, the degree of matching between the subject and the presentation template can be determined more accurately. Moreover, selecting the presentation template with the highest matching degree as the target presentation template can increase the probability of matching the target presentation template with the subject, so that the generated presentation better meets the user's needs and improves the user experience.
[0100] In the above S601, when determining the matching degree between the subject and the presentation template, it can be based on the template text in each text box included in the target presentation template, or based on the template text in some text boxes included in the target presentation template.
[0101] For example, the template text in the text box represented by the highest-level node can be obtained from the nested relationship tree of the presentation template as the first-level template text, and then the similarity between the subject and the first-level template text can be determined to obtain the matching degree between the subject and the presentation template.
[0102] Among them, the node with the highest level is the first-level node in the nested relationship tree.
[0103] As the node level is higher, the meaning of the template text in the text box corresponding to the node is more general, and the meaning of the template text in the text box corresponding to the node is more specific as the node level is lower. Therefore, the first-level template text can better reflect the scene adapted by the presentation template, so based on the first-level template text, the accuracy of the matching degree between the determined theme and the presentation template can be improved.
[0104] In the disclosed embodiment, the above-mentioned method of determining the similarity between the subject and the primary template text includes the following two methods.
[0105] Similarity determination method 1: perform vector conversion on the topic to obtain the topic vector, and perform vector conversion on the primary template text to obtain the template vector. Then determine the cosine similarity between the topic vector and the template vector.
[0106] You can use the preset vector conversion method to convert the topic and the first-level template text into vectors. For example, the preset vector conversion method is: One-hot encoding, bag-of-words model, word to vector (Word2Vec), or Bidirectional Encoder Representations from Transformers (BERT).
[0107] Since the subject and the primary template text are both text structures and difficult to compare directly, the embodiment of the present application performs vector conversion on the subject and the primary template text respectively, thereby quantizing the text into vectors, thereby facilitating the calculation of the similarity between texts and improving the accuracy and efficiency of determining the similarity.
[0108] Similarity determination method 2: Generate a third-level title based on the theme. Then perform vector conversion on the third-level title to obtain the title vector, and perform vector conversion on the first-level template text to obtain the template vector. Then determine the cosine similarity between the title vector and the template vector to obtain the matching degree between the theme and the presentation template.
[0109] The electronic device can input the topic into a pre-trained title generation model to obtain a three-level title output by the title generation model. The title generation model can be a large language model or other models with text processing capabilities.
[0110] The third-level titles include document title, chapter title and page title. For example, the theme is: Team Building Plan for the First Quarter of a Company; the third-level titles generated based on the theme are: Team Building Plan - Basic Information - Personnel and Time, Team Building Plan - Activity Arrangement - Activity Project 1, Team Building Plan - Activity Arrangement - Activity Project 2. Each third-level title corresponds to a presentation page.
[0111] Optionally, after receiving the subject input by the user, multiple third-level headings can be generated based on the subject, and the target presentation template can be determined based on each third-level heading, and a one-page presentation can be generated. Alternatively, after receiving the subject input by the user, multiple third-level headings can be generated based on the subject, and multiple third-level headings can be displayed. Whenever it is detected that the user selects a third-level heading, the target presentation template is determined for the third-level heading selected by the user, and a one-page presentation can be generated.
[0112] In the disclosed embodiment, a third-level title can be generated based on the subject input by the user. Compared with the subject, the third-level title contains richer content and describes the presentation requirements more specifically. Therefore, based on the third-level title, a more accurate target presentation template can be obtained.
[0113] In actual application scenarios, the coverage of applicable scenarios by various presentation templates may not be wide enough, or the topic input by the user may be too special, which may result in that various presentation templates are not suitable for generating presentations on this topic.
[0114] Therefore, in order to improve the accuracy of the generated presentation, before selecting the presentation template with the highest matching degree as the target presentation template in the above S602, the embodiment of the present disclosure can also determine whether the highest matching degree is greater than the preset matching degree among the matching degrees between the theme and each presentation template. The preset matching degree can be determined according to actual business needs. For example, when the matching degree value range is [0,1], the preset matching degree can be set to 0.8.
[0115] If so, the step of selecting the presentation template with the highest matching degree as the target presentation template is performed.
[0116] When the highest matching degree is greater than the preset matching degree, it means that the presentation template corresponding to the highest matching degree is suitable for generating a presentation on the subject. Therefore, the presentation template can be used as the target presentation template, thereby improving the adaptability of the generated presentation to the subject input by the user and improving the user experience.
[0117] On the other hand, if the highest matching degree is less than or equal to the preset matching degree, a structured text that conforms to the theme is generated, and then a target simple template that matches the structured text is selected from each simple template, and the structured text is filled into the text box included in the target simple template to obtain a presentation.
[0118] Among them, the simple template is a template for generating presentations without a nested relationship tree. It can be seen that the simple template is a PPT template with a simple layout structure, for example, a template in which there is no position nesting relationship and / or meaning nesting relationship between text boxes.
[0119] The topic can be input into a pre-trained structured text generation model to obtain a structured text in a preset format output by the structured text generation model. The structured text generation model can be a large language model or other models with text processing capabilities.
[0120] For example, the preset format may be: lightweight markup language (Markdown) or JSON and other formats.
[0121] When the highest matching degree is less than or equal to the preset matching degree, it means that the presentation template corresponding to the highest matching degree is not suitable for generating a presentation of the theme. Compared with a presentation template with a complex layout structure, since the layout structure of a simple template is simpler, it is more universal. Therefore, the embodiment of the present disclosure generates a presentation based on a simple template at this time, which can improve the compatibility of the generated presentation with the theme input by the user and improve the user experience.
[0122] In some embodiments of the present disclosure, the above S104 generates text content that conforms to the theme for the text box included in the target presentation template based on the theme and the nested relationship tree using the large language model, including the following steps:
[0123] Step 1: Based on the subject and the nested relationship tree, fill in the second prompt template to obtain the second prompt information. The second prompt template may be a prompt template, and the second prompt information may be called prompt.
[0124] The second prompt information is used to indicate that text content that meets the theme is generated for each node in the nested relationship tree.
[0125] For example, the second prompt message is:
[0126] "Task description: You are now going to do a ppt nested relationship tree imitation task: just imitate the string corresponding to the text key in the input nested relationship tree according to the given topic, and then output it. The following is a nested relationship tree of a ppt page. In the nested relationship tree, text represents the ppt text content, box represents the position and size information of the text content on the ppt page, and next represents the list of child nodes under the node. The child nodes and parent nodes conform to a certain logical order relationship.
[0127] Nested relationship tree: {text:aaa,
[0128] id:001,
[0129] box:(x,y),aa×bb,
[0130] next:[{secondary node 1},{secondary node 2},…,{secondary node n}]}
[0131] According to the structure and logical relationship of the above nested relationship tree, the string corresponding to the text key is rewritten into the text content of a PPT page titled "Team Building Plan for the First Quarter of a Certain Company".
[0132] Step 2: Input the second prompt information into the large language model to obtain the text content output by the large language model.
[0133] If the topic is directly input into the large language model to make the large language model generate the text content within the text box only based on the topic, it may cause the text content to be unable to adapt to the presentation template with a professional layout structure. Therefore, in the embodiments of the present disclosure, the second prompt information is generated based on the topic and the nested relationship tree, and the second prompt information is input into the large language model, so that the large language model can generate text content according to the nested relationship between the text boxes in the presentation template. Therefore, the generated text content can adapt to the professional layout structure, enabling the embodiments of the present disclosure to generate more professional, more logical, more aesthetic, and more diverse presentation documents.
[0134] It should be noted that after generating a page of the presentation document in S105 above, the style of the presentation document can be set to a specified style, where the specified style is the style of the generated presentation document pages in the document to which this page of the presentation document belongs, so as to keep the styles of the presentation document pages in the same document consistent. Among them, the style includes the page background image, the border color of the text box, and the font color of the text content, etc.
[0135] See Figure 7 , the overall process of the presentation document generation method based on the large model provided by the embodiments of the present disclosure will be described below in combination with the actual application scenario:
[0136] The presentation document generation method based on the large model includes two parts. The first part is the extraction of professional layout logical relationships, and the second part is the generation of professional layout text content.
[0137] In the extraction of professional layout logical relationships, for each presentation template, the attribute information such as the text box identifier, size, position in the presentation template, and template text within the text box of each text box is extracted from the code of the presentation template; then, based on the attribute information of each text box and the image of the presentation template, the first prompt template is filled to obtain the first prompt information, and the first prompt information is input into the multi-modal large language model to obtain the nested relationship tree of the presentation template output by the multi-modal large language model, and the presentation template and its nested relationship tree are stored in the ppt professional layout template library.
[0138] In the generation of professional layout text content, the subject input by the user is received, and then it is determined whether the subject recall is successful, that is, for the presentation templates in the ppt professional layout template library, the first-level template text of each presentation template is vectorized to obtain the template vector, and the subject is vectorized to obtain the subject vector, and the cosine similarity between the subject vector and each template vector is determined to obtain the matching degree between the subject and each presentation template; it is determined whether the highest matching degree is greater than the preset matching degree, and if so, it means that the subject recall is successful, and the presentation template with the highest matching degree is selected as the target presentation template. Then, based on the nested relationship tree between the subject and the target presentation template, the second prompt information is constructed, and the second prompt information is input into the large language model to obtain the text content that conforms to the subject generated by the large language model for each text box included in the target presentation template, and then each text content is filled into the text box included in the target presentation template to obtain a page of presentation. If not, it means that the subject recall fails, then a structured text that matches the subject is generated, and then a target simple template that matches the structured text is determined from each simple template, and the structured text is filled into the text box included in the target simple template to obtain a page of presentation.
[0139] It can be seen that the embodiment of the present disclosure provides an end-to-end presentation generation method, that is, a user can obtain an automatically generated presentation with a professional layout structure by inputting a topic. The embodiment of the present disclosure not only improves the layout diversity of the automatically generated presentation, but also improves the professionalism and aesthetics of the presentation layout.
[0140] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of the subject matter and presentation information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0141] It should be noted that the large language model in this embodiment is not a model for a specific user and cannot reflect the personal information of a specific user.
[0142] It should be noted that the presentation template in this embodiment may come from a public data set.
[0143] Based on the same inventive concept, corresponding to the above method embodiment, the embodiment of the present disclosure also provides a presentation document generation device based on a large model, such as Figure 8 As shown, the device includes: a receiving module 801, a selecting module 802, an acquiring module 803, a generating module 804 and a constructing module 805;
[0144] Receiving module 801, used for receiving a subject input by a user;
[0145] A selection module 802 is used to select a target presentation template matching the theme received by the receiving module 801 from multiple presentation templates;
[0146] An acquisition module 803 is used to acquire a nested relationship tree of the target presentation template selected by the selection module 802, wherein the nested relationship tree includes a plurality of nodes and logical nested relationships between the nodes, and each node represents a text box included in the target presentation template;
[0147] A generating module 804 is used to generate text content that meets the theme for the text box included in the target presentation template using a large language model based on the theme received by the receiving module 801 and the nested relationship tree obtained by the obtaining module 803;
[0148] The construction module 805 is used to fill the text content generated by the generation module 804 into the text box included in the target presentation template selected by the selection module 802 to obtain a presentation.
[0149] In some embodiments of the present disclosure, the device further comprises:
[0150] An extraction module is used to extract the attribute information of each text box from the code of each presentation template before receiving the subject input by the user;
[0151] The generation module 804 is further used to generate a nested relationship tree of the presentation template using a multimodal large language model based on the attribute information of each text box and the image of the presentation template.
[0152] In some embodiments of the present disclosure, the generating module 804 is specifically configured to:
[0153] Based on the attribute information of each text box and the picture of the presentation template, fill in the first prompt template to obtain first prompt information, where the first prompt information is used to indicate that the logical nested relationship between the text boxes is identified based on the attribute information and the picture, and a nested relationship tree is generated;
[0154] The first prompt information is input into the multimodal large language model to obtain a nested relationship tree of the presentation template output by the multimodal large language model.
[0155] In some embodiments of the present disclosure, the attribute information includes at least one of the following: a text box identifier, a size, a position in a presentation template, and a template text in the text box.
[0156] In some embodiments of the present disclosure, the generating module 804 is specifically configured to:
[0157] Based on the subject and the nested relationship tree, fill in a second prompt template to obtain second prompt information, where the second prompt information is used to indicate that text content that meets the subject is generated for each node in the nested relationship tree;
[0158] The second prompt information is input into the large language model to obtain text content output by the large language model.
[0159] In some embodiments of the present disclosure, the selection module 802 is specifically configured to:
[0160] For each presentation template, based on the theme and the template text in the text box included in the target presentation template, determine the matching degree between the theme and the presentation template;
[0161] Select the presentation template that best matches you as the target presentation template.
[0162] In some embodiments of the present disclosure, the selection module 802 is specifically configured to:
[0163] From the nested relationship tree of the presentation template, obtain the template text in the text box represented by the node with the highest level as the first-level template text;
[0164] The similarity between the subject and the primary template text is determined to obtain the matching degree between the subject and the presentation template.
[0165] In some embodiments of the present disclosure, the selection module 802 is specifically configured to:
[0166] Perform vector transformation on the topic to obtain the topic vector;
[0167] Perform vector conversion on the first-level template text to obtain a template vector;
[0168] Determine the cosine similarity between the subject vector and the template vector.
[0169] In some embodiments of the present disclosure, the selection module 802 is specifically configured to:
[0170] Generate three-level headings based on the topic, including document title, chapter title and page title;
[0171] Perform vector conversion on the third-level title to obtain the title vector;
[0172] Perform vector conversion on the first-level template text to obtain a template vector;
[0173] Determine the cosine similarity between the title vector and the template vector.
[0174] In some embodiments of the present disclosure, the device further comprises:
[0175] A judgment module, used for determining whether the highest matching degree is greater than a preset matching degree among the matching degrees between the theme and each presentation template before selecting the presentation template with the highest matching degree as the target presentation template;
[0176] The calling module is used to call the selection module 802 to execute the step of selecting the presentation template with the highest matching degree as the target presentation template if the judgment result of the judgment module is yes.
[0177] In some embodiments of the present disclosure,
[0178] The generating module 804 is further used to determine whether the highest matching degree is greater than a preset matching degree among the matching degrees between the theme and each presentation template, and if not, generate a structured text that meets the theme;
[0179] The selection module 802 is further used to select a target simple template matching the structured text from the simple templates, wherein the simple template is a template for generating a presentation without a nested relationship tree;
[0180] The construction module 805 is also used to fill the structured text into the text box included in the target simple template to obtain a presentation.
[0181] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0182] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0183] like Fig. 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0184] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0185] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as a large model-based presentation generation method. For example, in some embodiments, the large model-based presentation generation method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the large model-based presentation generation method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the large model-based presentation generation method in any other appropriate manner (eg, by means of firmware).
[0186] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0187] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0188] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0189] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0190] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0191] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0192] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0193] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A presentation generation method based on a large model, comprising: Receive the topic input by the user; Select a target presentation template matching the theme from among a plurality of presentation templates; Acquire a nested relationship tree of the target presentation template, the nested relationship tree comprising a plurality of nodes and logical nested relationships between the nodes, each node representing a text box included in the target presentation template; Based on the subject and the nested relationship tree, using a large language model to generate text content that conforms to the subject for the text box included in the target presentation template; Fill the text content into the text box included in the target presentation template to obtain a presentation.
2. The method according to claim 1, before receiving the subject input by the user, further comprising: For each presentation template, extract the attribute information of each text box from the code of the presentation template; Based on the attribute information of each text box and the image of the presentation template, a nested relationship tree of the presentation template is generated using a multimodal large language model.
3. The method according to claim 2, wherein: The method of generating a nested relationship tree of the presentation template using a multimodal large language model based on the attribute information of each text box and the picture of the presentation template includes: Based on the attribute information of each text box and the picture of the presentation template, fill in the first prompt template to obtain first prompt information, wherein the first prompt information is used to indicate that the logical nested relationship between each text box is identified based on the attribute information and the picture, and a nested relationship tree is generated; The first prompt information is input into the multimodal large language model to obtain a nested relationship tree of the presentation template output by the multimodal large language model.
4. The method according to claim 2 or 3, wherein: The attribute information includes at least one of the following: a text box identifier, a size, a position in a presentation template, and a template text in the text box.
5. The method according to any one of claims 1 to 4, wherein: The step of generating text content that conforms to the theme for the text box included in the target presentation template based on the theme and the nested relationship tree using a large language model includes: Based on the subject and the nested relationship tree, fill in a second prompt template to obtain second prompt information, wherein the second prompt information is used to indicate that text content that conforms to the subject is generated for each node in the nested relationship tree; The second prompt information is input into the large language model to obtain the text content output by the large language model.
6. The method according to any one of claims 1 to 5, wherein: The step of selecting a target presentation template matching the theme from a plurality of presentation templates comprises: For each presentation template, based on the subject and the template text in the text box included in the target presentation template, determine the matching degree between the subject and the presentation template; The presentation template with the highest matching degree is selected as the target presentation template.
7. The method according to claim 6, wherein: The determining the degree of matching between the subject and the presentation template based on the subject and the template text in the text box included in the target presentation template includes: From the nested relationship tree of the presentation template, obtain the template text in the text box represented by the node with the highest level as the first-level template text; The similarity between the subject and the primary template text is determined to obtain a matching degree between the subject and the presentation template.
8. The method according to claim 7, wherein: Determining the similarity between the subject and the primary template text includes: Performing vector conversion on the subject to obtain a subject vector; Performing vector conversion on the primary template text to obtain a template vector; A cosine similarity between the subject vector and the template vector is determined.
9. The method according to claim 7, wherein: Determining the similarity between the subject and the primary template text includes: Based on the subject, generate a third-level title, wherein the third-level title includes a document title, a chapter title, and a page title; Performing vector conversion on the third-level title to obtain a title vector; Performing vector conversion on the primary template text to obtain a template vector; A cosine similarity between the title vector and the template vector is determined.
10. The method according to any one of claims 6 to 9, before selecting the presentation template with the highest matching degree as the target presentation template, further comprising: Among the matching degrees between the subject and each presentation template, determining whether the highest matching degree is greater than a preset matching degree; If so, the step of selecting the presentation template with the highest matching degree as the target presentation template is performed.
11. The method according to claim 10, after determining whether the highest matching degree among the matching degrees between the subject and each presentation template is greater than a preset matching degree, further comprises: If not, generate structured text that matches the topic; Selecting a target simple template matching the structured text from the simple templates, wherein the simple template is a template for generating a presentation without a nested relationship tree; Fill the structured text into the text box included in the target simple template to obtain a presentation.
12. A presentation document generation device based on a large model, comprising: A receiving module, used for receiving a topic input by a user; A selection module, configured to select a target presentation template matching the theme received by the receiving module from a plurality of presentation templates; An acquisition module, used for acquiring a nested relationship tree of the target presentation template selected by the selection module, wherein the nested relationship tree includes a plurality of nodes and logical nested relationships between the nodes, and each node represents a text box included in the target presentation template; A generating module, configured to generate text content that conforms to the subject for the text box included in the target presentation template using a large language model based on the subject received by the receiving module and the nested relationship tree acquired by the acquiring module; A construction module is used to fill the text content generated by the generation module into the text box included in the target presentation template selected by the selection module to obtain a presentation.
13. The apparatus according to claim 12, further comprising: An extraction module, used for extracting the attribute information of each text box from the code of each presentation template for each presentation template before receiving the subject input by the user; The generation module is also used to generate a nested relationship tree of the presentation template using a multimodal large language model based on the attribute information of each text box and the image of the presentation template.
14. The device according to claim 13, wherein: The generation module is specifically used for: Based on the attribute information of each text box and the picture of the presentation template, fill in the first prompt template to obtain first prompt information, wherein the first prompt information is used to indicate that the logical nested relationship between each text box is identified based on the attribute information and the picture, and a nested relationship tree is generated; The first prompt information is input into the multimodal large language model to obtain a nested relationship tree of the presentation template output by the multimodal large language model.
15. The device according to claim 13 or 14, wherein: The attribute information includes at least one of the following: a text box identifier, a size, a position in a presentation template, and a template text in the text box.
16. The device according to any one of claims 12 to 15, wherein: The generation module is specifically used for: Based on the subject and the nested relationship tree, fill in a second prompt template to obtain second prompt information, wherein the second prompt information is used to indicate that text content that conforms to the subject is generated for each node in the nested relationship tree; The second prompt information is input into the large language model to obtain the text content output by the large language model.
17. The device according to any one of claims 12 to 16, wherein: The selection module is specifically used for: For each presentation template, based on the subject and the template text in the text box included in the target presentation template, determine the matching degree between the subject and the presentation template; The presentation template with the highest matching degree is selected as the target presentation template.
18. The device according to claim 17, wherein: The selection module is specifically used for: From the nested relationship tree of the presentation template, obtain the template text in the text box represented by the node with the highest level as the first-level template text; The similarity between the subject and the primary template text is determined to obtain a matching degree between the subject and the presentation template.
19. The device according to claim 18, wherein The selection module is specifically used for: Performing vector conversion on the subject to obtain a subject vector; Performing vector conversion on the primary template text to obtain a template vector; A cosine similarity between the subject vector and the template vector is determined.
20. The device according to claim 18, wherein The selection module is specifically used for: Based on the subject, generate a third-level title, wherein the third-level title includes a document title, a chapter title, and a page title; Performing vector conversion on the third-level title to obtain a title vector; Performing vector conversion on the primary template text to obtain a template vector; A cosine similarity between the title vector and the template vector is determined.
21. The device according to any one of claims 17 to 20, further comprising: A judgment module, used for determining whether the highest matching degree among the matching degrees between the subject and each presentation template is greater than a preset matching degree before selecting the presentation template with the highest matching degree as the target presentation template; The calling module is used to call the selecting module to execute the step of selecting the presentation template with the highest matching degree as the target presentation template if the judgment result of the judging module is yes.
22. The device according to claim 21, The generating module is further configured to determine whether the highest matching degree among the matching degrees between the subject and each presentation template is greater than a preset matching degree, and if not, generate a structured text that conforms to the subject; The selection module is further used to select a target simple template matching the structured text from the simple templates, wherein the simple template is a template for generating a presentation without a nested relationship tree; The construction module is also used to fill the structured text into the text box included in the target simple template to obtain a presentation.
23. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
25. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Cited By
Credible traceability presentation file generation method and system based on document content
CN120493894A
A method and system for generating a trusted traceability presentation based on document content
CN120493894B