Method for generating PPT based on large model

By obtaining information on the front end, generating PPT outlines using microservices and large models and filling algorithms in the windows virtual machine to generate PPT, the problems of insufficient generated content and monotonous templates in the existing technology are solved, and more complex and higher-quality PPT generation effects are achieved.

CN120012712APending Publication Date: 2025-05-16BEIYIN FINANCIAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510101982.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The word count and quality generated by existing AIppt products are insufficient, the structural level is low, the PPT template is simple and monotonous, the filling logic is single, and the generated PPT effect is not ideal.

Method used

Obtain input information through the front-end, use microservices and large models to generate outlines and content, and generate PPT through fill algorithms in Windows virtual machines, combine VLLM technology for inference deployment and acceleration, and use the PPT library provided by Microsoft for programming, optimize PPT templates and fill algorithms to improve generation effect.

Benefits of technology

The generated PPT content is more complex and high-quality, with stable structure, diverse templates, and rich filling logic, which reduces development costs and improves the controllability and stability of the generation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012712A_ABST
    Figure CN120012712A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating a PPT (Power Point) based on a large model. The method comprises the following steps: acquiring an input title and a file through a front end; front-end information is transmitted to the micro-service, input information of a user is transmitted to the large model through transfer of the micro-service, and an outline and content are generated; the micro-service sends the generated content to the windows virtual machine, and PPT is generated through a filling algorithm; and returning the generated PPT to a front-end page for a user to download. The effect and quality of the synthesized PPT are more stable and controllable, and the development cost of a large number of assemblies is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of PPT generation, and specifically is a method for generating PPT based on a large model. Background Art

[0002] The rise of the technology of generating PPT based on large models is the result of the rapid development of the field of artificial intelligence (AI). In particular, in recent years, the advancement of deep learning technology and the successful application of large-scale language models (such as GPT-3, BERT, etc.) have greatly promoted the research and application of natural language processing (NLP) and computer vision (CV).

[0003] The current product has two functions. The first is to generate PPT with one sentence, that is, generate PPT content through user input. Under normal circumstances, the user's input content is a specified title, such as the year-end summary of a multimodal algorithm engineer, how to effectively manage your body, etc.; the second function is to generate PPT based on the file and title provided by the user. The current input file formats supported are word, PPT, standard pdf, and non-standard pdf.

[0004] Existing AIppt products also generate PPTs based on large models, and are generally based on one-step or two-step prompts for outline construction and content generation. For example, in the two-step generation example of the year-end summary of a multimodal algorithm engineer, the outline is first generated, and then the corresponding content is generated through the outline, and finally the content is generated into the final PPT file through the PPT filling algorithm; in the one-step example, the outline and the corresponding content are generated directly; the difference between one-step and two-step is mainly reflected in the user's interaction logic. For example, in the two-step generation process, the user can change the generated outline and then generate content that better meets user needs; in the one-step generation product, the user will make modifications after the complete generation until the PPT generation is completed.

[0005] The current AIppt products have the following problems:

[0006] 1. The generated text content has few words, low quality and low structural level; this is a common problem of AIPPT products, mainly caused by the lack of ability of large models to generate long texts.

[0007] 2. The PPT template is simple and monotonous; this is also a common problem of AIPPT products, which is mainly caused by the simple text content generated by the first point and the lack of templates that support different structures.

[0008] 3. The filling logic of PPT is simple; this is also a problem caused by the above two points. The content generated by the large model does not have a complex hierarchical structure, so the PPT filling algorithm itself is relatively simple, and the filling logic for chapters, sub-chapters, and sub-nodes is relatively simple. Summary of the invention

[0009] In view of the above problems, the present invention is proposed to provide a method for generating a PPT based on a large model, which overcomes the above problems or at least partially solves the above problems.

[0010] To achieve the above object, the present invention adopts the following technical solutions:

[0011] A method for generating a PPT based on a large model, the method comprising:

[0012] Get input title and file through the front end;

[0013] Pass the front-end information to the microservice, and use the microservice as a transfer to send the user's input information to the big model and generate the outline and content;

[0014] The microservice sends the generated content to the Windows virtual machine and generates a PPT through a filling algorithm;

[0015] The generated PPT is returned to the front-end page for users to download.

[0016] Optionally, the method also includes training the large model to make the structure and content generated by the large model more complex and to produce a more complex PPT.

[0017] Optionally, the method further includes performing targeted training on the generated content to fix the document structure generated by the large model.

[0018] Optionally, the method further includes optimizing the user experience of PPT, and the optimization includes:

[0019] Dynamically adjust the text size according to the text box size, text length and font file to avoid the content exceeding the box due to excessive content generated by large models;

[0020] In the function of generating PPT by uploading files, the large model is set to dynamically adjust the number of chapters preset by the user during the training process to ensure the stability of the generated PPT structure;

[0021] In the function of generating PPTs when users upload files, the big model is set to optimize the user's document copy during the training process;

[0022] Use VLLM technology for inference deployment and acceleration.

[0023] Optionally, a large amount of randomness can be added to the filling algorithm so that a non-repeating template is used each time a PPT is generated.

[0024] Optionally, generating a PPT using a filling algorithm includes filling a PPT template, and the PPT template adopts a unified structure and processing method.

[0025] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0026] 1. The present invention can better control the generation effect by synthesizing PPT files through powerpoint in the windows virtual machine, can directly program powerpoint, and has built-in a large number of libraries provided by Microsoft, for example, adaptive scaling of text boxes and font sizes, copying specified panes to specified number of pages, and automatically generating various PPT animation effects. Compared with other current solutions for synthesizing PPT in the front end, the effect and quality of synthesizing PPT in the present invention are more stable and controllable, and the PPT library provided by Microsoft is directly called for programming, which reduces the development cost of a large number of components.

[0027] 2. The present invention conducts targeted training on the big model. For example, in controlling the complexity of PPT structure, numerical control of text box characters, standardized format control, and various PPT directions, compared with the big models used in other similar schemes, the present invention conducts targeted training on the trained big model to generate more complex and higher-quality text content, which is helpful to generate a more ideal PPT. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A flowchart of a method for generating a PPT based on a large model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0030] See also Figure 1 This embodiment provides a method for generating a PPT based on a large model, the method comprising:

[0031] S1. Get the input title and file through the front end.

[0032] S2. Pass the front-end information to the microservice, and use the microservice as a transfer to send the user's input information to the big model and generate the outline and content.

[0033] S3. Train the large model to make the structure and content generated by the large model more complex and produce a more complex PPT.

[0034] S4. Conduct targeted training on the generated content to fix the document structure generated by the large model.

[0035] Due to the randomness of the general large model, it is difficult to guarantee the success rate of PPT generation if it is generated using the general large model. Among all the large models tested by the author, only 20% of them meet the prompt conditions, and only chatgpt can generate the format that meets the second question (with a requirement for the number of words generated per day), and the accuracy rate can only be maintained at 30%. At present, the number of QA pairs for training large models is 5,000, which are manually synthesized from 100,000 data (various general large model products); secondly, for the generation of content, the problem of more complex structure is similar to the first point. Regardless of whether the prompt project or prompt words are used to standardize the output conditions and structure, the randomness of the large model will cause the generated PPT to fail. Finally, regarding filling, we have innovatively designed multiple sets of sophisticated templates to meet different needs. Each template currently has 36 pages, and each page contains different types of structures, such as a specific cover page, directory page, content page, and end page. The content page contains 1-6 sub-nodes of content, and a large amount of randomness is added to the filling algorithm, so that each generation can use non-repetitive templates as much as possible.

[0036] S5. The microservice sends the generated content to the Windows virtual machine and generates a PPT through the filling algorithm.

[0037] A large amount of randomness is added to the filling algorithm so that each time a PPT is generated, a non-repeated template is used. Generating a PPT through a filling algorithm includes filling a PPT template, and the PPT template adopts a unified structure and processing method. For example, the number of pages of all templates is the same (39 pages), where the first page of the PPT is the homepage, the second page is the directory page, the last page is the end page, the third to sixth pages are single title pages, and the content includes a sub-chapter title, a content title and corresponding content, and the seventh to twelfth chapters are double title pages, and the content includes a sub-chapter title, two content titles and two corresponding contents, and so on.

[0038] When the outline title generated by the large model is a single section with four titles, a random selection is made for the existing templates. However, in order to ensure the generation quality, try to avoid using previously selected templates when generating a single PPT.

[0039] S6. Optimize the user experience of PPT, the optimization includes:

[0040] The text size is dynamically adjusted according to the text box size, text length and font file to avoid the content exceeding the box due to excessive content generated by large models.

[0041] In the function of generating PPT by users uploading files, the large model is set to dynamically adjust the number of chapters preset by users during the training process to ensure the stability of the generated PPT structure.

[0042] In the function of generating PPT by users uploading files, the big model is set to optimize the user document copy during the training process to ensure the generation effect.

[0043] Using VLLM technology for inference deployment and acceleration can increase the speed of inference acceleration to more than twice the original speed compared to conventional deployment.

[0044] S7. Return the generated PPT to the front-end page for users to download.

[0045] This embodiment can better control the generation effect by synthesizing PPT files through powerpoint in the windows virtual machine, can directly program powerpoint, and has built-in a large number of libraries provided by Microsoft, for example, adaptive scaling of text boxes and font sizes, copying specified panes to specified number of pages, and automatically generating various PPT animation effects. Compared with other current solutions for synthesizing PPT in the front end, the effect and quality of synthesizing PPT in the present invention are more stable and controllable, and the PPT library provided by Microsoft is directly called for programming, which reduces the development cost of a large number of components.

[0046] This embodiment conducts targeted training on the big model. For example, in controlling the complexity of PPT structure, numerical control of text box characters, standardized format control, and various PPT directions, compared with the big models used in other similar solutions, the present invention can generate more complex and higher-quality text content for the trained big model, which helps to generate a more ideal PPT.

[0047] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A method for generating PPT based on a large model, characterized in that: The method comprises: Get input title and file through the front end; Pass the front-end information to the microservice, and use the microservice as a transfer to send the user's input information to the big model and generate the outline and content; The microservice sends the generated content to the Windows virtual machine and generates a PPT through a filling algorithm; The generated PPT is returned to the front-end page for users to download.

2. A method for generating PPT based on a large model as claimed in claim 1, characterized in that: The method also includes training the large model to make the structure and content generated by the large model more complex and to produce a more complex PPT.

3. A method for generating PPT based on a large model as claimed in claim 1, characterized in that: The method also includes performing targeted training on the generated content to fix the document structure generated by the large model.

4. A method for generating PPT based on a large model as claimed in claim 1, characterized in that: The method further includes optimizing the user experience of PPT, wherein the optimization includes: Dynamically adjust the text size according to the text box size, text length and font file to avoid the content exceeding the box due to excessive content generated by large models; In the function of generating PPT by uploading files, the large model is set to dynamically adjust the number of chapters preset by the user during the training process to ensure the stability of the generated PPT structure; In the function of generating PPTs when users upload files, the big model is set to optimize the user's document copy during the training process; Use VLLM technology for inference deployment and acceleration.

5. A method for generating PPT based on a large model as claimed in claim 1, characterized in that: Add a lot of randomness to the filling algorithm to ensure that each PPT generated uses a non-repeated template.

6. A method for generating PPT based on a large model as claimed in claim 1, characterized in that: Generating a PPT through a filling algorithm includes filling a PPT template, and the PPT template adopts a unified structure and processing method.