Automatic Generation of Demonstration Document Training Method, Device and System

By training multimodal large models in stages, gradually generating demonstration documents, solving the problem of high generation complexity and improving the generation quality and stability.

CN119990067BActive Publication Date: 2025-07-18HANGZHOU SHAOKE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510465839.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the prior art, the generation of presentation documents is high in complexity, resulting in poor generation quality.

Method used

The multimodal large model is gradually optimized by using a staged training method, and a demonstration document containing only black and white text, a color demonstration document with text style, and a single-page complete demonstration document are generated, and the next stage model is initialized using the final weight of the previous stage.

Benefits of technology

Reduces the complexity of the generation task and improves the stability and quality of the generation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990067B_ABST
    Figure CN119990067B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, and system for training to automatically generate a presentation document. Among them, the method for training to automatically generate a presentation document includes: in the first-stage training, in response to a generation instruction, a training model generates an HTML corresponding to a single-page presentation document containing only black-and-white text according to a Prompt text; in the second-stage training, the training model generates an HTML corresponding to a single-page color presentation document with text styles; in the third-stage training, the training model generates an HTML corresponding to a single-page presentation document; in the fourth-stage training, the training model generates an HTML corresponding to a multi-page presentation document, and constructs a presentation document according to the HTML corresponding to the multi-page presentation document. Compared with directly training a model to generate a multi-page presentation document, the present application can reduce the complexity of the model learning to automatically generate a presentation document through a phased and progressive training method, thereby generating a presentation document with stable quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning, and particularly to a method, device, and system for training to automatically generate presentation documents. Background Art

[0002] Presentation documents are an effective communication medium that can not only improve the effect of information transmission but also enhance the persuasiveness of the speaker. However, manually creating presentation documents requires a great deal of time and effort. In the era of artificial intelligence, it is already possible to automatically generate presentation documents based on the text provided by the user.

[0003] In the prior art, generative adversarial networks and variational autoencoders have been applied to image generation for presentation documents. However, due to the overly complex task of generating presentation documents and poor model convergence, the quality of the generated presentation documents is not good. Therefore, how to reduce the complexity of generating presentation documents and improve the quality of the generated presentation documents has become a key problem to be solved urgently. Summary of the Invention

[0004] Embodiments of this application provide a method, device, and system for training to automatically generate presentation documents to solve the problem in the related art of how to reduce the complexity of generating presentation documents and improve the quality of the generated presentation documents.

[0005] In a first aspect, embodiments of this application provide a method for training to automatically generate presentation documents. The method is applied to a multimodal large model and includes:

[0006] Obtain a reference presentation document, and based on the reference presentation document, obtain a training data set;

[0007] In the first-stage training, in response to a generation instruction, train the multimodal large model to generate HTML corresponding to a single-page presentation document with only black and white text based on the training data set, and save the final weights of the first stage as the initial weights of the second stage;

[0008] In the second-stage training, the multimodal large model is initialized by loading the initial weights of the second stage. Based on the training data set, train the model to generate HTML corresponding to a single-page colored presentation document with text styles, and save the final weights of the second stage as the initial weights of the third stage;

[0009] In the third-stage training, based on the initial weights of the third stage and the training data set, train the multimodal large model to generate HTML corresponding to a single-page presentation document, and save the final weights of the third stage as the initial weights of the fourth stage;

[0010] In the fourth-stage training, based on the initial weights of the fourth stage and the training data set, train the multimodal large model to generate HTML corresponding to a multi-page presentation document, and construct a presentation document based on the HTML corresponding to the multi-page presentation document.

[0011] In one embodiment, according to a reference demonstration document and a multimodal large model, a training dataset is obtained, including:

[0012] Taking the page image of the reference demonstration document as the input of the multimodal large model, obtaining the Prompt text and the corresponding HTML converted from the page image of the reference demonstration document;

[0013] Generating training data for four stages based on the HTML;

[0014] Adjusting the Prompt text according to the task requirements of each stage to generate Prompt texts for four stages;

[0015] Pairing the Prompt text and the HTML of the same stage to form a training dataset.

[0016] In one embodiment, for the HTML, generating training data for four stages, including:

[0017] Processing the HTML to obtain the complete HTML, and taking the complete HTML as the HTML training data for the fourth stage;

[0018] Splitting the HTML training data of the fourth stage, retaining the complete image and text information, obtaining the HTML content corresponding to a single-page demonstration document, and saving the HTML content corresponding to the single-page demonstration document as the HTML training data for the third stage;

[0019] Removing the image elements from the HTML content of the third stage, retaining only the text content, obtaining the HTML corresponding to the text single-page demonstration document, and saving the HTML with text styles corresponding to the single-page color demonstration document as the HTML training data for the second stage;

[0020] Removing the text styles from the HTML of the second stage, retaining only the black and white text content, obtaining the HTML corresponding to the black and white text single-page demonstration document, and saving the HTML corresponding to the black and white text single-page demonstration document as the HTML training data for the first stage.

[0021] In one embodiment, in response to a generation instruction, training the multimodal large model to generate the HTML corresponding to a single-page demonstration document with only black and white text according to the training dataset, and saving the final weight of the first stage as the initial weight of the second stage, including:

[0022] In response to the generation instruction, obtaining the Prompt text of the first stage;

[0023] Obtaining a random initial weight, and initializing the multimodal large model based on the initial weight;

[0024] Based on the HTML training data in the first stage and the Prompt text in the first stage, train the multimodal large model to generate HTML corresponding to a demonstration document containing only black and white text, and save the final training weights in the first stage as the initial weights in the second stage for initializing the weights of the multimodal large model in the second stage.

[0025] In one embodiment, in the second-stage training, the multimodal large model loads the initial weights in the second stage for initialization, and trains the model to generate HTML corresponding to a single-page color demonstration document with text styles according to the training dataset, and saves the final weights in the second stage as the initial weights in the third stage, including:

[0026] In response to the generation instruction, obtain the Prompt text in the second stage;

[0027] The multimodal large model loads the initial weights in the second stage and initializes according to the initial weights in the second stage;

[0028] According to the HTML training data in the second stage and the Prompt text in the second stage, train the multimodal large model to generate HTML corresponding to a single-page color demonstration document with text styles, and save the final training weights in the second stage as the initial weights in the third stage for initializing the weights of the multimodal large model in the third stage.

[0029] In one embodiment, based on the initial weights in the third stage and the training dataset, train the multimodal large model to generate HTML corresponding to a single-page demonstration document, and save the final weights in the third stage as the initial weights in the fourth stage, including:

[0030] In response to the generation instruction, obtain the Prompt text in the third stage;

[0031] The multimodal large model loads the initial weights in the third stage and initializes according to the initial weights in the third stage;

[0032] According to the HTML training data in the third stage and the Prompt text in the third stage, train the initialized multimodal large model to generate HTML corresponding to a single-page demonstration document, and use the final training weights in the third stage as the initial weights in the fourth stage for initializing the weights of the multimodal large model in the fourth stage.

[0033] In one embodiment, according to the initial weights in the fourth stage and the training dataset, train the multimodal large model to generate HTML corresponding to a multi-page demonstration document, and construct a demonstration document based on the HTML corresponding to the multi-page demonstration document, including:

[0034] In response to the generation instruction, obtain the Prompt text in the fourth stage;

[0035] Initialize the multi-modal large model by loading the initial weights in the fourth stage;

[0036] According to the HTML training data in the fourth stage and the Prompt text in the fourth stage, train the initialized multi-modal large model to generate the HTML corresponding to the multi-page presentation document, parse the HTML corresponding to the multi-page presentation document, determine the layout and content of each page of the presentation document according to the parsing result, and obtain the multi-page presentation document. In a second aspect, an embodiment of the present application provides an automatic presentation document generation training device, including:

[0037] A training dataset acquisition module, configured to acquire a reference presentation document and, according to the reference presentation document, acquire a training dataset;

[0038] A first-stage training module, configured to, in the first-stage training, in response to a generation instruction, train the multi-modal large model to generate the HTML corresponding to a single-page presentation document containing only black and white text according to the training dataset, and save the final weights of the first stage as the initial weights of the second stage;

[0039] A second-stage training module, configured to, in the second-stage training, initialize the multi-modal large model by loading the initial weights of the second stage, and train the model to generate the HTML corresponding to a single-page color presentation document with text styles according to the training dataset, and save the final weights of the second stage as the initial weights of the third stage;

[0040] A third-stage training module, configured to, in the third-stage training, based on the initial weights of the third stage and the training dataset, train the multi-modal large model to generate the HTML corresponding to a single-page presentation document, and save the final weights of the third stage as the initial weights of the fourth stage;

[0041] A fourth-stage training module, configured to, in the fourth-stage training, train the multi-modal large model to generate the HTML corresponding to a multi-page presentation document according to the initial weights of the fourth stage and the training dataset, and construct a presentation document according to the HTML corresponding to the multi-page presentation document.

[0042] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the automatic presentation document generation training method as described in the first aspect above.

[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the automatic presentation document generation training method as described in the first aspect above.

[0044] The automatic generation of a demonstration document training method, apparatus, and system provided by the embodiments of the present application at least have the following technical effects.

[0045] By dividing the demonstration document generation task into multiple stages and gradually optimizing based on different training data, the model sequentially generates the HTML corresponding to the demonstration document containing only black and white text, the HTML corresponding to the colored demonstration document with text styles, the HTML corresponding to the single-page complete demonstration document, and finally generates the HTML corresponding to the multi-page demonstration document. The final weights trained in the previous stage are used to initialize the model weights of the next stage to achieve progressive learning. This phased progressive training for generating demonstration documents effectively reduces the complexity of the generation task. Compared with the method of directly training to generate multi-page demonstration documents, the present application generates content through phased learning, thereby reducing the learning complexity of the model and improving the stability and quality of the generation results.

[0046] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0048] Figure 1 is a flowchart of an automatic generation of a demonstration document training method shown according to an exemplary embodiment;

[0049] Figure 2 is a process of an automatic generation of a demonstration document training method shown according to an exemplary embodiment;

[0050] Figure 3 is a block diagram of an automatic generation of a demonstration document training apparatus shown according to an exemplary embodiment;

[0051] Figure 4 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided by the present application without creative efforts fall within the scope of protection of the present application.

[0053] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.

[0054] In the present application, the mention of "embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0055] Unless otherwise defined, the technical terms or scientific terms involved in the present application should have the ordinary meaning understood by those of ordinary skill in the technical field to which the present application belongs. The words such as "a", "an", "one kind", "the" and the like involved in the present application do not represent a quantity limitation and can represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0056] In a first aspect, an embodiment of the present application provides an automatic generation method for training a presentation document. Figure 1 It is a flowchart of an automatic generation method for training a presentation document shown according to an exemplary example, as Figure 1As shown in the figure, the method for training an automatically generated demonstration document includes:

[0057] Step S101: Obtain a reference demonstration document, and based on the reference demonstration document, obtain a training data set.

[0058] The reference demonstration document is a complete demonstration document, including a cover page, a table of contents page, content, and an ending page, and there is specific text content on any page. The reference demonstration document also needs to meet the following requirements: the content is complete, the expression is clear, and the style is unified, so as to ensure the quality of the generated demonstration document.

[0059] In deep learning, the initial weight is the process of assigning an initial value to the weight, which is the starting point of deep learning. Through the training and learning of the model, the weight will also be continuously optimized, enabling the model to obtain better data fitting.

[0060] The training data set includes training data for each stage. Among them, obtaining the training data set specifically includes:

[0061] Step S100: Use the page image of the reference demonstration document as the input of the multi-modal large model to obtain the Prompt text and the corresponding HTML converted from the page image of the reference demonstration document.

[0062] Step S200: Based on the HTML, generate training data for four stages.

[0063] Step S300: According to the task requirements of each stage, adjust the Prompt text to generate Prompt texts for four stages.

[0064] Step S400: Pair the Prompt text and HTML of the same stage to form a training data set.

[0065] Input the page image of the reference demonstration document into the multi-modal large model to obtain the HTML and the Prompt text. And process the obtained HTML to obtain the training data for four training stages. Since the processing tasks of each stage are different, according to the task requirements of each stage, adjust the obtained Prompt text to obtain the Prompt texts for four stages. Pair the Prompt text and HTML of the same stage to form the training data set for each stage.

[0066] Regarding obtaining the training data for each stage, it specifically includes:

[0067] Step S201: Process the HTML to obtain the complete HTML, and use the complete HTML as the HTML training data for the fourth stage.

[0068] Format the HTML to obtain all the elements in the reference demonstration document. Among them, all the elements include: the original image, text content, and layout. And obtain the HTML of all the elements, and use the HTML containing all the elements as the HTML training data for the fourth stage.

[0069] Step S202: Split the HTML training data for the fourth stage, retain the complete image and text information, obtain the HTML content corresponding to a single-page demonstration document, and save the HTML content corresponding to the single-page demonstration document as the HTML training data for the third stage.

[0070] All the elements are obtained based on a complete demonstration document, specifically including the content of each page of the demonstration document. Split all the elements to obtain the HTML corresponding to each page of the demonstration document, and save the HTML content corresponding to the single-page demonstration document as the HTML training data for the third stage.

[0071] Step S203: Remove the image elements in the HTML content of the third stage, only retain the text content, obtain the HTML corresponding to the text single-page demonstration document, and save the HTML corresponding to the single-page color demonstration document with text styles as the HTML training data for the second stage.

[0072] The single-page elements include the original image and text content. In the demonstration document, the text content not only has black and white text, but also contains various text styles, which can play the role of emphasizing key content and improving readability. Therefore, the text content includes text and text styles. Among them, the text styles include font style, font color, and font size.

[0073] Obtaining the HTML training data for the second stage includes: removing the original image in the single-page elements, retaining the text content, and obtaining the HTLM data of the text content, and using the HTML corresponding to the demonstration document containing the text content as the training data for the second stage.

[0074] Step S204: Remove the text styles in the HTML of the second stage, only retain the black and white text content, obtain the HTML corresponding to the black and white text single-page demonstration document, and save the HTML corresponding to the black and white text single-page demonstration document as the HTML training data for the first stage.

[0075] Remove the text styles in the text content to obtain black and white text elements, which are also the initial text. And obtain the HTML of the black and white text elements, and use the HTML of the black and white text elements as the HTML training data for the first stage.

[0076] By obtaining the training data for each stage through the above steps S201 to S202, it is possible to systematically generate phased training data and provide effective support for the subsequent training of multi-modal large models.

[0077] Continuing to refer to steps S100 to S400, after data processing, obtain the Prompt text and HTML, pair them according to the same stage, and form the training dataset for each stage. When generating demonstration documents in subsequent phases, call different training data for generating demonstration documents in different phases to improve the accuracy of generating demonstration documents.

[0078] Step S102, in the first-stage training, in response to the generation instruction, train the multi-modal large model to generate the HTML corresponding to a single-page demonstration document containing only black and white text, and save the final weights of the first stage as the initial weights of the second stage.

[0079] During the process of generating the demonstration document, before obtaining the final demonstration document, all the obtained multi-type demonstration documents are displayed in HTML.

[0080] According to the Prompt text in the training dataset, train the model to generate the HTML of a single-page black and white text demonstration document, specifically including:

[0081] Step S121, in response to the generation instruction, obtain the Prompt text of the first stage.

[0082] In the first stage, in response to the generation instruction, the Prompt text includes the generation instruction and the prompt. Among them, the prompt is: "Please generate the HTML corresponding to a single-page demonstration document containing only black and white text according to the generation instruction." For example, if the generation instruction is "Help me generate the table of contents of a demonstration document on the introduction to Transformer, which includes four parts. The first part is the uses of Transformer, the second part is the basic structure of Transformer, the third part is the advantages of Transformer, and the fourth part is the deficiencies of Transformer.", then the Prompt text input to the multi-modal large model is: "Generation instruction: Help me generate the table of contents of a demonstration document on the introduction to Transformer, which includes four parts. The first part is the uses of Transformer, the second part is the basic structure of Transformer, the third part is the advantages of Transformer, and the fourth part is the deficiencies of Transformer. Please generate the HTML corresponding to a single-page demonstration document containing only black and white text according to the generation instruction."

[0083] Step S122, obtain the initial weights, and the multi-modal large model is initialized based on the initial weights.

[0084] Randomly initialize the weights to obtain the initial weights. The multi-modal large model randomly initializes the weights in the first-stage training to prepare for training the multi-modal large model to generate the HTML corresponding to a single-page black-and-white text presentation document.

[0085] Step S123: According to the HTML training data in the second stage and the Prompt text in the first stage, train the multi-modal large model to generate the HTML corresponding to a single-page color presentation document with text styles, and save the final training weights in the first stage as the initial weights in the second stage for initializing the weights of the multi-modal large model in the second stage.

[0086] By learning the HTML training data in the second stage, the multi-modal large model can generate the HTML corresponding to a single-page black-and-white text presentation document according to the Prompt text in the first stage.

[0087] In the process of generating the HTML corresponding to the single-page black-and-white text presentation document in the first stage, save the final trained weights in the first stage as the initial weights in the second stage, and transfer the initial weights in the second stage to the next generation stage, that is, the second stage of generating the HTML corresponding to the single-page color presentation document with text styles.

[0088] In step S102, through learning the HTML training data in the first stage, the multi-modal large model has the ability to generate the HTML corresponding to a presentation document with black-and-white text. And transfer the initial weights in the second stage to the second stage, and continue the ability to generate a presentation document with black-and-white text to the next stage, so that the second stage can start training from a higher starting point, thereby more effectively generating a presentation document with text styles.

[0089] Step S103: In the second-stage training, the multi-modal large model loads the initial weights in the second stage for initialization, and trains the model to generate the HTML corresponding to a single-page color presentation document with text styles according to the training dataset, and saves the final weights in the second stage as the initial weights in the third stage.

[0090] In the first stage of generating a presentation document with black-and-white text, the multi-modal large model already has the ability to generate the HTML corresponding to a presentation document that only contains black-and-white text. Therefore, in the stage of generating a single-page presentation document with text styles, use the final weights in the first stage as the initialization weights, and combine the Prompt text containing "text styles" and the HTML of text styles as the training data to further train the model to generate the HTML corresponding to a single-page color presentation document with text styles. Among them, the text styles include font styles (such as boldface, regular script), font colors, and font sizes.

[0091] The steps for training a multi-modal large model to generate a single-page color presentation document with text styles specifically include the following steps:

[0092] Step S131: In response to a generation instruction, obtain the second-stage Prompt text.

[0093] In the second stage, in response to a generation instruction, the Prompt text includes the generation instruction and a prompt. In the stage of generating a single-page color presentation document with text styles, the Prompt text input to the model includes an instruction and a prompt. The prompt is "Please generate the HTML corresponding to a single-page color presentation document without images but with text styles according to the generation instruction."

[0094] In one embodiment, the second-stage Prompt text specifically includes: If the generation instruction is "Help me generate a table of contents for a presentation document on the introduction to Transformers, which includes four parts. The first part is the uses of Transformers, the second part is the basic structure of Transformers, the third part is the advantages of Transformers, and the fourth part is the disadvantages of Transformers.", then the Prompt text input to the multi-modal large model is: "Generation instruction: Help me generate a table of contents for a presentation document on the introduction to Transformers, which includes four parts. The first part is the uses of Transformers, the second part is the basic structure of Transformers, the third part is the advantages of Transformers, and the fourth part is the disadvantages of Transformers. Please generate the HTML corresponding to a single-page color presentation document without images but with text styles according to the generation instruction."

[0095] Step S132: The multi-modal large model loads the initial weights of the second stage and initializes according to the initial weights of the second stage.

[0096] The multi-modal large model loads the initial weights of the second stage and initializes based on the initial weights of the second stage to prepare for training the multi-modal large model to generate the HTML corresponding to a single-page color presentation document with text styles.

[0097] Step S133: According to the second-stage HTML training data and the second-stage Prompt text, train the multi-modal large model to generate the HTML corresponding to a single-page color presentation document with text styles, and save the final training weights of the second stage as the initial weights of the third stage for initializing the weights of the multi-modal large model in the third stage.

[0098] The multi-modal large model learns from the HTML training data in the second stage to train the model to generate the HTML corresponding to a single-page color presentation document with text styles according to the Prompt text in the second stage.

[0099] In the second stage of generating the HTML corresponding to a single-page color presentation document with text styles, the weights finally trained in the second stage are saved as the initial weights in the third stage, and the initial weights in the third stage are passed to the next generation stage, that is, the third stage of generating the HTML corresponding to a single-page presentation document.

[0100] Continuing to refer to step S103, the second stage is a further optimization based on the first stage. That is, the stage of generating a single-page text-style presentation document is a further optimization based on the stage of generating a single-page black-and-white text presentation document. Compared with the method of directly generating a presentation document containing text styles, the learning difficulty of the model in generating a presentation document containing text styles is significantly reduced on the premise that the model already has the ability to generate a black-and-white text presentation document. By initializing with the weights of the first stage, the model can learn based on its existing capabilities, avoid repeated learning, and effectively reduce training time and resource consumption. In addition, this phased training strategy enables the model to be more focused on learning style features when generating HTML containing text styles, thereby further reducing the overall training difficulty.

[0101] Step S104: In the third stage of training, based on the initial weights and training data set in the third stage, train the multi-modal large model to generate the HTML corresponding to a single-page presentation document, and save the final weights in the third stage as the initial weights in the fourth stage.

[0102] The multi-modal large model can process and understand the relationships between different modalities. Therefore, in the stage of generating a single-page presentation document, the multi-modal large model can also understand and process the image information added to the HTML. The specific steps for generating a single-page presentation document are as follows:

[0103] Step S141: In response to the generation instruction, obtain the Prompt text in the third stage.

[0104] In the third stage, in response to the generation instruction, the Prompt text includes the generation instruction and a prompt. In the stage of generating a single-page presentation document, the Prompt text input to the model includes an instruction and a prompt. The prompt is "Please generate the HTML corresponding to a single-page presentation document according to the generation instruction."

[0105] In one embodiment, in the third stage, that is, the stage of generating a single-page presentation document, the Prompt text input to the multimodal large model is "Generation instruction: Help me generate a table of contents for a presentation document on the introduction to Transformer, which includes four parts. The first part is the uses of Transformer, the second part is the basic structure of Transformer, the third part is the advantages of Transformer, and the fourth part is the deficiencies of Transformer. Please generate the HTML corresponding to the single-page presentation document according to the generation instruction. Please generate the HTML corresponding to the single-page presentation document according to the generation instruction", then the output data of the multimodal large model is the corresponding HTML described in the Prompt text.

[0106] When the multimodal large model receives the presentation document page image as input and outputs HTML, the model represents the image features through quantization coding and maps these special encodings to special tokens. These tokens are embedded into the HTML generated by the model to identify the corresponding image content. Therefore, to obtain the HTML training data containing images, the HTML generated by the model must be processed. First, parse out the special tokens identifying the images from the generated HTML; then, use the image quantization coding model to decode the special tokens to generate the actual image data; finally, insert these images into the corresponding image tag positions in the HTML.

[0107] Step S142: The multimodal large model loads the initial weights of the third stage and initializes according to the initial weights of the third stage.

[0108] The multimodal large model loads the initial weights of the third stage and initializes based on the initial weights of the third stage to prepare for training the multimodal large model to generate the HTML corresponding to the single-page presentation document.

[0109] Step S143: According to the HTML training data of the third stage, train the model to generate the HTML corresponding to the single-page presentation document and save the final training weights of the third stage as the initial weights of the fourth stage for initializing the model weights of the fourth stage.

[0110] Based on the initial weights of the third stage and the HTML training data of the third stage, train the multimodal large model to generate the HTML corresponding to the single-page presentation document according to the Prompt text of the third stage. The HTML of the single-page presentation document includes the content of the complete single-page presentation document, such as text, images, and page layout, making the generated HTML closer to the presentation form of the actual application presentation document.

[0111] Since the multimodal large model can receive both text and images as input. The multimodal large model extracts image features through the visual encoder and combines the HTML training data of the third stage to learn to generate HTML files containing images. During the training process, the multimodal large model continuously optimizes the weights by learning the HTML training data of the third stage, so that the generated HTML is closer to the style of the real presentation document. The final training weights of the third stage are saved as the initial weights of the fourth stage, and the initial weights of the fourth stage are passed to the next generation stage, that is, the fourth stage of generating the HTML corresponding to the multi-page presentation document.

[0112] In one embodiment, an image of a table of contents page of a presentation document is input into the multimodal model, and the prompt output by the model is described as follows: This is a table of contents page of a presentation document, the background color is mainly orange, with white Chinese "目录" and English "Contents" on the left. There are six orange rectangular boxes on the right, each with a white serial number and text, and the specific content is as follows:

[0113] 01 Transformer Overview

[0114] 02 Self-Attention Mechanism

[0115] 03 Encoder Detailed Explanation

[0116] 04 Decoder Details

[0117] 05 Positional Encoding

[0118] 06 Loss Function and Optimization

[0119] Continuing to refer to step S104, the third stage is further optimized based on the HTML generated in the second stage, that is, the single-page presentation document generation stage is further optimized based on the HTML corresponding to the generated single-page color presentation document with text style. Before the third stage training begins, the model already has the ability to generate the HTML corresponding to the single-page presentation document with text style, so the difficulty of learning to generate the HTML corresponding to the complete single-page presentation document is significantly reduced.

[0120] Step S105: In the fourth stage of training, the training model generates HTML corresponding to the multi-page presentation document according to the initial weights of the fourth stage and the training data set, and constructs a presentation document according to the HTML corresponding to the multi-page presentation document.

[0121] The specific steps to obtain a multi-page presentation document are as follows:

[0122] Step S151, in response to the generation instruction, obtaining the Prompt text of the fourth stage;

[0123] In the fourth stage, in response to the generation instruction, the Prompt text includes the generation instruction and a prompt. Among them, the prompt is: "Please generate the HTML corresponding to the multi-page presentation document according to the generation instruction." For example, if the generation instruction is "Help me generate the table of contents of a presentation document for an introductory explanation of Transformers, which includes four parts. The first part is the uses of Transformers, the second part is the basic structure of Transformers, the third part is the advantages of Transformers, and the fourth part is the deficiencies of Transformers.", then the Prompt text input to the multi-modal large model is: "Generation instruction: Help me generate the table of contents of a presentation document for an introductory explanation of Transformers, which includes four parts. The first part is the uses of Transformers, the second part is the basic structure of Transformers, the third part is the advantages of Transformers, and the fourth part is the deficiencies of Transformers. Please generate the HTML corresponding to the multi-page presentation document according to the generation instruction."

[0124] Step S152: The multi-modal large model loads the initial weights of the fourth stage for initialization.

[0125] The multi-modal large model loads the initial weights of the fourth stage for initialization to prepare for training to generate the HTML corresponding to the multi-page presentation document.

[0126] Step S153: According to the HTML training data of the fourth stage and the Prompt text of the fourth stage, train the model to generate the HTML corresponding to the multi-page presentation document, parse the HTML corresponding to the multi-page presentation document, and determine the layout and content of each page of the presentation document according to the parsing result to obtain the multi-page presentation document.

[0127] Based on the initial weights of the fourth stage and the HTML training data of the fourth stage, train the model to generate the corresponding HTML of the multi-page presentation document according to the Prompt text of the fourth stage.

[0128] The HTML training data of the fourth stage is generated by inputting the reference presentation document into the multi-modal large model. Therefore, the HTML training data of the fourth stage includes the structure and content information of the multi-page presentation document. In the fourth stage, the model is further trained on the basis of the previous training to learn the ability to generate multi-page presentation documents. By training on the basis of the final weights in the third stage, it carries the learning results of the previous stage, thus accelerating the training process.

[0129] Parse the HTML of the multi-page presentation document and construct the presentation document page by page according to the parsed content. The constructed content includes the layout and specific content of each page of the presentation document to form the final multi-page presentation document.

[0130] In another embodiment, Figure 2 is a flowchart of a method for training an automatically generated presentation document according to an exemplary embodiment, as Figure 2 shown. Step 1: Screen open-source presentation documents, convert the page images of the presentation documents into text descriptions through a multimodal large model, and convert the page images of the presentation documents into HTML to construct a dataset. Step 2: In the first stage, generate the HTML corresponding to a single-page presentation document containing only black and white text, and use the final weights trained in this stage for the initialization of the next stage. Step 3: In the second stage, generate the HTML corresponding to a single-page color presentation document without images, and use the final weights trained in this stage for the initialization of the next stage. Step 4: In the third stage, generate the HTML corresponding to a single-page presentation document containing all elements, and apply the final weights trained in this stage to the initialization of the next stage. Step 5: In the fourth stage, train the model to generate the HTML corresponding to a multi-page presentation document.

[0131] In summary, the method for training an automatically generated presentation document provided by the embodiments of the present application, through a phased training strategy, selects corresponding training data according to the generation tasks specific to each stage. Starting from the stage of generating text styles, each stage continues to optimize on the basis of the training results of the previous stage, thus avoiding independent and repeated learning in each stage. Compared with direct training, this step-by-step learning method effectively reduces the content that the model needs to learn, reduces the learning difficulty of the model, and the quality of the generated presentation documents is more stable.

[0132] In a second aspect, an embodiment of the present application provides an apparatus for training an automatically generated presentation document. Figure 3 is a block diagram of an apparatus for training an automatically generated presentation document according to an exemplary embodiment. As Figure 3 shown, the apparatus for training an automatically generated presentation document includes:

[0133] A training dataset acquisition module, configured to acquire a reference presentation document and, according to the reference presentation document, acquire a training dataset;

[0134] A first-stage training module, configured to, in the first-stage training, in response to a generation instruction, train the model to generate the HTML corresponding to a single-page presentation document containing only black and white text according to the Prompt text, and save the final weights of the first stage as the initial weights of the second stage;

[0135] A second-stage training module, configured to, in the second-stage training, initialize the model by loading the initial weights of the second stage, train the model to generate the HTML corresponding to a single-page color presentation document with text styles according to the Prompt text, and save the final weights of the second stage as the initial weights of the third stage;

[0136] The third-stage training module is used to train the multi-modal large model to generate the HTML corresponding to the single-page presentation document based on the initial weights of the third stage and the training data set in the third-stage training, and save the final weights of the third stage as the initial weights of the fourth stage;

[0137] The fourth-stage training module is used to train the model to generate the HTML corresponding to the multi-page presentation document according to the initial weights of the fourth stage and the training data set in the fourth-stage training, and construct the presentation document based on the HTML corresponding to the multi-page presentation document.

[0138] In summary, the automatic presentation document generation training device provided in this application divides the generation task in the generation instruction into multiple stages, and gradually optimizes based on different training data. The model sequentially generates the HTML corresponding to the presentation document containing only black and white text, the HTML corresponding to the presentation document with text styles, the HTML corresponding to the single-page complete presentation document, and finally generates the HTML corresponding to the multi-page presentation document. The final weights trained in the previous stage are used to initialize the model weights of the next stage in each stage to achieve progressive learning. This phased progressive training for generating presentation documents effectively reduces the complexity of the generation task. Compared with the method of directly training to generate multi-page presentation documents, this application generates content through phased learning, thereby significantly reducing the learning complexity of the model and improving the stability and quality of the generation results.

[0139] It should be noted that the automatic presentation document generation training device provided in this embodiment is used to implement the above implementation manners, and those that have been described will not be repeated. As used above, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the above embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0140] In a third aspect, an embodiment of the present application provides an electronic device, Figure 4 which is a block diagram of an electronic device shown according to an exemplary embodiment. As Figure 4 shown, the electronic device may include a processor 81 and a memory 82 storing computer program instructions.

[0141] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.

[0142] Among them, the memory 82 may include a mass storage for data or instructions. By way of example and not limitation, the memory 82 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 82 may include removable or non-removable (or fixed) media. In a suitable case, the memory 82 may be internal or external to the data processing device. In a particular embodiment, the memory 82 is a non-volatile memory. In a particular embodiment, the memory 82 includes a read-only memory (ROM) and a random access memory (RAM). In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory (FLASH), or a combination of two or more of these. In a suitable case, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0143] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0144] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the automatic generation of presentation document training methods in the above embodiments.

[0145] In one embodiment, the automatic generation of presentation document training device may further include a communication interface 83 and a bus 80. Among them, as Figure 4 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.

[0146] The communication interface 83 is used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application. The communication interface 83 can also implement data communication with other components, such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.

[0147] Bus 80 includes hardware, software, or both, and couples the components of the demonstration document training device automatically generated together with each other. Bus 80 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, Bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In a suitable case, Bus 80 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0148] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the automatic generation demonstration document training method provided in the first aspect is implemented.

[0149] Among them, more specifically, the readable storage medium may include, but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, a wipeable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0150] In a possible implementation manner, the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps of implementing the automatic generation of a presentation document training method provided in the first aspect.

[0151] Among them, the program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0152] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0153] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. An automatic generation method for training demonstration documents, characterized in that, The method applied to the multimodal large model includes: Obtain a reference demonstration document, use the page images of the reference demonstration document as the input of the multimodal large model, and obtain the Prompt text and the corresponding HTML converted from the page images of the reference demonstration document; Process the HTML to obtain the complete HTML, and use the complete HTML as the HTML training data for the fourth stage; Split the HTML training data for the fourth stage, retain the complete image and text information, obtain the HTML content corresponding to the single-page demonstration document, and save the HTML content corresponding to the single-page demonstration document as the HTML training data for the third stage; Remove the image elements in the HTML content for the third stage, retain only the text content, obtain the HTML corresponding to the text single-page demonstration document, and save the HTML corresponding to the single-page color demonstration document with text styles as the HTML training data for the second stage; Remove the text styles of the HTML for the second stage, retain only the black-and-white text content, obtain the HTML corresponding to the black-and-white text single-page demonstration document, and save the HTML corresponding to the black-and-white text single-page demonstration document as the HTML training data for the first stage; Adjust the Prompt text according to the task requirements of each stage to generate the Prompt text for the four stages; Pair the Prompt text and HTML in the same stage to form a training data set; In the first-stage training, in response to the generation instruction, train the multimodal large model to generate the HTML corresponding to the single-page demonstration document with only black-and-white text according to the training data set, and save the final weight of the first stage as the initial weight of the second stage; In the second-stage training, the multimodal large model is initialized by loading the initial weight of the second stage. According to the training data set, train the model to generate the HTML corresponding to the single-page color demonstration document with text styles, and save the final weight of the second stage as the initial weight of the third stage; In the third-stage training, based on the initial weight of the third stage and the training data set, train the multimodal large model to generate the HTML corresponding to the single-page demonstration document, and save the final weight of the third stage as the initial weight of the fourth stage; In the fourth-stage training, according to the initial weight of the fourth stage and the training data set, train the multimodal large model to generate the HTML corresponding to the multi-page demonstration document, and construct a demonstration document according to the HTML corresponding to the multi-page demonstration document.

2. The automatic generation of a demonstration document training method according to claim 1, wherein The step of, in response to the generation instruction, training the multimodal large model to generate the HTML corresponding to the single-page demonstration document with only black-and-white text according to the training data set, and saving the final weight of the first stage as the initial weight of the second stage, includes: In response to the generation instruction, obtain the Prompt text for the first stage; Obtain a random initial weight, and the multimodal large model is initialized based on the initial weight; Based on the HTML training data in the first stage and the Prompt text in the first stage, train the multimodal large model to generate HTML corresponding to a black-and-white text-only presentation document, and save the final training weights in the first stage as the initial weights in the second stage for initializing the weights of the multimodal large model in the second stage.

3. The automatic generation of a demonstration document training method according to claim 1, characterized in that, In the second-stage training, the multimodal large model loads the initial weights in the second stage for initialization, and according to the training dataset, trains the model to generate HTML corresponding to a single-page color presentation document with text styles, and saves the final weights in the second stage as the initial weights in the third stage, including: In response to the generation instruction, obtain the second-stage Prompt text; The multimodal large model loads the initial weights in the second stage and initializes according to the initial weights in the second stage; According to the HTML training data in the second stage and the second-stage Prompt text, train the multimodal large model to generate HTML corresponding to a single-page color presentation document with text styles, and save the final training weights in the second stage as the initial weights in the third stage for initializing the weights of the multimodal large model in the third stage.

4. The automatic generation of a demonstration document training method according to claim 1, characterized in that Based on the initial weights in the third stage and the training dataset, train the multimodal large model to generate HTML corresponding to a single-page presentation document, and save the final weights in the third stage as the initial weights in the fourth stage, including: In response to the generation instruction, obtain the third-stage Prompt text; The multimodal large model loads the initial weights in the third stage and initializes according to the initial weights in the third stage; According to the HTML training data in the third stage and the third-stage Prompt text, train the initialized multimodal large model to generate HTML corresponding to a single-page presentation document, and use the final training weights in the third stage as the initial weights in the fourth stage for initializing the weights of the multimodal large model in the fourth stage.

5. The automatic generation of a demonstration document training method according to claim 1, wherein According to the initial weights in the fourth stage and the training dataset, train the multimodal large model to generate HTML corresponding to a multi-page presentation document, and construct a presentation document based on the HTML corresponding to the multi-page presentation document, including: In response to the generation instruction, obtain the fourth-stage Prompt text; The multimodal large model loads the initial weights in the fourth stage for initialization; According to the HTML training data in the fourth stage and the fourth-stage Prompt text, train the initialized multimodal large model to generate HTML corresponding to a multi-page presentation document, and parse the HTML corresponding to the multi-page presentation document, determine the layout and content of each page of the presentation document according to the parsing result, and obtain a multi-page presentation document.

6. An automatic demonstration document generation training device, characterized in that, Including: A module for obtaining Prompt text and HTML, which is used to obtain a reference presentation document, use the page image of the reference presentation document as the input of the multimodal large model, and obtain the Prompt text and the corresponding HTML converted from the page image of the reference presentation document; A module for obtaining four - stage HTML training data, which is used to process the HTML, obtain complete HTML, and use the complete HTML as the HTML training data for the fourth stage; A module for obtaining three - stage HTML training data, which is used to split the HTML training data of the fourth stage, retain complete image and text information, obtain the HTML content corresponding to a single - page presentation document, and save the HTML content corresponding to the single - page presentation document as the HTML training data for the third stage; A module for obtaining two - stage HTML training data, which is used to remove image elements from the HTML content of the third stage, retain only text content, obtain the HTML corresponding to a text single - page presentation document, and save the HTML corresponding to a single - page color presentation document with text styles as the HTML training data for the second stage; A module for obtaining one - stage HTML training data, which is used to remove the text styles of the HTML of the second stage, retain only black - and - white text content, obtain the HTML corresponding to a black - and - white text single - page presentation document, and save the HTML corresponding to the black - and - white text single - page presentation document as the HTML training data for the first stage; A module for generating four - stage Prompt texts, which is used to adjust the Prompt texts according to the task requirements of each stage and generate four - stage Prompt texts; A module for obtaining a training data set, which is used to pair the Prompt texts and HTML of the same stage to form a training data set; A first - stage training module, which is used in the first - stage training, in response to a generation instruction, to train a multi - modal large model to generate the HTML corresponding to a single - page presentation document with only black - and - white text according to the training data set, and save the final weights of the first stage as the initial weights of the second stage; A second - stage training module, which is used in the second - stage training, the multi - modal large model loads the initial weights of the second stage for initialization, and according to the training data set, trains the model to generate the HTML corresponding to a single - page color presentation document with text styles, and saves the final weights of the second stage as the initial weights of the third stage; A third - stage training module, which is used in the third - stage training, based on the initial weights of the third stage and the training data set, trains the multi - modal large model to generate the HTML corresponding to a single - page presentation document, and saves the final weights of the third stage as the initial weights of the fourth stage; A fourth - stage training module, which is used in the fourth - stage training, according to the initial weights of the fourth stage and the training data set, trains the multi - modal large model to generate the HTML corresponding to a multi - page presentation document, and constructs a presentation document according to the HTML corresponding to the multi - page presentation document.

7. An electronic device, characterized in that, Comprising a memory, a processor, and A computer program stored on the memory and executable on the processor, when the processor executes the computer program, it implements the automatic generation of a presentation document training method as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the automatic generation of a presentation document training method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text processing method, article generation method and text processing model training method

    CN115994522A

  • Content augmentation with machine generated content to meet content gaps during interaction with target entities

    US20230121711A1