Training method, device and system for automatically generating demonstration document
Through a phased progressive training strategy, the demonstration document generation task is gradually optimized, which solves the problems of high complexity and poor quality of generated demonstration documents in the existing technology, and achieves more stable and high-quality demonstration document generation.
Patent Information
- Application Number
- CN202510465839.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In the prior art, the complexity of generating presentation documents is high, resulting in poor convergence of the model and poor quality of the generated presentation documents.
By dividing the demonstration document generation task into multiple stages and gradually optimizing based on different training data, the model generates presentation documents containing only black and white text, color presentation documents with text styles, and a single page complete presentation document, and finally generates a multi-page presentation document. Each stage uses the final weights of the previous stage to initialize the model weights of the next stage for progressive learning.
It effectively reduces the complexity of the generation task, reduces the learning complexity of the model, and improves the stability and quality of the generated results.
Smart Images

Figure CN119990067A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning, and in particular to a method, device and system for automatically generating training presentation documents. Background Art
[0002] Presentation documents are an effective communication medium that not only improves the effectiveness of information transmission, but also enhances the speaker's persuasiveness. However, manually making presentation documents requires a lot of time and energy. In the era of artificial intelligence, presentation documents can be automatically generated based on the text provided by the user.
[0003] In the prior art, generative adversarial networks and variational autoencoders are used in image generation demonstration documents. However, since the task of generating demonstration documents is too complex, the model convergence is poor, resulting in poor quality of the generated demonstration documents. Therefore, how to reduce the complexity of generating demonstration documents and improve the quality of generated demonstration documents has become a key issue that needs to be solved urgently. Summary of the invention
[0004] The embodiments of the present application provide a method, device and system for automatically generating a presentation document training to solve the problem in the related art of how to reduce the complexity of generating a presentation document and improve the quality of the generated presentation document.
[0005] In a first aspect, an embodiment of the present application provides a method for automatically generating a demonstration document training method, wherein the method is applied to a multimodal large model and includes: Obtain a reference demonstration document, and obtain a training data set based on the reference demonstration document; In the first stage of training, in response to the generation instruction, the multimodal large model is trained to generate HTML corresponding to a single-page presentation document containing only black and white text according to the training data set, and the final weight of the first stage is saved as the initial weight of the second stage; In the second stage of training, the multimodal large model loads the initial weights of the second stage for initialization. Based on the training data set, the training model generates HTML corresponding to a single-page color presentation document with text style, and saves the final weights of the second stage as the initial weights of the third stage. In the third stage of training, based on the initial weights and training data set of the third stage, the multimodal large model is trained to generate HTML corresponding to the single-page presentation document, and the final weights of the third stage are saved as the initial weights of the fourth stage; In the fourth stage of training, based on the initial weights and training data set of the fourth stage, the multimodal large model is trained to generate HTML corresponding to the multi-page presentation document, and the presentation document is constructed based on the HTML corresponding to the multi-page presentation document.
[0006] In one embodiment, obtaining a training data set according to a reference demonstration document and a multimodal large model includes: The page image of the reference demonstration document is used as the input of the multimodal large model to obtain the corresponding HTML converted from the prompt text and the page image of the reference demonstration document; Generate four stages of training data based on HTML; Adjust the prompt text according to the task requirements of each stage and generate prompt texts for four stages; The prompt text and HTML of the same stage are paired to form a training data set.
[0007] In one embodiment, in HTML, four stages of training data are generated, including: Process the HTML to obtain the complete HTML, and use the complete HTML as the HTML training data in the fourth stage; Split the HTML training data of the fourth stage, retain the complete image and text information, obtain the HTML content corresponding to the single-page presentation document, and save the HTML content corresponding to the single-page presentation document as the HTML training data of the third stage; Remove the image elements in the third-stage HTML content, retain only the text content, obtain the HTML corresponding to the text single-page presentation document, and save the HTML corresponding to the single-page color presentation document containing the text style as the second-stage HTML training data; The text style of the HTML in the second stage is removed, and only the black and white text content is retained to obtain the HTML corresponding to the black and white text single-page presentation document, and the HTML corresponding to the black and white text single-page presentation document is saved as the HTML training data of the first stage.
[0008] In one embodiment, in response to the generation instruction, the multimodal large model is trained to generate HTML corresponding to a single-page presentation document containing only black and white text according to the training data set, and the final weight of the first stage is saved as the initial weight of the second stage, including: In response to the generation instruction, the prompt text of the first stage is obtained; Obtain random initial weights, and initialize the multimodal large model based on the initial weights; Based on the HTML training data of the first stage and the Prompt text of the first stage, the multimodal large model is trained to generate HTML corresponding to the presentation document containing only black and white text, and the final training weights of the first stage are saved as the initial weights of the second stage to initialize the weights of the multimodal large model of the second stage.
[0009] In one embodiment, in the second stage of training, the multimodal large model loads the initial weights of the second stage for initialization, and the training model generates HTML corresponding to a single-page color presentation document containing a text style according to the training data set, and saves the final weights of the second stage as the initial weights of the third stage, including: In response to the generation instruction, obtaining the second stage Prompt text; The multimodal large model loads the initial weights of the second stage and is initialized according to the initial weights of the second stage; Based on the HTML training data of the second stage and the prompt text of the second stage, the multimodal large model is trained to generate the HTML corresponding to the single-page color presentation document with text style, and the final training weights of the second stage are saved as the initial weights of the third stage to initialize the weights of the multimodal large model of the third stage.
[0010] In one embodiment, based on the initial weights and training data set of the third stage, the multimodal large model is trained to generate HTML corresponding to the single-page presentation document, and the final weights of the third stage are saved as the initial weights of the fourth stage, including: In response to the generation instruction, the prompt text of the third stage is obtained; The multimodal large model loads the initial weights of the third stage and is initialized according to the initial weights of the third stage; According to the HTML training data of the third stage and the Prompt text of the third stage, the initialized multimodal large model is trained to generate HTML corresponding to the single-page presentation document, and the final training weights of the third stage are used as the initial weights of the fourth stage to initialize the weights of the multimodal large model of the fourth stage.
[0011] In one embodiment, according to the initial weights and training data set of the fourth stage, the multimodal large model is trained to generate HTML corresponding to the multi-page presentation document, and the presentation document is constructed according to the HTML corresponding to the multi-page presentation document, including: In response to the generation instruction, the prompt text of the fourth stage is obtained; The multimodal large model loads the initial weights of the fourth stage for initialization; According to the fourth stage HTML training data and the fourth stage prompt text, the initialized multimodal large model is trained to generate HTML corresponding to the multi-page presentation document, and the HTML corresponding to the multi-page presentation document is parsed, and the layout and content of each page of the presentation document are determined according to the parsing result to obtain the multi-page presentation document. In a second aspect, the embodiment of the present application provides a training device for automatically generating presentation documents, including: A training data set acquisition module is used to acquire a reference demonstration document, and acquire a training data set according to the reference demonstration document; A first-stage training module, for, in the first-stage training, in response to a generation instruction, training the multimodal large model to generate HTML corresponding to a single-page presentation document containing only black and white text according to the training data set, and saving the final weight of the first stage as the initial weight of the second stage; The second stage training module is used for loading the initial weights of the second stage into the multimodal large model for initialization in the second stage training, and generating HTML corresponding to a single-page color presentation document containing a text style based on the training data set by the training model, and saving the final weights of the second stage as the initial weights of the third stage; A third-stage training module, used for training the multimodal large model to generate HTML corresponding to a single-page presentation document based on the initial weights of the third stage and the training data set in the third-stage training, and saving the final weights of the third stage as the initial weights of the fourth stage; The fourth stage training module is used to train the multimodal large model to generate HTML corresponding to a multi-page presentation document according to the initial weights of the fourth stage and the training data set in the fourth stage training, and to construct a presentation document according to the HTML corresponding to the multi-page presentation document.
[0012] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for automatically generating presentation document training as described in the first aspect above is implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for automatically generating a presentation document training as described in the first aspect above.
[0014] The method, device and system for automatically generating presentation document training provided in the embodiments of the present application have at least the following technical effects.
[0015] By dividing the presentation document generation task into multiple stages and gradually optimizing based on different training data, the model sequentially generates HTML corresponding to a presentation document containing only black and white text, HTML corresponding to a color presentation document with text styles, and HTML corresponding to a single-page complete presentation document, and finally generates HTML corresponding to a multi-page presentation document. Each stage uses the final weights trained in the previous stage to initialize the model weights for the next stage to achieve progressive learning. This staged progressive training to generate presentation documents effectively reduces the complexity of the generation task. Compared with the method of directly training to generate multi-page presentation documents, this application generates content through staged learning, thereby reducing the learning complexity of the model and improving the stability and quality of the generation results.
[0016] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 is a flow chart of a method for automatically generating a presentation document training according to an exemplary embodiment; Figure 2 The present invention is a process of automatically generating a demonstration document training method according to an exemplary embodiment; Figure 3 is a block diagram of a training device for automatically generating presentation documents according to an exemplary embodiment; Figure 4 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0019] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.
[0020] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0021] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantity limitation, and may indicate the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0022] In a first aspect, the present application provides a method for automatically generating a presentation document training method. Figure 1 is a flowchart of a method for automatically generating a presentation document training according to an exemplary embodiment, such as Figure 1 As shown, the automatic generation of demonstration document training method includes: Step S101: Obtain a reference demonstration document, and obtain a training data set based on the reference demonstration document.
[0023] The reference presentation document is a complete presentation document, including the cover page, directory page, content, and end page, and there is specific text content on any page. The reference presentation document must also meet the following requirements: complete content, clear expression, and unified style to ensure the quality of the generated presentation document.
[0024] In deep learning, initial weights are the process of assigning initial values to weights, which is the starting point of deep learning. Through model training and learning, weights will be continuously optimized, allowing the model to better fit the data.
[0025] The training data set includes the training data of each stage, wherein obtaining the training data set specifically includes: Step S100: Use the page image of the reference presentation document as the input of the multimodal large model to obtain the corresponding HTML converted from the Prompt text and the page image of the reference presentation document.
[0026] Step S200: Generate four stages of training data based on HTML.
[0027] Step S300: Adjust the prompt text according to the task requirements of each stage, and generate prompt texts for four stages.
[0028] Step S400: Pair the Prompt text and HTML of the same stage to form a training data set.
[0029] Input the page image of the reference demonstration document into the multimodal large model to obtain HTML and prompt text. Then process the obtained HTML to obtain training data for four training stages. Since the processing tasks of each stage are different, adjust the obtained prompt text according to the task requirements of each stage to obtain prompt texts for four stages. Pair the prompt text and HTML of the same stage to form a training data set for each stage.
[0030] To obtain the training data in each stage, it includes: Step S201: Process HTML to obtain complete HTML, and use the complete HTML as HTML training data in the fourth stage.
[0031] Format the HTML to obtain all elements in the reference presentation document. All elements include: original image, text content, and layout. Obtain the HTML of all elements, and use the HTML containing all elements as the HTML training data for the fourth stage.
[0032] Step S202: split the HTML training data of the fourth stage, retain the complete image and text information, obtain the HTML content corresponding to the single-page presentation document, and save the HTML content corresponding to the single-page presentation document as the HTML training data of the third stage.
[0033] All elements are obtained based on a complete presentation document, including the content of each page of the presentation document. All elements are split to obtain the HTML corresponding to each page of the presentation document, and the HTML content corresponding to the single page presentation document is saved as the HTML training data of the third stage.
[0034] Step S203, remove the image elements in the third-stage HTML content, retain only the text content, obtain the HTML corresponding to the text single-page presentation document, and save the HTML corresponding to the single-page color presentation document containing the text style as the second-stage HTML training data.
[0035] Single-page elements include original images and text content. In presentation documents, text content includes not only black and white text, but also various text styles, which can emphasize key content and improve readability. Therefore, text content includes text content and text style. Text style includes font style, font color, and font size.
[0036] Obtaining the second-stage HTML training data includes: removing the original image in the single-page element, retaining the text content, and obtaining the HTML data of the text content, and using the HTML corresponding to the presentation document containing the text content as the second-stage training data.
[0037] Step S204, remove the text style of the HTML of the second stage, retain only the black and white text content, obtain the HTML corresponding to the black and white text single-page presentation document, and save the HTML corresponding to the black and white text single-page presentation document as the HTML training data of the first stage.
[0038] The text style in the text content is removed to obtain the black and white text element, which is also the initial text. The HTML of the black and white text element is obtained and used as the HTML training data of the first stage.
[0039] By obtaining the training data for each stage through the above steps S201 to S202, stage-by-stage training data can be systematically generated, and effective support can be provided for subsequent multimodal large model training.
[0040] Continuing to refer to step S100 to step S400, after data processing, Prompt text and HTML are obtained, and paired according to the same stage to form a training data set for each stage. When the presentation document is generated in subsequent stages, different training data are called for presentation documents generated in different stages to improve the accuracy of generating the presentation document.
[0041] Step S102, in the first stage of training, in response to the generation instruction, the multimodal large model is trained to generate HTML corresponding to a single-page presentation document containing only black and white text, and the final weight of the first stage is saved as the initial weight of the second stage.
[0042] In the process of generating the presentation document, before obtaining the final presentation document, the obtained multiple types of presentation documents are all displayed in HTML.
[0043] Based on the prompt text in the training data set, the training model generates the HTML of a single-page black-and-white text presentation document, including: Step S121, in response to the generation instruction, obtain the Prompt text of the first stage.
[0044] In the first stage, in response to the generation instruction, the Prompt text includes the generation instruction and the prompt. The prompt is: "Please generate the HTML corresponding to the presentation document containing only black and white text according to the generation instruction." For example, the generation instruction is "Help me generate a presentation document directory for the introductory explanation of Transformer, which contains four parts. The first part is the purpose of Transformer, the second part is the basic structure of Transformer, the third part is the advantages of Transformer, and the fourth part is the shortcomings of Transformer." Then the Prompt text for inputting the multimodal large model is: "Generation instruction: Help me generate a presentation document directory for the introductory explanation of Transformer, which contains four parts. The first part is the purpose of Transformer, the second part is the basic structure of Transformer, the third part is the advantages of Transformer, and the fourth part is the shortcomings of Transformer. Please generate the HTML corresponding to the single-page presentation document containing only black and white text according to the generation instruction." Step S122: obtaining initial weights, and initializing the multimodal large model based on the initial weights.
[0045] The weights are randomly initialized to obtain the initial weights. The multimodal large model randomly initializes the weights in the first stage of training to prepare for training the multimodal large model to generate the HTML corresponding to the single-page black and white text presentation document.
[0046] Step S123: train the multimodal large model according to the HTML training data of the second stage and the Prompt text of the first stage, generate the HTML corresponding to the single-page color presentation document containing the text style, and save the final training weights of the first stage as the initial weights of the second stage to initialize the multimodal large model weights of the second stage.
[0047] By learning from the HTML training data in the second stage, the multimodal large model can generate the HTML corresponding to a single-page presentation document containing black and white text based on the prompt text in the first stage.
[0048] In the HTML corresponding to the single-page black-and-white text presentation document generated in the first stage, the final trained weights of the first stage are saved as the initial weights of the second stage, and the initial weights of the second stage are passed to the next generation stage, that is, the second stage of generating the HTML corresponding to the single-page color presentation document with text style.
[0049] In step S102, the multimodal large model has the ability to generate HTML corresponding to the presentation document with black and white text by learning the HTML training data of the first stage. The initial weight of the second stage is passed to the second stage, and the ability to generate black and white text presentation documents is continued to the next stage, so that the second stage is carried out at a higher training starting point, thereby more effectively generating presentation documents with text styles.
[0050] Step S103: In the second stage of training, the multimodal large model loads the initial weights of the second stage for initialization. According to the training data set, the training model generates HTML corresponding to a single-page color presentation document containing a text style, and saves the final weights of the second stage as the initial weights of the third stage.
[0051] In the first stage of generating a presentation document containing black and white text, the multimodal large model has the ability to generate HTML corresponding to a presentation document containing only black and white text. Therefore, in the stage of generating a single-page presentation document containing text style, the final weight of the first stage is used as the initialization weight, and combined with the prompt text containing "text style" and the HTML of the text style as training data, the model can be further trained to generate the HTML corresponding to a single-page color presentation document containing text style. Among them, the text style includes the font style (such as bold, regular), font color, and font size.
[0052] Training a large multimodal model to generate a single-page color presentation document with text styles specifically includes the following steps: Step S131, in response to the generation instruction, obtain the second stage Prompt text.
[0053] In the second stage, in response to the generation instruction, the Prompt text includes the generation instruction and the prompt. In the stage of generating a single-page color presentation document with text style, the Prompt text of the input model includes the instruction and the prompt. The prompt is "please generate the HTML corresponding to the single-page color presentation document without images but with text style according to the generation instruction."
[0054] In one embodiment, the prompt text of the second stage specifically includes: if the generation instruction is "Help me generate a presentation document directory about the Transformer introduction, which contains four parts, the first part is the purpose of the Transformer, the second part is the basic structure of the Transformer, the third part is the advantages of the Transformer, and the fourth part is the shortcomings of the Transformer.", then the prompt text for inputting the multimodal large model is: "Generation instruction: Help me generate a presentation document directory about the Transformer introduction, which contains four parts, the first part is the purpose of the Transformer, the second part is the basic structure of the Transformer, the third part is the advantages of the Transformer, and the fourth part is the shortcomings of the Transformer. Please generate the HTML corresponding to the single-page color presentation document without images but with text styles according to the generation instruction." Step S132: The multimodal large model loads the initial weights of the second stage and is initialized according to the initial weights of the second stage.
[0055] The multimodal large model loads the initial weights of the second stage, and the multimodal large model is initialized based on the initial weights of the second stage to prepare for training the multimodal large model to generate HTML corresponding to a single-page color presentation document with a text style.
[0056] Step S133: According to the HTML training data of the second stage and the Prompt text of the second stage, train the multimodal large model to generate HTML corresponding to a single-page color presentation document containing a text style, and save the final training weights of the second stage as the initial weights of the third stage, so as to initialize the multimodal large model weights of the third stage.
[0057] The multimodal large model learns from the HTML training data of the second stage, and the training model generates the HTML corresponding to the single-page color presentation document with text style according to the prompt text of the second stage.
[0058] In the HTML corresponding to the single-page color presentation document with text style generated in the second stage, the final trained weights of the second stage are saved as the initial weights of the third stage, and the initial weights of the third stage are passed to the next generation stage, that is, the third stage of generating the HTML corresponding to the single-page presentation document.
[0059] Continuing to refer to step S103, the second stage is further optimized on the basis of the first stage, that is, the stage of generating a single-page text style presentation document is further optimized on the basis of the stage of generating a single-page black and white text presentation document. Compared with the method of directly generating a presentation document containing text styles, the learning difficulty of generating a presentation document containing text styles is significantly reduced on the premise that the model already has the ability to generate black and white text presentation documents. By initializing with the weights of the first stage, the model can learn based on the existing capabilities, avoid repeated learning, and can also effectively reduce training time and resource consumption. In addition, this phased training strategy enables the model to focus more on learning style features when generating HTML containing text styles, thereby further reducing the overall training difficulty.
[0060] Step S104: In the third stage of training, based on the initial weights and training data set of the third stage, the multimodal large model is trained to generate HTML corresponding to the single-page presentation document, and the final weights of the third stage are saved as the initial weights of the fourth stage.
[0061] The multimodal model can process and understand the relationship between different modes. Therefore, in the stage of generating a single-page presentation document, the multimodal model can also understand and process the image information added to the HTML. The specific steps to generate a single-page presentation document are as follows: Step S141, in response to the generation instruction, obtain the Prompt text of the third stage.
[0062] In the third stage, in response to the generation instruction, the Prompt text includes the generation instruction and the prompt. In the stage of generating a single-page presentation document, the Prompt text of the input model includes the instruction and the prompt. The prompt is "Please generate the HTML corresponding to the single-page presentation document according to the generation instruction."
[0063] In one embodiment, in the third stage, that is, the stage of generating a single-page presentation document, the Prompt text input into the multimodal large model is "Generation instructions: Help me generate a presentation document directory about the introductory explanation of Transformer, which contains four parts. The first part is the purpose of Transformer, the second part is the basic structure of Transformer, the third part is the advantages of Transformer, and the fourth part is the shortcomings of Transformer. Please generate the HTML corresponding to the single-page presentation document according to the generation instructions. Please generate the HTML corresponding to the single-page presentation document according to the generation instructions." Then the output data of the multimodal large model is the corresponding HTML described by the Prompt text.
[0064] When the multimodal large model receives a presentation document page image as input, when outputting HTML, the model represents the image features through quantization encoding and maps these special encodings to special tokens. These tokens are embedded in the HTML generated by the model to identify the corresponding image content. Therefore, to obtain HTML training data containing images, the HTML generated by the model must be processed. First, the special tokens that identify the image are parsed from the generated HTML; then, the image quantization encoding model is used to decode the special tokens to generate the actual image data; finally, these images are inserted into the corresponding image tag positions in the HTML.
[0065] Step S142: The multimodal large model loads the initial weights of the third stage and is initialized according to the initial weights of the third stage.
[0066] The multimodal large model loads the initial weights of the third stage and is initialized based on the initial weights of the third stage in preparation for training the multimodal large model to generate HTML corresponding to a single-page presentation document.
[0067] Step S143: Based on the HTML training data of the third stage, the training model generates HTML corresponding to the single-page presentation document, and saves the final training weights of the third stage as the initial weights of the fourth stage to initialize the model weights of the fourth stage.
[0068] Based on the initial weights of the third stage and the HTML training data of the third stage, the multimodal large model is trained to generate the HTML corresponding to the single-page presentation document according to the prompt text of the third stage. The HTML of the single-page presentation document includes the content of the complete single-page presentation document, such as text, images, and page layout, making the generated HTML closer to the presentation form of the actual application presentation document.
[0069] Since the multimodal large model can receive both text and images as input. The multimodal large model extracts image features through the visual encoder and combines the HTML training data of the third stage to learn to generate HTML files containing images. During the training process, the multimodal large model continuously optimizes the weights by learning the HTML training data of the third stage, so that the generated HTML is closer to the style of the real presentation document. The final training weights of the third stage are saved as the initial weights of the fourth stage, and the initial weights of the fourth stage are passed to the next generation stage, that is, the fourth stage of generating the HTML corresponding to the multi-page presentation document.
[0070] In one embodiment, the table of contents page image of a presentation document is input into a multi-modal model, and the Prompt description output by the model is: This is the table of contents page of a presentation document. The background color is mainly orange. There are Chinese characters "目录" (Table of Contents) and the English word "Contents" in white on the left. There are six orange rectangular boxes on the right. Each box contains a white serial number and text, and the specific content is as follows: 01 Transformer Overview 02 Self-Attention Mechanism 03 Encoder Details 04 Decoder Details 05 Position Encoding 06 Loss Function and Optimization Continuing to refer to step S104, the third stage is to further optimize based on the HTML generated in the second stage. That is, the stage of generating a single-page presentation document is to further optimize based on the HTML corresponding to the generated single-page color presentation document with text styles. Before the start of training in the third stage, the model already has the ability to generate the HTML corresponding to a single-page presentation document with text styles. Therefore, the difficulty of learning to generate the HTML corresponding to a complete single-page presentation document is significantly reduced.
[0071] Step S105, in the fourth stage of training, according to the initial weights of the fourth stage and the training data set, train the model to generate the HTML corresponding to a multi-page presentation document, and construct a presentation document according to the HTML corresponding to the multi-page presentation document.
[0072] The specific steps to obtain a multi-page presentation document are as follows: Step S151, in response to a generation instruction, obtain the Prompt text of the fourth stage; In the fourth stage, in response to the generation instruction, the Prompt text includes the generation instruction and the prompt. The prompt is: "Please generate the HTML corresponding to the multi-page presentation document according to the generation instruction." For example, the generation instruction is "Help me generate a presentation document directory about the Transformer introduction, which contains four parts. The first part is the purpose of the Transformer, the second part is the basic structure of the Transformer, the third part is the advantages of the Transformer, and the fourth part is the shortcomings of the Transformer." Then the Prompt text for inputting the multimodal large model is: "Generation instruction: Help me generate a presentation document directory about the Transformer introduction, which contains four parts. The first part is the purpose of the Transformer, the second part is the basic structure of the Transformer, the third part is the advantages of the Transformer, and the fourth part is the shortcomings of the Transformer. Please generate the HTML corresponding to the multi-page presentation document according to the generation instruction." Step S152: The multimodal large model loads the initial weights of the fourth stage for initialization.
[0073] The multimodal large model loads the initial weights of the fourth stage for initialization in preparation for training to generate HTML corresponding to the multi-page presentation document.
[0074] Step S153: Based on the HTML training data of the fourth stage and the Prompt text of the fourth stage, the training model generates HTML corresponding to the multi-page presentation document, and parses the HTML corresponding to the multi-page presentation document, determines the layout and content of each page of the presentation document according to the parsing results, and obtains the multi-page presentation document.
[0075] Based on the initial weights of the fourth stage and the HTML training data of the fourth stage, the training model generates the corresponding HTML of the multi-page presentation document according to the prompt text of the fourth stage.
[0076] The HTML training data in the fourth stage is generated by inputting the reference presentation document into the multimodal large model. Therefore, the HTML training data in the fourth stage includes the structure and content information of the multi-page presentation document. In the fourth stage, the model is further trained on the basis of the previous training, and the ability to generate multi-page presentation documents is learned by the training data. By training based on the final weights of the third stage, the learning results of the previous stage are carried forward, thereby accelerating the training process.
[0077] Parse the HTML of the multi-page presentation document and build the presentation document page by page based on the parsed content. The built content includes the layout and specific content of each page of the presentation document to form the final multi-page presentation document.
[0078] In another embodiment, Figure 2 is a flowchart of a method for automatically generating a presentation document training according to an exemplary embodiment. Figure 2 As shown, step 1 is to screen open source presentation documents, and convert the presentation document page images into text descriptions through a multimodal large model, and convert the presentation document page images into HTML to construct a dataset. Step 2. In the first stage, generate the HTML corresponding to a single-page presentation document containing only black and white text, and use the final weights trained in this stage for initialization of the next stage. Step 3. In the second stage, generate the HTML corresponding to a single-page color presentation document that does not contain an image, and use the final weights trained in this stage for initialization of the next stage. Step 4. In the third stage, generate the HTML corresponding to a single-page presentation document containing all elements, and apply the final weights trained in this stage to the initialization of the next stage. Step 5. In the fourth stage, train the model to generate HTML corresponding to multiple pages of presentation documents.
[0079] In summary, the automatic generation of presentation document training method provided by the embodiment of the present application adopts a phased training strategy, generates tasks specifically according to each stage, and selects corresponding training data. Since the text style generation stage, each stage continues to optimize based on the training results of the previous stage, thereby avoiding independent repeated learning of each stage. Compared with direct training, this step-by-step learning method effectively reduces the content that the model needs to learn, reduces the learning difficulty of the model, and the quality of the generated presentation document is more stable.
[0080] In a second aspect, an embodiment of the present application provides a training device for automatically generating presentation documents. Figure 3 FIG. 1 is a block diagram of a device for automatically generating a presentation document training program according to an exemplary embodiment. Figure 3 As shown, the automatic generation of demonstration document training device includes: A training data set acquisition module is used to acquire a reference demonstration document, and acquire a training data set according to the reference demonstration document; A first-stage training module, used for, in the first-stage training, in response to a generation instruction, training the model to generate HTML corresponding to a single-page presentation document containing only black and white text according to the Prompt text, and saving the final weight of the first stage as the initial weight of the second stage; The second stage training module is used for initializing the model by loading the initial weights of the second stage in the second stage training, training the model to generate the HTML corresponding to the single-page color presentation document with text style according to the prompt text, and saving the final weights of the second stage as the initial weights of the third stage; The third stage training module is used to train the multimodal large model to generate HTML corresponding to the single-page presentation document based on the initial weights of the third stage and the training data set in the third stage training, and save the final weights of the third stage as the initial weights of the fourth stage; The fourth stage training module is used to train the model to generate HTML corresponding to a multi-page presentation document according to the initial weights of the fourth stage and the training data set during the fourth stage training, and to construct a presentation document according to the HTML corresponding to the multi-page presentation document.
[0081] In summary, the automatic presentation document generation training device provided by the present application divides the generation task in the generation instruction into multiple stages, and gradually optimizes based on different training data. The model sequentially generates HTML corresponding to a presentation document containing only black and white text, HTML corresponding to a presentation document with a text style, and HTML corresponding to a single-page complete presentation document, and finally generates HTML corresponding to a multi-page presentation document. Each stage uses the final weights trained in the previous stage to initialize the model weights of the next stage to achieve progressive learning. This staged progressive training to generate presentation documents effectively reduces the complexity of the generation task. Compared with the method of directly training to generate multi-page presentation documents, the present application generates content through staged learning, thereby significantly reducing the learning complexity of the model and improving the stability and quality of the generation results.
[0082] It should be noted that the automatic generation of presentation document training device provided in this embodiment is used to implement the above-mentioned implementation mode, and the description has been made no further. As used above, the terms "module", "unit", "subunit" and the like can implement a combination of software and / or hardware for a predetermined function. Although the device described in the above embodiment is preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and conceivable.
[0083] In a third aspect, an embodiment of the present application provides an electronic device, Figure 4 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Figure 4 As shown, the electronic device may include a processor 81 and a memory 82 storing computer program instructions.
[0084] Specifically, the processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0085] Among them, the memory 82 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 82 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 82 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 82 may be inside or outside the data processing device. In a specific embodiment, the memory 82 is a non-volatile memory. In a specific embodiment, the memory 82 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0086] The memory 82 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81 .
[0087] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the automatic generation of presentation document training methods in the above embodiments.
[0088] In one embodiment, the automatic generation of presentation document training equipment may further include a communication interface 83 and a bus 80. Figure 4 As shown, the processor 81, the memory 82, and the communication interface 83 are connected via a bus 80 and communicate with each other.
[0089] The communication interface 83 is used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. The communication interface 83 can also implement data communication with other components such as: external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.
[0090] The bus 80 includes hardware, software or both, and couples the components of the automatic presentation document generation training device to each other. The bus 80 includes but is not limited to at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, bus 80 may include one or more buses. Although embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.
[0091] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a program stored thereon, and when the program is executed by a processor, the method for automatically generating a presentation document training provided in the first aspect is implemented.
[0092] The readable storage medium may include but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device or any suitable combination of the above.
[0093] In a possible implementation, the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps of the automatic generation of presentation document training method provided in the first aspect.
[0094] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on a user device, partially on a user device, as an independent software package, partially on a user device and partially on a remote device, or entirely on a remote device.
[0095] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0096] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for automatically generating a demonstration document training, characterized in that: The method is applied to a multimodal large model including: Obtain a reference demonstration document, and obtain a training data set based on the reference demonstration document; In the first stage of training, in response to the generation instruction, the multimodal large model is trained to generate HTML corresponding to a single-page presentation document containing only black and white text according to the training data set, and the final weight of the first stage is saved as the initial weight of the second stage; In the second stage of training, the multimodal large model loads the initial weights of the second stage for initialization, and the training model generates HTML corresponding to a single-page color presentation document containing a text style according to the training data set, and saves the final weights of the second stage as the initial weights of the third stage; In the third stage of training, based on the initial weights of the third stage and the training data set, the multimodal large model is trained to generate HTML corresponding to a single-page presentation document, and the final weights of the third stage are saved as initial weights of the fourth stage; In the fourth stage of training, the multimodal large model is trained to generate HTML corresponding to a multi-page presentation document based on the initial weights of the fourth stage and the training data set, and a presentation document is constructed based on the HTML corresponding to the multi-page presentation document.
2. The method for automatically generating presentation documents according to claim 1, characterized in that: The step of obtaining a training data set according to the reference demonstration document and the multimodal large model includes: Using the page image of the reference presentation document as the input of the multimodal large model, obtaining the corresponding HTML converted from the Prompt text and the page image of the reference presentation document; Based on the HTML, generate four stages of training data; According to the task requirements of each stage, the prompt text is adjusted to generate prompt texts of four stages; The Prompt text and HTML at the same stage are paired to form the training data set.
3. The method for automatically generating presentation documents according to claim 2, characterized in that: The HTML generates four stages of training data, including: Processing the HTML to obtain complete HTML, and using the complete HTML as HTML training data in the fourth stage; Splitting the HTML training data of the fourth stage, retaining complete image and text information, obtaining HTML content corresponding to a single-page presentation document, and saving the HTML content corresponding to the single-page presentation document as the HTML training data of the third stage; Remove the image elements in the third-stage HTML content, retain only the text content, obtain the HTML corresponding to the text single-page presentation document, and save the HTML corresponding to the single-page color presentation document containing the text style as the second-stage HTML training data; The text style of the HTML of the second stage is removed, and only the black and white text content is retained to obtain the HTML corresponding to the black and white text single-page presentation document, and the HTML corresponding to the black and white text single-page presentation document is saved as the HTML training data of the first stage.
4. The method for automatically generating presentation documents according to claim 3, characterized in that: In response to the generating instruction, the multimodal large model is trained to generate HTML corresponding to a single-page presentation document containing only black and white text according to the training data set, and the final weight of the first stage is saved as the initial weight of the second stage, including: In response to the generation instruction, obtaining the prompt text of the first stage; Obtaining random initial weights, the multimodal large model is initialized based on the initial weights; Based on the HTML training data of the first stage and the Prompt text of the first stage, the multimodal large model is trained to generate HTML corresponding to the presentation document containing only black and white text, and the final training weights of the first stage are saved as the initial weights of the second stage to initialize the multimodal large model weights of the second stage.
5. The method for automatically generating presentation documents according to claim 3, characterized in that: In the second stage training, the multimodal large model loads the initial weights of the second stage for initialization, and according to the training data set, the training model generates HTML corresponding to a single-page color presentation document containing a text style, and saves the final weights of the second stage as the initial weights of the third stage, including: In response to the generation instruction, obtaining a second-stage Prompt text; The multimodal large model loads the initial weights of the second stage and is initialized according to the initial weights of the second stage; According to the HTML training data of the second stage and the Prompt text of the second stage, the multimodal large model is trained to generate HTML corresponding to a single-page color presentation document containing a text style, and the final training weights of the second stage are saved as the initial weights of the third stage to initialize the multimodal large model weights of the third stage.
6. The method for automatically generating presentation documents according to claim 3, characterized in that: The step of training the multimodal large model to generate HTML corresponding to a single-page presentation document based on the initial weights of the third stage and the training data set, and saving the final weights of the third stage as the initial weights of the fourth stage, includes: In response to the generation instruction, the prompt text of the third stage is obtained; The multimodal large model loads the initial weights of the third stage and is initialized according to the initial weights of the third stage; According to the HTML training data of the third stage and the Prompt text of the third stage, the initialized multimodal large model is trained to generate HTML corresponding to a single-page presentation document, and the final training weights of the third stage are used as the initial weights of the fourth stage to initialize the weights of the multimodal large model of the fourth stage.
7. The method for automatically generating presentation documents according to claim 3, characterized in that: The step of training the multimodal large model to generate HTML corresponding to a multi-page presentation document according to the initial weights of the fourth stage and the training data set, and constructing a presentation document according to the HTML corresponding to the multi-page presentation document includes: In response to the generation instruction, obtaining the Prompt text of the fourth stage; The multimodal large model loads the initial weights of the fourth stage for initialization; According to the HTML training data of the fourth stage and the Prompt text of the fourth stage, the initialized multimodal large model is trained to generate HTML corresponding to a multi-page presentation document, and the HTML corresponding to the multi-page presentation document is parsed. The layout and content of each page of the presentation document are determined according to the parsing results to obtain a multi-page presentation document.
8. A training device for automatically generating presentation documents, characterized in that: include: A training data set acquisition module is used to acquire a reference demonstration document, and acquire a training data set according to the reference demonstration document; A first-stage training module, used for, in the first-stage training, in response to a generation instruction, training the multimodal large model to generate HTML corresponding to a single-page presentation document containing only black and white text according to the training data set, and saving the final weight of the first stage as the initial weight of the second stage; The second stage training module is used for loading the initial weights of the second stage into the multimodal large model for initialization in the second stage training, and generating HTML corresponding to a single-page color presentation document containing a text style based on the training data set by the training model, and saving the final weights of the second stage as the initial weights of the third stage; A third-stage training module, used for training the multimodal large model to generate HTML corresponding to a single-page presentation document based on the initial weights of the third stage and the training data set in the third-stage training, and saving the final weights of the third stage as the initial weights of the fourth stage; The fourth stage training module is used to train the multimodal large model to generate HTML corresponding to a multi-page presentation document according to the initial weights of the fourth stage and the training data set in the fourth stage training, and to construct a presentation document according to the HTML corresponding to the multi-page presentation document.
9. An electronic device, characterized in that: including memory, processor, and A computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for automatically generating a presentation document training method as claimed in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for automatically generating a presentation document training method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text generation model training method, text generation method and respective devices
CN114997395A
Text processing method, article generation method and text processing model training method
CN115994522A
Automatic generation method and device of presentation document, equipment and storage medium
CN117852500A
Document generation method and device based on multi-modal large model, equipment and medium
CN119203944A
Content augmentation with machine generated content to meet content gaps during interaction with target entities
US20230121711A1