Presentation generation method and training method based on large model
By analyzing presentation requirements text using a large model, generating a set of building elements and laying out the page, the problem of presentation template limitations is solved, and more efficient and flexible presentation generation is achieved.
Patent Information
- Application Number
- CN202410599478.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-05-14
AI Technical Summary
In existing technologies, the flexibility and practicality of presentation templates are lacking, resulting in low efficiency and quality of presentation generation methods.
The presentation requirements text is obtained from a large model, analyzed and generated into a set of related building elements, and the page layout is performed to generate the target presentation.
It improves the efficiency and quality of presentation generation, reduces the difficulty of user operation, enhances the flexibility and practicality of the generation method, and avoids dependence on templates.
Smart Images

Figure CN118428332B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, and in particular to the fields of artificial intelligence technologies such as natural language processing and deep learning. Background Technology
[0002] With the development of technology, people have higher and higher requirements for the efficiency and quality of presentation creation. Optionally, the creation of presentations needs to be based on layout design, pagination, and illustrations.
[0003] In related technologies, existing presentation templates can be called up, and presentations can be created by filling the blank spaces in the template with content. However, because presentation templates have certain limitations on the content that can be filled, the presentation generation method based on presentation templates is not very flexible or practical. Summary of the Invention
[0004] This disclosure proposes a presentation generation method based on a large model and a training method for the large model.
[0005] According to a first aspect of this disclosure, a presentation generation method based on a large model is proposed, comprising: obtaining presentation requirement text from a user; inputting the presentation requirement text into a large model, and obtaining a set of associated presentation building elements based on the presentation requirement text using the model capabilities of the large model; and performing page layout on the set of presentation building elements using the model capabilities of the large model to generate a target presentation corresponding to the presentation requirement text.
[0006] According to a second aspect of this disclosure, a training method for a large model is proposed, comprising: obtaining a candidate large model to be optimized; obtaining a sample requirement text and a sample initial presentation corresponding to the sample requirement text, and calibrating the sample initial presentation to obtain a sample candidate presentation; obtaining sample presentation building elements of the sample candidate presentation to generate training samples for the candidate large model; optimizing and training the candidate large model using the training samples until the training is completed to obtain an optimized target large model, wherein the target large model is used to implement the presentation generation method based on the large model proposed in the first aspect above.
[0007] According to a third aspect of this disclosure, a presentation generation device based on a large model is proposed, comprising: a first acquisition module for acquiring presentation requirement text from a user terminal; a second acquisition module for inputting the presentation requirement text into a large model and obtaining a set of associated presentation building elements based on the presentation requirement text using the model capabilities of the large model; and a first generation module for performing page layout on the set of presentation building elements using the model capabilities of the large model to generate a target presentation corresponding to the presentation requirement text.
[0008] According to the fourth aspect of this disclosure, a training apparatus for a large model is proposed, comprising: a third acquisition module for acquiring candidate large models to be optimized; a fourth acquisition module for acquiring sample requirement text from a sample user terminal and a sample initial presentation corresponding to the sample requirement text, and calibrating the sample initial presentation to obtain a sample candidate presentation from the sample user terminal; a second generation module for acquiring sample presentation building elements of the sample candidate presentation to generate training samples for the candidate large model; and a training module for optimizing and training the candidate large model using the training samples until training is completed to obtain an optimized target large model, wherein the target large model is used to implement the large model-based presentation generation apparatus proposed in the third aspect above.
[0009] According to a fifth aspect of this disclosure, an electronic device is proposed, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the presentation generation method based on a large model proposed in the first aspect above and / or the training method for a large model proposed in the second aspect above.
[0010] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is proposed, wherein the computer instructions are used to cause the computer to execute the presentation generation method based on a large model proposed in the first aspect and / or the training method of a large model proposed in the second aspect.
[0011] According to the seventh aspect of this disclosure, a computer program product is proposed, comprising a computer program that, when executed by a processor, implements the presentation generation method based on a large model as proposed in the first aspect and / or the training method for a large model as proposed in the second aspect.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1 This is a flowchart illustrating a method for generating presentations based on a large model, according to an embodiment of this disclosure.
[0015] Figure 2 This is a flowchart illustrating another embodiment of the presentation generation method based on a large model.
[0016] Figure 3 This is a flowchart illustrating another embodiment of the presentation generation method based on a large model.
[0017] Figure 4 This is a schematic flowchart of a large model training method according to an embodiment of the present disclosure;
[0018] Figure 5 This is a flowchart illustrating a training method for a large model according to another embodiment of this disclosure;
[0019] Figure 6 This is a schematic diagram of a presentation page according to an embodiment of the present disclosure;
[0020] Figure 7 This is a schematic diagram of a presentation page for another embodiment of this disclosure;
[0021] Figure 8 This is a schematic diagram of a presentation page for another embodiment of this disclosure;
[0022] Figure 9 This is a schematic diagram of a presentation page for another embodiment of this disclosure;
[0023] Figure 10 This is a schematic diagram of the structure of a presentation generation device based on a large model according to an embodiment of the present disclosure;
[0024] Figure 11 This is a schematic diagram of the structure of a training device for a large model according to an embodiment of the present disclosure;
[0025] Figure 12 This is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0027] Data processing is a fundamental aspect of systems engineering and automatic control. Data is a form of expression of facts, concepts, or instructions, which can be processed manually or by automated devices. After data is interpreted and given meaning, it becomes information. Data processing involves the acquisition, storage, retrieval, processing, transformation, and transmission of data. The basic purpose of data processing is to extract and derive valuable and meaningful data from large amounts of potentially chaotic and difficult-to-understand data.
[0028] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Its main applications include machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, and Chinese OCR.
[0029] Artificial Intelligence (AI) is a new branch of computer science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. It attempts to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Since its inception, AI has matured in both theory and technology, and its applications have expanded continuously. It is conceivable that future AI-driven technological products will serve as "containers" of human wisdom. AI can simulate the information processes of human consciousness and thought.
[0030] Figure 1 This is a flowchart illustrating a method for generating presentations based on a large model, according to an embodiment of this disclosure. Figure 1 As shown, the method includes:
[0031] S101, Obtain the presentation request text from the user.
[0032] In daily work, people can use presentations (PowerPoint, PPT) to display content. In this scenario, people can create PPTs through layout design, pagination, and the addition of images.
[0033] In this embodiment of the disclosure, presentations can be created by calling the model capabilities of a large model. In this scenario, the user can provide the large model with the corresponding PPT generation requirements, and then generate the presentation by calling the model capabilities of the large model.
[0034] Among them, the generation requirement text provided by the user to the large model for generating PPT can be marked as the user's presentation requirement text.
[0035] S102, input the presentation requirement text into the large model, and obtain the associated set of presentation building elements based on the presentation requirement text using the model capabilities of the large model.
[0036] In this embodiment of the disclosure, after obtaining the presentation requirement text from the user, the obtained presentation requirement text can be input into the pre-acquired large model. The large model's modeling capabilities are used to analyze the presentation requirement text, thereby generating the presentation.
[0037] Optionally, a presentation can be constructed based on a combination of various elements such as images, icons, text, and lines. The elements used to construct and generate the presentation can be defined as the building elements of the presentation.
[0038] In this scenario, the large model's capabilities can be used to analyze the presentation requirement text. Based on the analysis results, the building elements required to construct the presentation corresponding to the presentation requirement text can be obtained and marked as presentation building elements associated with the presentation requirement text. The set of these presentation building elements can then be marked as the presentation building element set.
[0039] It should be noted that the large model proposed in this embodiment can be a large language model or other large models capable of generating presentations; no specific limitation is made here.
[0040] S103 utilizes the modeling capabilities of a large model to perform page layout on the presentation building element set, generating the target presentation corresponding to the presentation requirement text.
[0041] In this embodiment of the disclosure, after obtaining the presentation building element set, the page layout of each element in the presentation building element set can be performed using the modeling capabilities of the large model.
[0042] This can be understood as using the model capabilities of a large model to determine the layout position of each presentation building element in the presentation building element set within the presentation, and then laying out each presentation building element according to its layout position to obtain the presentation generated by combining the presentation building elements, and finally identifying this presentation as the target presentation corresponding to the presentation requirement text.
[0043] This disclosure presents a presentation generation method based on a large model. It obtains the presentation requirement text from the user, leverages the model capabilities of the large model to acquire the set of presentation building elements associated with the requirement text, and then uses these capabilities to perform page layout on the set of building elements to obtain the target presentation. This method improves the efficiency and quality of presentation generation by generating the target presentation through a large model. The user only needs to input the presentation requirement text, reducing the operational difficulty. The method uses the model capabilities of the large model to perform page layout on the set of building elements to obtain the target presentation. Compared to related technologies that rely on calling existing presentation templates, this method avoids the limitations imposed by presentation templates on the presentation generation method, improving the flexibility and practicality of the presentation generation method and optimizing the user experience.
[0044] In the above embodiments, the acquisition of the presentation building element set can be combined with... Figure 2 To understand further, Figure 2 This is a flowchart illustrating another embodiment of the presentation generation method based on a large model, as shown below. Figure 2 As shown, the method includes:
[0045] S201 utilizes the capabilities of the large model to segment the presentation requirement text, resulting in multiple segmented requirement sub-texts.
[0046] In this embodiment of the disclosure, the presentation requirement text may include a variety of information. The presentation requirement text can be segmented based on a preset presentation requirement text segmentation strategy using the model capabilities of a large model, and the multiple texts obtained after segmentation are marked as multiple requirement sub-texts.
[0047] Optionally, the presentation requirement text can be analyzed using the modeling capabilities of a large model. Based on the analysis results, the page information included in each of the multiple presentation pages required to build the target presentation can be determined. The presentation requirement text can then be segmented based on the page information to obtain multiple segmented requirement sub-texts.
[0048] Optionally, the modeling capabilities of a large model can be used to analyze the heading levels in the presentation requirement text, and then the text can be segmented according to a preset presentation requirement text segmentation method based on heading levels to obtain multiple segmented requirement sub-texts.
[0049] Specifically, it can obtain the set of text titles in the presentation requirement text and the title level of each text title in the set of text titles. Then, through the model capabilities of the large model, it can segment the presentation requirement text according to the title level to obtain multiple segmented requirement sub-texts.
[0050] In this embodiment of the disclosure, the presentation requirement text may include multiple title texts, and the title texts have different title levels. In this scenario, the presentation requirement text can be segmented by title level.
[0051] As an example, the presentation requirements text is set to a lightweight markup language (Markdown) format. The Markdown presentation requirements text includes the following headings: First heading "Changes in population growth rate among different population groups", Second heading "Urban and rural population growth", Third heading "Differences in knowledge level and physical fitness", Fourth heading "Genomic quality and population differences", and Fifth heading "Differences in population growth rate". The first heading is the first heading level, and the second, third, fourth, and fifth headings are all second heading levels.
[0052] In this example, the segmentation strategy for the presentation requirement text is set to segment the text information under the headings of the second heading level. In this example, the text information under the second, third, fourth, and fifth headings in the presentation requirement text can be segmented to obtain multiple segmented requirement subtexts consisting of the first requirement subtext, the second requirement subtext, the third requirement subtext, and the fourth requirement subtext, as shown below.
[0053] The first requirement subtext: "Urban and rural population growth".
[0054] The second requirement subtext: "Differences in knowledge level and physical condition".
[0055] The third requirement subtext: "Genome quality and population differences".
[0056] The fourth requirement subtext: "Differences in population growth rate".
[0057] S202, through the modeling capabilities of the large model, traverses multiple requirement sub-texts to obtain a subset of presentation building elements for each requirement sub-text.
[0058] In this embodiment of the disclosure, for any requirement subtext, the set of presentation building elements required to construct the presentation page corresponding to the requirement subtext can be determined as the subset of presentation building elements of the requirement subtext.
[0059] Optionally, for any requirement subtext, the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements associated with the requirement subtext are obtained through the model capabilities of the large model, and a subset of presentation building elements of the requirement subtext is obtained based on the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements.
[0060] In this embodiment of the disclosure, the presentation building elements may include shape elements, textbox elements, image elements, icon elements, and line elements. Shape elements may be marked as shape building elements, textbox elements as textbox building elements, image elements as image building elements, icon elements as icon building elements, and line elements as line building elements.
[0061] In this scenario, for any requirement subtext, the model capabilities of the large model can be used to obtain the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements required to construct the presentation page corresponding to the requirement subtext. The set of elements composed of the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements is marked as the presentation building element subset of the requirement subtext.
[0062] S203, based on the subset of presentation building elements of each requirement subtext, obtain the set of presentation building elements associated with the presentation requirement text.
[0063] In this embodiment of the disclosure, after obtaining the presentation building element subsets of each of the multiple requirement sub-texts, these presentation building element subsets can be combined to obtain a presentation building element set.
[0064] Optionally, the text order of multiple requirement sub-texts in the presentation requirement text can be obtained, and the presentation building element subsets of each of the multiple requirement sub-texts can be sorted and combined according to the text order, and the set obtained by sorting and combining can be determined as the presentation building element set associated with the presentation requirement text.
[0065] Optionally, the presentation order among the various requirement sub-texts can be obtained, and the presentation building element subsets of each of the multiple requirement sub-texts can be sorted and combined according to the presentation order to obtain the presentation building element set associated with the presentation requirement text.
[0066] This disclosure presents a presentation generation method based on a large model. It segments the presentation requirement text using the modeling capabilities of the large model to obtain multiple requirement sub-texts. Then, it obtains a subset of presentation building elements for each of these sub-texts, resulting in a set of presentation building elements associated with the presentation requirement text. This disclosure improves the efficiency and accuracy of presentation requirement text segmentation by leveraging the modeling capabilities of the large model, thereby enhancing the accuracy of the adaptation between the target presentation and the presentation requirement text. This improves the generation efficiency and quality of the target presentation. Furthermore, obtaining the set of presentation building elements from the presentation requirement text through the modeling capabilities of the large model provides support for subsequent page layout using these capabilities. This allows the target presentation generation method to function without relying on existing presentation templates, increasing the flexibility and practicality of the presentation generation method.
[0067] In the above embodiments, the acquisition of the target presentation document can be combined with... Figure 3 To understand further, Figure 3 This is a flowchart illustrating another embodiment of the presentation generation method based on a large model, as shown below. Figure 3 As shown, the method includes:
[0068] S301, for any requirement subtext, obtain the subset of presentation building elements corresponding to the requirement subtext in the presentation building element set.
[0069] For details regarding step S301 in this embodiment, please refer to the relevant information in the above embodiments, which will not be repeated here.
[0070] S302, obtain the element description tags of each presentation building element in the presentation building element subset, and obtain the target page layout file of the presentation building element subset.
[0071] In this embodiment of the disclosure, the presentation building element subset includes multiple presentation building elements. In order to realize the page layout of each presentation building element, the relevant layout information of each presentation building element can be obtained through the model capabilities of the large model.
[0072] Optionally, for any requirement subtext, obtain the candidate block tags of each presentation building element of the requirement subtext.
[0073] In this embodiment of the disclosure, for any requirement subtext, after obtaining the presentation building element subset of the requirement subtext, block tags can be generated for each presentation building element in the presentation building element subset based on the block tag generation algorithm in the related technology, and the block tags of each presentation building element are marked as candidate block tags.
[0074] It should be noted that the candidate block tags for shape building elements, text box building elements, and line building elements can be tags in Hyper Text Markup Language (HTML) code snippets. The candidate block tags for image building elements and symbol building elements can be tags in HTML code snippets. () can also be other types of tags, without specific restrictions here.
[0075] Optionally, for any presentation building element, obtain the element description information of the presentation building element, and fill the element description information into the candidate block tag to generate the element description tag of the presentation building element.
[0076] In this embodiment of the disclosure, the description information of the presentation building elements in the presentation can be marked as element description information. It can be understood that the appearance of the presentation building elements in the presentation can be determined through the element description information.
[0077] As an example, for an image building element, its element description information can include at least style attribute information such as shape type, shape width, height, and background color.
[0078] For text box building elements, their element description information can include at least the relevant text content, as well as style attribute information such as font, font size, color, and alignment of the text in the text content.
[0079] For image building elements, their element description information can include at least the image's network link (uniform resource location, URL), image width, image height, and other attribute information.
[0080] For symbol building elements, their element description information can include at least the symbol's network link (uniform resource location, URL), symbol width, symbol height, and other attribute information.
[0081] For line-building elements, their element description information can include at least the coordinates of the start and end points of the line, attribute information such as border style and position, and information on the simulated effect of the line-building element in the presentation.
[0082] Optionally, for any requirement subtext, obtain the element layout information set of each element subset of the presentation of the requirement subtext, and generate the first candidate layout file and the second candidate layout file of the requirement subtext based on the element layout information set.
[0083] In this embodiment of the disclosure, each presentation building element in the subset of presentation building elements has layout information, which can be marked as the element layout information of each presentation building element, and the set of this part of the element layout information is marked as the element layout information set of each presentation building element.
[0084] The element layout information of each presentation building element can include at least the position information of each presentation building element during the presentation building process, such as coordinate information, width information, and height information, etc., without specific limitations here.
[0085] Optionally, a subset of presentation building elements for each of the multiple requirement sub-files can be obtained. For the process of obtaining the subset of presentation building elements for each of the multiple requirement sub-files, please refer to the relevant content in the above embodiments for understanding, and it will not be repeated here.
[0086] Optionally, for any requirement subtext, shared element layout information is extracted from the element layout information set of each subset of elements in the presentation of the requirement subtext to generate a first candidate layout file for the requirement subtext.
[0087] In this embodiment of the disclosure, for any requirement subtext, the element layout information set of each subset of the presentation elements of the requirement subtext may contain the same or similar element layout information.
[0088] As an example, if all presentation building elements of the requirement subtext are set to be black, then in this example, the color-related layout information in the element layout information set of each subset of presentation building elements can be determined to be the same or similar element layout information.
[0089] In this scenario, the same or similar element layout information in each element layout information set can be marked as shared element layout information in each element layout information set.
[0090] Furthermore, shared element layout information can be extracted from the set of element layout information, and a layout file that can be shared by each presentation element can be constructed based on the shared element layout information, and the layout file can be marked as the first candidate layout file for the requirement subtext.
[0091] As an example, the first candidate layout file may include the absolute positioning information (position:absolute), left margin information (left), top margin information (top), width information (width), and height information (height) of each presentation building element of the corresponding requirement subtext, and may also include other shared element layout information, which is not specifically limited here.
[0092] It should be noted that, for any requirement subtext, the first candidate layout file for that requirement subtext can be a code snippet (style) in the HTML code snippet corresponding to the target presentation of that requirement subtext, or it can be an independent Cascading Style Sheets (CSS) file; no specific limitation is made here.
[0093] Optionally, for any presentation building element in the subset of presentation building elements, exclusive element layout information other than shared element layout information is obtained from the element layout information set of the presentation building elements, and a second candidate layout file for the required subtext is generated based on the exclusive element layout information.
[0094] In this embodiment of the disclosure, after extracting shared element layout information from the element layout information sets of each subset of presentation building elements, the remaining element layout information in each element layout information set other than the extracted shared element layout information can be marked as the exclusive element layout information of each presentation building element.
[0095] This can be understood as the exclusive element layout information for any presentation building element, which can only be interpreted based on the position of that presentation building element during the presentation building process.
[0096] Furthermore, based on the preset layout file construction method, the exclusive element layout information of each presentation building element is constructed into a file, and the layout file constructed in this scenario is marked as the second candidate layout file of the requirement subtext to which each presentation building element belongs.
[0097] It should be noted that, for any requirement subtext, the second candidate layout file for that requirement subtext can be a code snippet (style) in the HTML code snippet corresponding to the target presentation of that requirement subtext, or it can be an independent Cascading Style Sheets (CSS) file; no specific limitation is made here.
[0098] Optionally, a target page layout file for the presentation element subset of the required subtext is obtained based on the first candidate layout file and the second candidate layout file.
[0099] In this embodiment of the disclosure, for any requirement sub-text, the first candidate layout file and the second candidate layout file of the requirement sub-text can be integrated based on a preset integration method, and the integrated file is determined as the target page layout file of the presentation building element subset of the requirement sub-text.
[0100] This can be understood as follows: by leveraging the modeling capabilities of a large model, based on the target page layout file, the position information of each presentation building element in the presentation building element subset of the requirement subtext can be determined during the presentation building process, thereby determining the layout of each presentation building element.
[0101] As an example, if the first and second candidate layout files are set as style code snippets in HTML code snippets, then in this example, the first and second candidate layout files can be integrated based on the integration method of HTML code to obtain the corresponding target page layout file.
[0102] As another example, if the first candidate layout file and the second candidate layout file are set to CSS files, then in this example, the first candidate layout file and the second candidate layout file can be integrated based on the CSS file integration method in related technologies to obtain the corresponding target page layout file.
[0103] S303, through the modeling capabilities of the large model, lays out each presentation building element according to the element description tags and the target page layout file, and generates presentation subpages of the required subtext.
[0104] In this embodiment of the disclosure, the modeling capabilities of the large model can be used to obtain the appearance of each presentation building element based on the element description tags, and the layout information of each presentation building element can be determined based on the target page layout file.
[0105] In this scenario, for any requirement subtext, the model capabilities of the large model can be used to lay out each presentation building element in the subset of presentation building elements of the requirement subtext, thereby generating the page corresponding to the requirement subtext and marking the page as the presentation subpage of the requirement subtext.
[0106] S304. Based on the presentation subpages of each requirement subtext, obtain the target presentation of the presentation requirement text.
[0107] In this embodiment of the disclosure, after obtaining the presentation subpages of each requirement subtext, the presentation subpages can be sorted and combined to obtain the target presentation of the presentation requirement text.
[0108] This involves obtaining the text order of each requirement sub-text within the presentation requirement text, and then sorting and combining each presentation sub-page based on this text order to obtain the target presentation.
[0109] Furthermore, it can obtain the presentation order between the required sub-texts and sort and combine the presentation sub-pages based on this presentation order to obtain the target presentation.
[0110] Optionally, the presentation subpage code of each requirement subtext can be obtained through the modeling capabilities of the large model to obtain the target presentation code of the presentation requirement text.
[0111] In this embodiment of the disclosure, the presentation subpage of each requirement subtext can be HTML code. In this scenario, each presentation subpage can be marked as its own presentation subpage code.
[0112] Optionally, the code of each presentation subpage can be integrated based on the HTML code integration method in related technologies to obtain an integrated code snippet, and the integrated code snippet can be marked as the target presentation code.
[0113] Further, the target presentation code is rendered to obtain the target presentation.
[0114] In this embodiment of the disclosure, the target presentation code can be rendered based on the rendering method in the related technology, so as to display the rendered target presentation on the user terminal.
[0115] It should be noted that the target presentation code can be rendered using rendering tools in a browser, or it can be rendered using other rendering tools; no specific limitation is made here.
[0116] This disclosure presents a presentation generation method based on a large model. It obtains a subset of presentation building elements for each requirement sub-text, acquires element description tags for each subset, and obtains the target page layout file for each subset. Leveraging the capabilities of the large model, it performs page layout for each presentation building element based on the element description tags and the target page layout file, generating presentation sub-pages for each requirement sub-text, thereby obtaining the target presentation for the presentation requirement text. This disclosure improves the accuracy of appearance and layout information for each presentation building element by obtaining the appearance and layout information through element description tags and the target page layout file. The target page layout file includes a shared first candidate layout file and a unique second candidate layout file. In scenarios where the target presentation is rendered based on HTML code, this reduces the character (token) usage of the HTML code snippets in the target presentation, thereby reducing the processing workload of the HTML code. Furthermore, by utilizing the capabilities of the large model to lay out each presentation building element, it improves the efficiency and quality of obtaining the target presentation.
[0117] This disclosure also proposes a training method for large models, which can be combined with Figure 4 understand, Figure 4 This is a flowchart illustrating a large model training method according to an embodiment of the present disclosure, as shown below. Figure 4 As shown, the method includes:
[0118] S401, Obtain candidate large models to be optimized.
[0119] In this embodiment of the disclosure, the large model before training can be marked as a candidate large model to be optimized. The candidate large model can be a large language model or other types of large models, which are not specifically limited here.
[0120] S402, obtain the sample requirement text and the corresponding initial sample presentation, and calibrate the initial sample presentation to obtain the sample candidate presentation.
[0121] In this embodiment of the disclosure, the requirement text used for training and optimizing the candidate large model can be marked as sample requirement text, and the unlabeled and uncalibrated presentation corresponding to the sample requirement text can be marked as the sample initial presentation corresponding to the sample requirement text.
[0122] Optionally, the initial sample presentation can be calibrated and annotated based on a preset calibration and annotation method to match the preset presentation conditions, and the initial sample presentation that meets the preset presentation conditions can be marked as a candidate sample presentation.
[0123] S403, Obtain the sample presentation building elements of the sample candidate presentations to generate training samples for the candidate large model.
[0124] In this embodiment of the disclosure, the sample candidate presentation can be traversed to obtain the presentation building elements that build the sample candidate presentation, and these elements can be marked as sample presentation building elements.
[0125] In this scenario, the sample requirement text and sample presentation building elements can be processed using sample generation algorithms in related technologies. Then, training samples composed of sample requirement text and sample presentation building elements can be obtained based on the results of the algorithm processing, and these training samples can be marked as training samples for candidate large models.
[0126] S404 optimizes the candidate large model by using training samples until the training is completed, and obtains the optimized target large model.
[0127] The target large model is used to achieve the above. Figure 1 The presentation method based on a large model proposed in the embodiments shown in the figure.
[0128] In this embodiment of the disclosure, training samples can be input into candidate large models for model optimization training. For any round of model training, when the large model after the end of the training round meets the preset model training termination condition, the model training optimization of the candidate large model can be terminated, and the large model obtained after the end of the training round is determined as the optimized target large model.
[0129] Optionally, corresponding training termination conditions can be set according to the training rounds. For any round of model training, when the total number of training rounds of the candidate large model meets the preset training termination conditions after the training of that round, the model training of the candidate large model can be terminated, and the large model obtained after the training of that round is determined as the target large model.
[0130] Optionally, a corresponding training termination condition can be set according to the training output. For any round of model training, when the output of the candidate large model meets the preset training termination condition after the training of that round, the training of the candidate large model can be terminated, and the large model obtained after the training of that round can be determined as the target large model.
[0131] The large-scale model training method proposed in this disclosure involves acquiring sample requirement text and initial sample presentations, annotating and calibrating the initial sample presentations to obtain candidate sample presentations, and obtaining training samples based on the sample requirement text and candidate sample presentations. The candidate large-scale model is then trained using these training samples until training is complete, resulting in an optimized target large-scale model. In this disclosure, the training samples for the candidate large-scale model are obtained through the sample presentation building elements of the candidate sample presentations. This allows the candidate large-scale model to learn the relationship between the sample requirement text and the sample presentation building elements, thereby enabling the trained target large-scale model to learn how to arrange the sample presentation building elements to obtain the corresponding presentation. This optimizes the learning effect of the large-scale model and improves the quality and effectiveness of the presentations obtained based on the target large-scale model.
[0132] In the above embodiments, the acquisition of the target large model can also be combined with... Figure 5 To understand further, Figure 5 This is a flowchart illustrating a training method for a large model according to another embodiment of this disclosure, as shown below. Figure 5 As shown, the method includes:
[0133] S501, Obtain candidate large models to be optimized.
[0134] For details regarding step S501, please refer to the relevant information in the above embodiments, which will not be repeated here.
[0135] S502, obtain the sample requirement text and the corresponding initial sample presentation text from the sample user terminal, and calibrate the initial sample presentation text to obtain the sample candidate presentation text from the sample user terminal.
[0136] Optionally, it is determined whether the initial sample presentation meets the preset presentation conditions. In response to the determination that the initial sample presentation does not meet the preset presentation conditions, abnormal presentation elements of the initial sample presentation are obtained and the abnormal presentation elements are calibrated to calibrate the initial sample presentation and obtain the sample candidate presentation for the sample user.
[0137] In this embodiment of the disclosure, the initial sample presentation can be evaluated and reviewed based on preset presentation conditions, thereby identifying whether the initial sample presentation meets the preset presentation conditions.
[0138] Presentation requirements may include at least whether the referenced images match the text content, whether the page layout is reasonable, and whether the referenced images are accurate.
[0139] In this scenario, the detailed conditions in the presentation conditions can be compared one by one with the corresponding elements of the initial sample presentation. When there are elements in the initial sample presentation that do not match the presentation conditions, it can be determined that the initial sample presentation does not meet the preset presentation conditions.
[0140] Optionally, sample initial presentations that do not meet the preset presentation conditions can be calibrated, wherein elements in the sample initial presentations that do not match the presentation conditions can be marked as abnormal presentation elements in the sample initial presentations.
[0141] In this scenario, the presentation conditions that the abnormal presentation element needs to match can be obtained, and the abnormal presentation element can be calibrated according to the presentation conditions so that the abnormal presentation element can match the presentation conditions, thereby calibrating the initial sample presentation and marking the calibrated initial sample presentation as a candidate sample presentation.
[0142] As an example, when the abnormal presentation element is an inaccurate page division, calibration operations such as merging or splitting the inaccurately divided pages in the initial presentation can be performed to calibrate the sample initial presentation.
[0143] As another example, when the abnormal presentation element is an image element, in response to the abnormal presentation element being an abnormal image element, a corresponding calibration query vector is constructed to obtain the recall image element corresponding to the calibration query vector, and the abnormal image element is updated based on the recall image element to calibrate the abnormal image element.
[0144] In this embodiment of the disclosure, image elements that are abnormal presentation elements can be marked as abnormal image elements. In this scenario, a corresponding query vector can be constructed based on the abnormal image element and marked as the calibration query vector of the abnormal image element.
[0145] In this scenario, corresponding image elements from a preset image database can be retrieved based on the calibration query vector, and the retrieved image elements can be marked as recalled image elements of abnormal image elements.
[0146] Furthermore, the recalled image element is overlaid on the abnormal image element, or the recalled image element is placed in the corresponding position on the page to which the abnormal image element belongs, so as to calibrate the abnormal image element on the presentation page, thereby calibrating the initial presentation.
[0147] S503, Obtain the sample presentation building elements of the sample candidate presentations to generate training samples for the candidate large model.
[0148] Optionally, the sample candidate presentations are divided into pages to generate a set of sample presentation pages for the sample candidate presentations. Then, each sample presentation page in the set of sample presentation pages is traversed to obtain the sample candidate presentation elements in each sample presentation page.
[0149] In this embodiment of the disclosure, the sample candidate presentation includes presentation page information. The sample candidate presentation can be divided into pages, and the set of multiple pages obtained from the division is labeled as the sample presentation page set.
[0150] In this scenario, the sample presentation pages in the sample presentation page collection can be traversed to extract the element information that constitutes each sample presentation page, which may include image elements, line elements, symbol elements, text box elements, and shape elements, etc.
[0151] In this scenario, the building elements of each sample presentation page obtained through iteration can be marked as sample candidate presentation elements for each sample presentation page.
[0152] Optionally, the sample element description tag and sample element layout tag of the sample candidate presentation element are obtained. Specifically, for any sample presentation page, the sample candidate page main tag of the sample presentation page is obtained, and the sample candidate presentation element, sample element description tag and sample element layout tag are filled into the sample candidate page main tag to obtain the sample target page main tag.
[0153] In this embodiment of the disclosure, the tags generated from the appearance description information of the sample candidate presentation elements can be marked as sample element description tags of the sample presentation building elements. The appearance effect of the corresponding sample presentation building elements can be determined based on the sample element description tags.
[0154] Additionally, it can obtain the position and layout information of the sample presentation building elements in the corresponding sample presentation page, and determine the tags generated based on the position and layout information as the sample element layout tags of the sample candidate presentation elements.
[0155] Optionally, for any sample presentation page, a main tag can be constructed for the sample presentation page based on the main tag (tag) construction method in related technologies, thereby obtaining the main tag of the sample presentation page and marking it as the main tag of the sample candidate page of the sample presentation page.
[0156] In this scenario, for any sample presentation page, the sample element description tags and sample element layout tags of each sample presentation building element in the sample presentation page can be filled into the sample candidate page main tag of the sample presentation page, and the tag obtained after filling is determined as the sample target page main tag of the sample presentation page.
[0157] It should be noted that for any sample presentation page, the sample element layout tags of each sample presentation building element include tag information composed of element layout information shared by each sample presentation building element, as well as tag information composed of element layout information unique to each sample presentation building element.
[0158] Optionally, sample presentation tags for the sample requirement text can be obtained based on the sample target page main tags of each sample presentation page, in order to obtain training samples for the candidate large model.
[0159] In this embodiment of the disclosure, the main tags of the target page of each sample presentation page can be integrated, and the integrated tag information can be used as the training sample tags of the sample requirement text and marked as the sample presentation tags of the sample requirement text.
[0160] Furthermore, based on the sample construction algorithm in related technologies, the sample requirement text and sample presentation labels are processed by the algorithm, and then the training labels of the candidate large model are obtained based on the results of the algorithm processing.
[0161] As an example, based on data conversion algorithms in related technologies, sample candidate presentations can be converted into data in a data format (JSON), and the sample candidate presentations in JSON format can be divided into pages, generating tags for each sample presentation page.
[0162] Furthermore, HTML tags are mapped to the sample presentation pages in JSON format to obtain the sample candidate presentation elements in HTML format for each sample presentation page, as well as the sample description tags for each sample candidate presentation element in HTML format. Tags and / or Label.
[0163] Additionally, it can obtain the style.css code snippet in HTML format, which is constructed based on the shared layout information of each sample candidate presentation element. This snippet is the shared layout tag information of each sample candidate presentation element, and the code snippet in HTML format, which is the exclusive layout tag information of each sample candidate presentation element, thereby obtaining the sample element layout tags of the sample presentation page.
[0164] In this example, for any given sample presentation page, the candidate sample presentation elements within that page can be... Tags and / or The tags, along with the sample element layout tags, are wrapped within the tags of the sample presentation page to obtain the sample target page main tags of the sample presentation page.
[0165] In this scenario, the main tags of the sample target page for each sample presentation page are code snippets in HTML format. The code snippets of the main tags of each sample target page can be integrated according to the code integration methods in related technologies, and the integrated code snippets can be used as the sample presentation tags of the sample requirement text.
[0166] S504 optimizes and trains the candidate large model using training samples until the training is complete, resulting in the optimized target large model.
[0167] Optionally, the training samples are input into the candidate large model to obtain the predicted presentation output by the candidate large model based on the sample demand text in the training samples, and the loss value of the predicted presentation based on the sample presentation label is obtained. The candidate large model is then adjusted according to the loss value, and the next sample demand text is obtained to continue to optimize and train the adjusted candidate large model until the training ends, and the optimized target large model is obtained.
[0168] In this embodiment of the disclosure, training samples can be input into a candidate large model, and the candidate large model can generate a presentation document corresponding to the sample requirement text as the output predicted presentation document.
[0169] Furthermore, based on the loss value acquisition algorithm in related technologies, the predicted presentation and the sample presentation labels are processed to obtain the loss value of the predicted presentation based on the sample presentation labels, and this loss value is used as the training loss of the candidate large model.
[0170] Optionally, the model parameters of the candidate large model are adjusted according to the training loss, and the candidate large model with adjusted parameters is returned to continue training until the training ends, and the optimized target large model after training is obtained.
[0171] The proposed large-scale model training method involves acquiring sample requirement text and initial sample presentations, annotating and calibrating the initial sample presentations to obtain candidate sample presentations, and then obtaining training samples based on the sample requirement text and candidate sample presentations. The candidate large-scale model is then trained on these training samples until training is complete, resulting in an optimized target large-scale model. In this method, the sample presentation tags are HTML code snippets, reducing the difficulty of acquiring these tags and increasing the richness and quantity of training samples. The candidate large-scale model can output HTML code snippets for predicted presentations, enabling it to utilize the layout capabilities inherent in HTML format. This reduces the rendering difficulty of the predicted presentation code snippets and improves their rendering effect, thereby optimizing the training performance of the candidate large-scale model.
[0172] To better understand the above embodiments, the following comparative examples can be used as a reference:
[0173] As an example, such as Figure 6 As shown, Figure 6 The image shown is a presentation page obtained using a presentation generation method based on related technologies. Figure 7 The presentation page is obtained based on the large model-based presentation generation method proposed in the embodiments of this disclosure.
[0174] in, Figure 6 The image shown is Figure 6 The displayed text content is not compatible, making Figure 6 The presentation slides shown are not up to par.
[0175] like Figure 7 As shown, you can choose based on Figure 7 The presentation avoids using images, relying solely on textual information to display information about changes in population growth rates among different groups, thus preventing... Figure 6 The presentation was unsatisfactory due to a mismatch between the images and text content.
[0176] As another example, such as Figure 8 and Figure 9 As shown, Figure 8 The presentation slides shown demonstrate references to relevant technologies. Figure 8 The images referenced in the displayed presentation slides did not match the cited references well, resulting in a subpar presentation.
[0177] In this scenario, the following can be adopted: Figure 9 The presentation page shown is obtained based on the large model-based presentation generation method proposed in the embodiments of this disclosure. Only the text content of the cited document titles is displayed to avoid... Figure 8 The situation shown occurs.
[0178] An embodiment of this disclosure also proposes a presentation generation device based on a large model. Since the presentation generation device based on a large model proposed in this disclosure corresponds to the presentation generation method based on a large model proposed in the above embodiments, the implementation methods of the above-mentioned presentation generation methods based on a large model are also applicable to the presentation generation device based on a large model proposed in this disclosure. It will not be described in detail in the following embodiments.
[0179] Figure 10 This is a schematic diagram of the structure of a large-scale model presentation generation apparatus according to an embodiment of the present disclosure, as shown below. Figure 10 As shown, the presentation generation device 1000 based on a large model includes a first acquisition module 1001, a second acquisition module 1002, and a first generation module 1003, wherein:
[0180] The first acquisition module 1001 is used to acquire the presentation request text from the user's end.
[0181] The second acquisition module 1002 is used to input the presentation requirement text into the large model, and obtain the associated set of presentation building elements based on the presentation requirement text through the model capabilities of the large model.
[0182] The first generation module 1003 is used to perform page layout on the presentation building element set through the model capabilities of the large model, and generate the target presentation corresponding to the presentation requirement text.
[0183] In this embodiment of the disclosure, the second acquisition module 1002 is further configured to: segment the presentation requirement text using the model capabilities of the large model to obtain multiple segmented requirement sub-texts; traverse the multiple requirement sub-texts using the model capabilities of the large model to obtain a subset of presentation building elements for each requirement sub-text; and obtain a set of presentation building elements associated with the presentation requirement text based on the subset of presentation building elements for each requirement sub-text.
[0184] In this embodiment of the disclosure, the second acquisition module 1002 is further configured to: acquire a set of text titles in the presentation requirement text and the title level of each text title in the text title set. Using the model capabilities of the large model, the presentation requirement text is segmented according to the title level to obtain multiple segmented requirement sub-texts.
[0185] In this embodiment of the disclosure, the second acquisition module 1002 is further configured to: for any requirement sub-text, acquire the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements associated with the requirement sub-text through the model capabilities of the large model. Based on the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements, obtain a subset of presentation building elements of the requirement sub-text.
[0186] In this embodiment of the disclosure, the first generation module 1003 is further configured to: for any requirement subtext, obtain a subset of presentation building elements corresponding to the requirement subtext in the presentation building element set; obtain the element description tags of each presentation building element in the subset of presentation building elements, and obtain the target page layout file of the subset of presentation building elements; utilize the model capabilities of the large model to lay out each presentation building element according to the element description tags and the target page layout file, generating a presentation subpage of the requirement subtext; and obtain the target presentation of the presentation requirement text based on the presentation subpages of each requirement subtext.
[0187] In this embodiment of the disclosure, the first generation module 1003 is further configured to: for any requirement subtext, obtain candidate block tags for each presentation building element of the requirement subtext; for any presentation building element, obtain element description information of the presentation building element, and fill the element description information into the candidate block tags to generate element description tags for the presentation building element.
[0188] In this embodiment of the disclosure, the first generation module 1003 is further configured to: for any requirement subtext, obtain the element layout information set of each subset of presentation building elements of the requirement subtext; generate a first candidate layout file and a second candidate layout file for the requirement subtext based on the element layout information set; and obtain the target page layout file of the subset of presentation building elements of the requirement subtext based on the first candidate layout file and the second candidate layout file.
[0189] In this embodiment of the disclosure, the first generation module 1003 is further configured to: obtain a subset of presentation building elements for each of the multiple requirement sub-files; for any requirement sub-text, obtain exclusive element layout information (excluding shared element layout information) from the element layout information set of the presentation building elements subset of the requirement sub-text for any presentation building element in the presentation building element subset, and generate a second candidate layout file for the requirement sub-text based on the exclusive element layout information.
[0190] In this embodiment of the disclosure, the first generation module 1003 is further configured to: obtain the presentation subpage code of each requirement subtext through the model capability of the large model, so as to obtain the target presentation code of the presentation requirement text; and render the target presentation code to obtain the target presentation.
[0191] This disclosure presents a presentation generation device based on a large model. It obtains the presentation requirement text from the user, acquires the set of presentation building elements associated with the requirement text through the modeling capabilities of the large model, and performs page layout on the set of presentation building elements using the modeling capabilities of the large model to obtain the target presentation. In this disclosure, generating the target presentation through a large model improves the efficiency and quality of presentation generation. The user only needs to input the presentation requirement text, reducing the operational difficulty for the user. The target presentation is obtained by performing page layout on the set of presentation building elements through the modeling capabilities of the large model. Compared to related technologies that create presentations by calling existing presentation building templates, this avoids the limitations imposed by presentation templates on the presentation generation method, improves the flexibility and practicality of the presentation generation method, and optimizes the user experience.
[0192] Corresponding to the large model training methods proposed in the above embodiments, an embodiment of this disclosure also proposes a large model training device. Since the large model training device proposed in this disclosure corresponds to the large model training methods proposed in the above embodiments, the implementation methods of the above large model training methods are also applicable to the large model training device proposed in this disclosure, and will not be described in detail in the following embodiments.
[0193] Figure 11 This is a schematic diagram of the structure of a large model training device according to an embodiment of the present disclosure, as shown below. Figure 11 As shown, the training device 1100 for the large model includes a third acquisition module 111, a fourth acquisition module 112, a second generation module 113, and a training module 114, wherein:
[0194] The third acquisition module 111 is used to acquire candidate large models to be optimized.
[0195] The fourth acquisition module 112 is used to acquire the sample requirement text and the sample initial presentation text corresponding to the sample requirement text from the sample user terminal, and to calibrate the sample initial presentation text to obtain the sample candidate presentation text from the sample user terminal.
[0196] The second generation module 113 is used to obtain the sample presentation building elements of the sample candidate presentation to generate training samples for the candidate large model.
[0197] Training module 114 is used to optimize and train the candidate large model using training samples until training is complete, obtaining the optimized target large model, wherein the target large model is used to achieve the above. Figure 10 The embodiment proposes a presentation generation device based on a large model.
[0198] In this embodiment of the disclosure, the fourth acquisition module 112 is further configured to: identify whether the initial sample presentation meets preset presentation conditions. In response to identifying that the initial sample presentation does not meet the preset presentation conditions, the module acquires abnormal presentation elements of the initial sample presentation and calibrates the abnormal presentation elements to calibrate the initial sample presentation and obtain a sample candidate presentation for the sample user terminal.
[0199] In this embodiment of the present disclosure, the fourth acquisition module 112 is further configured to: in response to the abnormal presentation element being an abnormal image element, construct a corresponding calibration query vector to obtain the recall image element corresponding to the calibration query vector, update the abnormal image element based on the recall image element, and calibrate the abnormal image element.
[0200] In this embodiment of the disclosure, the second generation module 113 is further configured to: divide the candidate presentations into pages to generate a set of sample presentation pages for the candidate presentations; traverse each sample presentation page in the set of sample presentation pages to obtain the candidate presentation elements in each sample presentation page; obtain the sample element description tags and sample element layout tags of the candidate presentation elements; for any sample presentation page, obtain the candidate page main tag of the sample presentation page, and fill the candidate presentation elements, sample element description tags, and sample element layout tags into the candidate page main tag to obtain the target page main tag; and obtain the sample presentation tags of the sample requirement text based on the target page main tags of each sample presentation page to obtain the training samples for the candidate large model.
[0201] In this embodiment of the disclosure, the training module 114 is further configured to: input training samples into the candidate large model to obtain a predicted presentation output by the candidate large model based on the sample requirement text in the training samples; obtain the loss value of the predicted presentation based on the sample presentation tags, adjust the candidate large model according to the loss value, and return to obtain the next sample requirement text to continue optimizing the adjusted candidate large model until the training is completed, thereby obtaining the optimized target large model.
[0202] The large-scale model training device disclosed herein acquires sample requirement text and initial sample presentations, annotates and calibrates the initial sample presentations to obtain candidate sample presentations, and obtains training samples based on the sample requirement text and candidate sample presentations. The candidate large-scale model is then trained based on the training samples until training is complete, resulting in an optimized target large-scale model. In this disclosure, the training samples for the candidate large-scale model are obtained through the sample presentation building elements of the candidate sample presentations. This allows the candidate large-scale model to learn the relationship between the sample requirement text and the sample presentation building elements, thereby enabling the trained target large-scale model to learn how to arrange the sample presentation building elements to obtain the corresponding presentation. This optimizes the learning effect of the large-scale model and improves the quality and effectiveness of the presentations obtained based on the target large-scale model.
[0203] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0204] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0205] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1202 or a computer program loaded from storage unit 1209 into random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.
[0206] Multiple components in device 1200 are connected to I / O interface 1205, including: input unit 1206, such as keyboard, mouse, etc.; output unit 1206, such as various types of monitors, speakers, etc.; storage unit 1209, such as disk, optical disk, etc.; and communication unit 1209, such as network card, modem, wireless transceiver, etc. The communication unit 1209 allows device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0207] The computing unit 1201 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as large model-based presentation generation methods and / or large model training methods. For example, in some embodiments, the large model-based presentation generation methods and / or large model training methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1209. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by computing unit 1201, one or more steps of the large model-based presentation generation method and / or large model training method described above can be performed. Alternatively, in other embodiments, computing unit 1201 can be configured to perform the large model-based presentation generation method and / or large model training method by any other suitable means (e.g., by means of firmware).
[0208] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0209] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0210] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0211] To initiate interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user account can submit input to the computer. Other types of devices can also be used to initiate interaction with the user account; for example, feedback submitted to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including voice input, speech input, or tactile input).
[0212] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user account computer with a graphical user interface or web browser through which a user account can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0213] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0214] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0215] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A presentation generation method based on a large model, wherein, The method includes: Obtain the presentation request text from the user's client; The presentation requirement text is input into the large model, and the model capability of the large model is used to segment the presentation requirement text to obtain multiple segmented requirement sub-texts. By leveraging the model capabilities of the large model, the multiple requirement sub-texts are traversed to obtain a subset of presentation building elements for each requirement sub-text. Based on the subset of presentation building elements for each requirement sub-text, an associated set of presentation building elements is obtained, wherein the subset of presentation building elements includes shape building elements, text box building elements, image building elements, symbol building elements, and line building elements. Using the modeling capabilities of the large model, the page layout of the presentation building element set is performed to generate the target presentation corresponding to the presentation requirement text.
2. The method according to claim 1, wherein, The presentation requirement text is segmented using the capabilities of the large model to obtain multiple segmented requirement sub-texts, including: Obtain the set of text titles in the presentation document requirement text and the title level of each text title in the set of text titles; Using the capabilities of the large model, the presentation requirement text is segmented according to the title level to obtain multiple segmented requirement sub-texts.
3. The method according to claim 1, wherein, The process of using the model capabilities of the large model to traverse the multiple requirement sub-texts and obtain a subset of presentation building elements for the requirement sub-texts includes: For any requirement sub-text, the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements associated with the requirement sub-text are obtained through the model capabilities of the large model; Based on the shape building element, the text box building element, the image building element, the symbol building element, and the line building element, the presentation building element subset of the requirement subtext is obtained.
4. The method according to claim 1, wherein, The process of using the model capabilities of the large model to perform page layout on the presentation building element set and generate the target presentation corresponding to the presentation requirements includes: For any requirement subtext, obtain the subset of presentation building elements corresponding to the requirement subtext in the presentation building element set; Obtain the element description tags of each presentation building element in the presentation building element subset, and obtain the target page layout file of the presentation building element subset; Using the model capabilities of the large model, the presentation building elements are laid out according to the element description tags and the target page layout file to generate the presentation subpage of the requirement subtext; Based on the presentation subpages of each requirement subtext, the target presentation of the presentation requirement text is obtained.
5. The method according to claim 4, wherein, The step of obtaining the element description tags of each presentation building element in the subset of presentation building elements includes: For any given requirement subtext, obtain the candidate block tags for each presentation building element of the requirement subtext; For any presentation building element, obtain the element description information of the presentation building element, and fill the element description information into the candidate block tag to generate the element description tag of the presentation building element.
6. The method according to claim 4, wherein, The step of obtaining the target page layout file for the subset of elements in the presentation includes: For any requirement subtext, obtain the element layout information set of each of the presentation building element subsets of the requirement subtext; Based on the set of element layout information, generate a first candidate layout file and a second candidate layout file for the requirement sub-text; Based on the first candidate layout file and the second candidate layout file, the target page layout file of the presentation building element subset of the requirement subtext is obtained.
7. The method according to claim 6, wherein, The step of generating a first candidate layout file and a second candidate layout file for the required sub-text based on the element layout information set includes: Obtain the presentation build element subsets of each of the multiple requirement sub-files. For any requirement subtext, shared element layout information is extracted from the element layout information set of each of the presentation building element subsets of the requirement subtext to generate the first candidate layout file of the requirement subtext. For any presentation building element in the subset of presentation building elements, obtain exclusive element layout information other than the shared element layout information from the element layout information set of the presentation building element, and generate the second candidate layout file of the requirement subtext based on the exclusive element layout information.
8. The method according to claim 4, wherein, The process of obtaining the target presentation document of the presentation requirement text based on the presentation document subpages of each requirement subtext includes: By leveraging the modeling capabilities of the large model, the presentation subpage code of each requirement subtext is obtained, thereby yielding the target presentation code for the aforementioned presentation requirement text. The target presentation code is rendered to obtain the target presentation.
9. A training method for a large model, wherein, The method includes: Obtain candidate large models to be optimized; Obtain the sample requirement text and the corresponding initial sample presentation, and calibrate the initial sample presentation to obtain the sample candidate presentation. Obtain the sample presentation building elements of the candidate presentation to generate training samples for the candidate large model, wherein the sample presentation building elements include image elements, line elements, symbol elements, text box elements, and shape elements. The candidate large model is optimized and trained using the training samples until the training is completed, resulting in an optimized target large model. The target large model is used to implement the presentation generation method based on the large model as described in any one of claims 1-8.
10. The method according to claim 9, wherein, The process of obtaining the sample requirement text from the sample user client and the corresponding initial sample presentation, and calibrating the initial sample presentation to obtain the sample candidate presentation for the sample user client, includes: Identify whether the initial presentation of the sample meets the preset presentation conditions; In response to the identification that the initial sample presentation does not meet the preset presentation conditions, abnormal presentation elements of the initial sample presentation are obtained and the abnormal presentation elements are calibrated to calibrate the initial sample presentation and obtain the sample candidate presentation for the sample user.
11. The method according to claim 10, wherein, The method further includes: In response to the abnormal presentation element being an abnormal image element, a corresponding calibration query vector is constructed to obtain the recalled image element corresponding to the calibration query vector, and the abnormal image element is updated based on the recalled image element to calibrate the abnormal image element.
12. The method according to claim 9, wherein, The step of obtaining the sample presentation building elements of the candidate presentations to generate training samples for the candidate large model includes: The candidate sample presentations are divided into pages to generate a set of sample presentation pages for the candidate sample presentations. Traverse each sample presentation page in the sample presentation page set to obtain the sample candidate presentation elements in each sample presentation page; Obtain the sample element description tag and sample element layout tag of the sample candidate presentation element; For any sample presentation page, obtain the sample candidate page main tag of the sample presentation page, and fill the sample candidate presentation element, the sample element description tag and the sample element layout tag into the sample candidate page main tag to obtain the sample target page main tag; The sample presentation tags of the sample requirement text are obtained based on the sample target page main tags of each sample presentation page, so as to obtain the training samples of the candidate large model.
13. The method according to claim 9, wherein, The step of optimizing and training the candidate large model using the training samples until training is complete, to obtain the optimized target large model, includes: The training samples are input into the candidate large model to obtain the predicted presentation output by the candidate large model based on the sample demand text in the training samples; Obtain the loss value of the predicted presentation based on the sample presentation labels, adjust the candidate large model according to the loss value, and return to obtain the next sample requirement text to continue optimizing and training the adjusted candidate large model until the training ends, and obtain the optimized target large model.
14. A presentation generation device based on a large model, wherein, The device includes: The first acquisition module is used to acquire the presentation request text from the user's end; The second acquisition module is used to input the presentation requirement text into a large model, segment the presentation requirement text using the model capabilities of the large model to obtain multiple segmented requirement sub-texts; traverse the multiple requirement sub-texts using the model capabilities of the large model to obtain a subset of presentation building elements for each requirement sub-text, and obtain an associated set of presentation building elements based on the subset of presentation building elements for each requirement sub-text, wherein the subset of presentation building elements includes shape building elements, text box building elements, image building elements, symbol building elements, and line building elements; The first generation module is used to perform page layout on the presentation building element set using the model capabilities of the large model, and generate the target presentation corresponding to the presentation requirement text.
15. The apparatus according to claim 14, wherein, The second acquisition module is further configured to: Obtain the set of text titles in the presentation document requirement text and the title level of each text title in the set of text titles; Using the capabilities of the large model, the presentation requirement text is segmented according to the title level to obtain multiple segmented requirement sub-texts.
16. The apparatus according to claim 14, wherein, The second acquisition module is further configured to: For any requirement sub-text, the shape building elements, text box building elements, image building elements, symbol building elements, and line building elements associated with the requirement sub-text are obtained through the model capabilities of the large model; Based on the shape building element, the text box building element, the image building element, the symbol building element, and the line building element, the presentation building element subset of the requirement subtext is obtained.
17. The apparatus according to claim 14, wherein, The first generation module is further configured to: For any requirement subtext, obtain the subset of presentation building elements corresponding to the requirement subtext in the presentation building element set; Obtain the element description tags of each presentation building element in the presentation building element subset, and obtain the target page layout file of the presentation building element subset; Using the model capabilities of the large model, the presentation building elements are laid out according to the element description tags and the target page layout file to generate the presentation subpage of the requirement subtext; Based on the presentation subpages of each requirement subtext, the target presentation of the presentation requirement text is obtained.
18. The apparatus according to claim 17, wherein, The first generation module is further configured to: For any given requirement subtext, obtain the candidate block tags for each presentation building element of the requirement subtext; For any presentation building element, obtain the element description information of the presentation building element, and fill the element description information into the candidate block tag to generate the element description tag of the presentation building element.
19. The apparatus according to claim 17, wherein, The first generation module is further configured to: For any requirement subtext, obtain the element layout information set of each of the presentation building element subsets of the requirement subtext; Based on the set of element layout information, generate a first candidate layout file and a second candidate layout file for the requirement sub-text; Based on the first candidate layout file and the second candidate layout file, the target page layout file of the presentation building element subset of the requirement subtext is obtained.
20. The apparatus according to claim 19, wherein, The first generation module is further configured to: Obtain the presentation build element subsets of each of the multiple requirement sub-files; For any requirement subtext, shared element layout information is extracted from the element layout information set of each of the presentation building element subsets of the requirement subtext to generate the first candidate layout file of the requirement subtext. For any presentation building element in the subset of presentation building elements, obtain exclusive element layout information other than the shared element layout information from the element layout information set of the presentation building element, and generate the second candidate layout file of the requirement subtext based on the exclusive element layout information.
21. The apparatus according to claim 14, wherein, The first generation module is further configured to: By leveraging the modeling capabilities of the large model, the presentation subpage code of each requirement subtext is obtained, thereby yielding the target presentation code for the aforementioned presentation requirement text. The target presentation code is rendered to obtain the target presentation.
22. A training device for a large model, wherein, The device includes: The third acquisition module is used to acquire candidate large models to be optimized; The fourth acquisition module is used to acquire the sample requirement text of the sample user terminal and the sample initial presentation text corresponding to the sample requirement text, and to calibrate the sample initial presentation text to obtain the sample candidate presentation text of the sample user terminal. The second generation module is used to obtain the sample presentation building elements of the sample candidate presentation to generate training samples of the candidate large model, wherein the sample presentation building elements include image elements, line elements, symbol elements, text box elements and shape elements. The training module is used to optimize and train the candidate large model using the training samples until the training is completed, and to obtain the optimized target large model, wherein the target large model is used to implement the presentation generation device based on the large model as described in any one of claims 14-21.
23. The apparatus according to claim 22, wherein, The fourth acquisition module is also used for: Identify whether the initial presentation of the sample meets the preset presentation conditions; In response to the identification that the initial sample presentation does not meet the preset presentation conditions, abnormal presentation elements of the initial sample presentation are obtained and the abnormal presentation elements are calibrated to calibrate the initial sample presentation and obtain the sample candidate presentation for the sample user.
24. The apparatus according to claim 23, wherein, The fourth acquisition module is also used for: In response to the abnormal presentation element being an abnormal image element, a corresponding calibration query vector is constructed to obtain the recalled image element corresponding to the calibration query vector, and the abnormal image element is updated based on the recalled image element to calibrate the abnormal image element.
25. The apparatus according to claim 22, wherein, The second generation module is further configured to: The candidate presentations are divided into pages to generate a sample presentation page set of the sample candidate presentations; Traverse each sample presentation page in the sample presentation page set to obtain the sample candidate presentation elements in each sample presentation page; Obtain the sample element description tag and sample element layout tag of the sample candidate presentation element; For any sample presentation page, obtain the sample candidate page main tag of the sample presentation page, and fill the sample candidate presentation element, the sample element description tag and the sample element layout tag into the sample candidate page main tag to obtain the sample target page main tag; The sample presentation tags of the sample requirement text are obtained based on the sample target page main tags of each sample presentation page, so as to obtain the training samples of the candidate large model.
26. The apparatus according to claim 25, wherein, The training module is also used for: The training samples are input into the candidate large model to obtain the predicted presentation output by the candidate large model based on the sample demand text in the training samples; Obtain the loss value of the predicted presentation based on the sample presentation labels, adjust the candidate large model according to the loss value, and return to obtain the next sample requirement text to continue optimizing and training the adjusted candidate large model until the training ends, and obtain the optimized target large model.
27. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8 and / or claims 9-13.
28. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8 and / or claims 9-13.
29. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8 and / or claims 9-13.
Citation Information
Patent Citations
Powerpoint generation method and device
CN116579308A
Document image intelligent analysis and processing method based on multi-modal information
CN117173730A