Method and device for optimizing presentation file, method and device for finely adjusting large model, equipment and medium
Through large-scale fine-tuning technology, the page image classification and analysis of presentations are optimized, and the problem of poor optimization of presentations in the existing technology is solved, achieving more efficient display effects and user experience.
Patent Information
- Application Number
- CN202510053379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively optimize presentations, especially in the classification and analysis of page images, resulting in poor presentation and user experience of presentations.
Through the big model fine-tuning method, the training data set is obtained and the page image classification model and analysis model iteratively trains the page image classification model and the analysis model are determined to determine the page analysis results of the candidate page images, thereby optimizing the presentation.
It realizes efficient optimization of presentations, improves the quality and user experience of page display, and provides diverse design needs support for different users.
Smart Images

Figure CN119991877A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, in particular to the field of artificial intelligence and large model technology, and specifically to a method, device, equipment and medium for optimizing a presentation and fine-tuning a large model. Background Art
[0002] With the continuous development of computer technology, presentations (Power Point, PPT) have been widely used in conferences, lectures, teaching and other occasions; presentations can combine multiple information forms such as text, pictures, charts, etc., and present them in a multimedia way, which can help speakers present information and ideas more intuitively and vividly. Summary of the invention
[0003] The present disclosure provides a method, device, equipment and medium for optimizing presentations and fine-tuning large models.
[0004] According to one aspect of the present disclosure, a method for optimizing a presentation is provided, comprising:
[0005] determining candidate page images of the presentation to be optimized;
[0006] Determining a page analysis result of the candidate page image, and optimizing the target page image when the candidate page image is determined to be a target page image according to the page analysis result;
[0007] In response to the instruction to generate the page optimization results of all the target page images of the presentation, a target recommendation template is determined, and a target presentation is generated based on the target recommendation template and each of the page optimization results.
[0008] According to another aspect of the present disclosure, a method for fine-tuning a large model for page image classification is provided, comprising:
[0009] Acquire a first training data set; the first training data set includes page images of a plurality of presentations and classification results of each of the page images;
[0010] Determine a first model prompt word, and input the first model prompt word and the first training data set into a first large model for iterative training;
[0011] When the iteration stop condition is met, a large model for page image classification is obtained;
[0012] The page image classification model is used to determine the page analysis results of the candidate page images.
[0013] According to another aspect of the present disclosure, a method for fine-tuning a large model of page image parsing is provided, comprising:
[0014] Acquire a second training data set; the second training data set includes page images of a plurality of presentations and content analysis results of each of the page images;
[0015] Determine a second model prompt word, and input the second model prompt word and the second training data set into the second largest model for iterative training;
[0016] When the iteration stop condition is met, a large model of page image parsing is obtained;
[0017] The page image parsing large model is used to determine the page analysis result of the candidate page image.
[0018] According to another aspect of the present disclosure, there is provided a presentation optimization device, comprising:
[0019] A candidate page image determination module, used to determine the candidate page images of the presentation to be optimized;
[0020] a page analysis result determination module, configured to determine a page analysis result of the candidate page image, and optimize the target page image when the candidate page image is determined to be a target page image according to the page analysis result;
[0021] The target presentation generation module is used to determine a target recommendation template in response to a generation instruction of page optimization results of all target page images of the presentation, and to generate a target presentation based on the target recommendation template and each of the page optimization results.
[0022] According to another aspect of the present disclosure, a fine-tuning device for a large model of page image classification is provided, comprising:
[0023] A first training data set acquisition module, used to acquire a first training data set; the first training data set includes page images of a plurality of presentations and classification results of each of the page images;
[0024] An iterative training module, used for determining a first model prompt word, and inputting the first model prompt word and the first training data set into a first large model for iterative training;
[0025] A large model generation module is used to obtain a large model for page image classification when an iteration stop condition is met;
[0026] The page image classification model is used to determine the page analysis results of the candidate page images.
[0027] According to another aspect of the present disclosure, there is provided a fine-tuning device for a large model of page image parsing, comprising:
[0028] A second training data set acquisition module, used to acquire a second training data set; the second training data set includes page images of a plurality of presentations and content analysis results of each of the page images;
[0029] An iterative training module, used for determining a second model prompt word, and inputting the second model prompt word and the second training data set into a second large model for iterative training;
[0030] A large model generation module is used to obtain a large model of page image analysis when an iteration stop condition is met;
[0031] The page image parsing large model is used to determine the page analysis result of the candidate page image.
[0032] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device comprising:
[0033] at least one processor; and
[0034] a memory communicatively connected to the at least one processor; wherein,
[0035] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by any embodiment of the present disclosure.
[0036] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided by any embodiment of the present disclosure.
[0037] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method provided according to any embodiment of the present disclosure is implemented.
[0038] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0040] Figure 1 is a flow chart of a method for optimizing a presentation provided according to an embodiment of the present disclosure;
[0041] Figure 2is a flow chart of another method for optimizing a presentation provided according to an embodiment of the present disclosure;
[0042] Figure 3 is a flow chart of another method for optimizing a presentation provided according to an embodiment of the present disclosure;
[0043] Figure 4 is a flowchart of another method for optimizing a presentation provided according to an embodiment of the present disclosure;
[0044] Figure 5 It is a flow chart of a method for fine-tuning a large model of page image classification provided according to an embodiment of the present disclosure;
[0045] Figure 6 is a flow chart of another method for fine-tuning a large model of page image classification provided according to an embodiment of the present disclosure;
[0046] Figure 7 It is a flow chart of a method for fine-tuning a large model of page image parsing provided according to an embodiment of the present disclosure;
[0047] Figure 8 is a flow chart of another method for fine-tuning a large model of page image parsing provided according to an embodiment of the present disclosure;
[0048] Fig. 9 is a flow chart of another method for optimizing a presentation provided according to an embodiment of the present disclosure;
[0049] Fig.10 is a schematic diagram of a page optimization result provided according to an embodiment of the present disclosure;
[0050] Fig.11 It is a structural schematic diagram of a presentation optimization device provided according to an embodiment of the present disclosure;
[0051] Fig.12 It is a structural schematic diagram of a fine-tuning device for a large model of page image classification provided according to an embodiment of the present disclosure;
[0052] Fig.13 It is a structural schematic diagram of a device for fine-tuning a large model of page image parsing provided according to an embodiment of the present disclosure;
[0053] Fig.14 It is a block diagram of an electronic device used to implement the presentation optimization method, the page image classification model fine-tuning method or the page image analysis model fine-tuning method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0054] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0055] In one example, Figure 1 This is a flowchart of a method for optimizing a presentation provided in accordance with an embodiment of the present disclosure. This embodiment is applicable to the case of optimizing a presentation to be optimized. The method can be executed by a presentation optimization device, which can be implemented by software and / or hardware and can generally be integrated in an electronic device. The electronic device can be a terminal device, a server device, or an intelligent robot device, etc. The embodiment of the present disclosure does not limit the specific device type of the electronic device. Accordingly, Figure 1 As shown, the method includes the following:
[0056] S110: Determine candidate page images of the presentation to be optimized.
[0057] The presentation to be optimized may be a promotional presentation, a display presentation, a summary presentation, or a report presentation, etc., which is preliminarily designed by the user, and is not limited in this embodiment.
[0058] In this embodiment, the candidate page images of the presentation to be optimized may include: a cover image, a directory image, images of each chapter, images of each text, or an image of the last page of the presentation, which are not limited in this embodiment.
[0059] Optionally, in this embodiment, the electronic device can receive a presentation to be optimized uploaded by a user through a web page, an application, or a mini-program deployed in a target application; further, the acquired presentation can be parsed and converted into multiple page images; exemplarily, a display page of a presentation can be determined as a page image; or a page slide of a presentation can be determined as a page image, which is not limited in this embodiment.
[0060] Furthermore, each page image can be screened to obtain candidate page images; illustratively, a page image containing text information, picture information, video information or table information can be determined as a candidate page image; a blank page image can be determined as a non-candidate page image, which is not limited in this embodiment.
[0061] In an optional implementation of the present embodiment, after obtaining the presentation to be optimized, the presentation can be further input into a screenshot tool, and the presentation can be screenshotted based on the screenshot tool to obtain page images of the presentation. Furthermore, image recognition can be performed on each page image, and candidate page images can be determined based on the image recognition results. Exemplarily, page images with non-empty image recognition results can be determined as candidate page images; and page images with empty image recognition results can be determined as non-candidate page images.
[0062] In this embodiment, if the page image is determined to be a non-candidate page image, the page image may be deleted, thereby avoiding processing of the page image containing no content information and reducing the algorithm execution efficiency.
[0063] S120: Determine a page analysis result of the candidate page image, and when the candidate page image is determined to be a target page image according to the page analysis result, optimize the target page image.
[0064] The target page image may be a page image that needs to be optimized among the candidate page images, for example, a text page of a presentation to be optimized.
[0065] Optionally, in the present embodiment, after determining to obtain a candidate page image of the presentation, the page analysis result of the candidate page image can be further determined. Furthermore, it can be determined based on the page analysis result of the candidate page whether the candidate page image is a target page image that needs to be optimized. If it is determined that the candidate page image is a target page image that needs to be optimized, the target page image can be further optimized based on the page analysis result.
[0066] In an optional implementation of the present embodiment, after determining each candidate page image of the presentation to be optimized, the page analysis result of each candidate page image can be further determined. Furthermore, it can be determined based on the page analysis result whether the candidate page image is the target page image to be optimized. If it is the target page image, the optimization result can be further determined based on the page analysis result.
[0067] Optionally, in this embodiment, the page analysis result of the candidate page image may include a classification result of the candidate page image, and may also include a parsing result of the candidate page image, which is not limited in this embodiment.
[0068] In an optional implementation of the present embodiment, after determining each candidate page image of the presentation to be optimized, the classification result and the analysis result of each candidate page image can be further determined; further, it can be determined whether the candidate page image is the target page image based on the classification result of the candidate page image; exemplarily, the classification result of the candidate page image can be matched with the target classification result (for example, the main text page), and if the two are consistent, it can be determined that the candidate page image is the target page image, and further, if it is determined that the candidate page image is the target page image based on the classification result, then the page optimization result of the candidate page image can be further determined based on the analysis result of the candidate page image; exemplarily, the text in the candidate target page image can be summarized, the unclear image can be replaced, or the table can be converted into a bar chart, etc., according to the analysis result, which is not limited in the present embodiment.
[0069] S130. In response to the instruction for generating page optimization results of all target page images of the presentation, determine a target recommendation template, and generate a target presentation based on the target recommendation template and each of the page optimization results.
[0070] The target recommendation template may include a theme, color scheme, font, layout, or animation effect, etc., which is not limited in this embodiment.
[0071] Optionally, in the present embodiment, after all target page images to be optimized are determined from among the candidate page images and optimized to obtain the optimization results of each page, a target recommendation template can be further determined. For example, the target recommendation template can be determined based on the page content of each target page image; further, each page optimization result can be rendered based on the theme, color, font, layout or animation effect of the target recommendation template to obtain a target presentation.
[0072] The scheme of this embodiment determines the candidate page images of the presentation to be optimized; determines the page analysis results of the candidate page images, and when the candidate page images are determined to be the target page images according to the page analysis results, optimizes the target page images; determines a target recommendation template in response to a generation instruction of page optimization results of all target page images of the presentation, and generates a target presentation based on the target recommendation template and each of the page optimization results. The target page images that need to be optimized can be quickly determined, and the target page images can be optimized, thereby improving the optimization efficiency and beautification effect of the presentation, and providing help in meeting the diverse needs of different users.
[0073] Figure 21 is a flowchart of another method for optimizing a presentation provided according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the optimization methods of presentation include the following:
[0074] S210: Determine candidate page images of the presentation to be optimized.
[0075] Optionally, in this embodiment, determining the candidate page images of the presentation to be optimized may include: acquiring the presentation to be optimized, and parsing the presentation; and determining the candidate page images according to the parsing result.
[0076] In an optional implementation of the present embodiment, after the presentation to be optimized is obtained, the presentation can be further parsed. For example, each slide, content data, image data or chart data of the presentation can be determined. Furthermore, the candidate page image can be determined based on the parsed results of the presentation. Exemplarily, in the present embodiment, a slide can be determined as a candidate page image; content data and image data of the same category can also be determined as a candidate page image; content data or chart data of the same category can also be determined as a candidate page image.
[0077] In another optional implementation of this embodiment, after the presentation to be optimized is obtained, the presentation can be input into a predetermined document analysis tool, and the presentation to be optimized can be segmented based on the document analysis tool to obtain candidate page images.
[0078] The advantage of such a setting is that the candidate page images of the presentation to be optimized can be quickly determined, which helps in the subsequent optimization of the presentation.
[0079] Optionally, in this embodiment, determining the candidate page images of the presentation to be optimized may also include: obtaining the presentation to be optimized, and converting the presentation into a target data format; determining each page image based on the page structure and element content of the target data format, and screening each of the page images to obtain the candidate page images.
[0080] Among them, the target data format can be a structured data format, for example, it can be JSON (JavaScript Object Notation) data format, XML (eXtensible Markup Language) data format, HDF5 (Hierarchical Data Format version 5) data format or other self-developed structured data formats, etc., which are not limited in this embodiment.
[0081] In an optional implementation of the present embodiment, after obtaining the presentation to be optimized, the obtained presentation can be further converted into a target data format, for example, into a JSON data format or into an XML data format, etc.; further, the page structure and element content of the presentation in the target data format can be obtained; further, each page image can be determined based on the page structure and element content, and each page image can be screened to obtain each candidate page image.
[0082] Among them, the page structure can be the category of pages of the presentation, for example, main text page, table of contents page, cover page or ending page, etc., which is not limited in this embodiment; the element content can include elements of different levels such as pages, titles, paragraphs, lists, tables, charts, pictures, etc., which is not limited in this embodiment.
[0083] In an optional implementation of the present embodiment, after each page image is determined based on the page structure and element content of the presentation converted into the target data format, it can be further determined whether each page image contains content, for example, whether it contains text content, picture content or chart content, etc.; exemplarily, if the first page image contains text content, the first page image can be determined as a candidate page image; if the second page image contains text content and picture content, the second page image can also be determined as a candidate page image; if the third page image contains text content, picture content and chart content, the third page image can also be determined as a candidate page image.
[0084] The solution of this embodiment, after obtaining the presentation to be optimized, can further convert the presentation into a target data format, and determine each candidate page image based on the presentation in the target data format, so that the candidate page image can be quickly determined and help be provided for the optimization of subsequent candidate page images.
[0085] S220: Determine a page analysis result of the candidate page image.
[0086] Among them, the page analysis result of the candidate page image may include the page classification result of the candidate page image; in this embodiment, the page classification result is: cover page, directory page, text page, chapter page or last page, which is not limited in this embodiment.
[0087] Correspondingly, in this embodiment, determining the page analysis result of the candidate page image may include: determining the page content of the candidate page image, and determining the page classification result of the candidate page image based on the page content of the candidate page image; or, inputting the candidate page image into a pre-fine-tuned page image classification model to obtain the page classification result of the candidate page image.
[0088] In an optional implementation of this embodiment, after acquiring each candidate page image of the presentation to be optimized, the page content of each candidate page image may be further determined, and the page classification result of each candidate page image may be determined according to the page content.
[0089] Optionally, in this embodiment, the candidate page image can be identified, the text content in the candidate page image can be extracted, and the extracted text content can be semantically understood. For example, the extracted text content can be input into a natural language processing model, and the semantic information matching the text content can be determined by the natural language processing model. Furthermore, the page classification result of the candidate page image can be determined based on the semantic information.
[0090] In an example of this embodiment, after obtaining each candidate page image, the text content in the first candidate page image can be extracted and input into a natural language processing model. The natural language processing model is used to determine that the semantic information matching the text content is “Chapter 1..., Chapter 2..., Chapter 3...”, then the page classification result of the first candidate page image can be determined as “Table of Contents Page”.
[0091] In another optional implementation of the present embodiment, after obtaining each candidate page image of the presentation to be optimized, each candidate page image can be further input into a pre-fine-tuned large model for page classification, and the page classification result of each candidate page image can be determined based on the model capability of the large model for page classification. For example, the image features and text features of each candidate page image can be extracted based on the model capability of the large model for page classification, and the image features and text features can be fused, and then the page classification result of the candidate page image can be output based on the preset model prompt words.
[0092] In an example of this embodiment, if the presentation to be optimized is parsed, a total of 5 candidate page images are obtained, which are input into the page classification model for processing. The output result of the page classification model can be "cover page, directory page, text page, main text page, and last page."
[0093] The solution of this embodiment can obtain the page classification results of the candidate page images in different ways, thereby improving the flexibility of the system and providing a basis for subsequently determining whether the candidate page images need to be optimized.
[0094] S230: Determine whether the candidate page image is a target page image according to the page classification result of the candidate page image.
[0095] Optionally, in this embodiment, after the page classification result of the candidate page image is determined, it may be further determined whether the candidate page image is the target page image based on the page classification result of the candidate page image.
[0096] In an optional implementation of the present embodiment, each page classification result may be defined in advance. For example, a candidate page image whose page classification result is "text page" may be defined as a target page image. Furthermore, if the page classification result of the candidate page image is consistent with a predefined page classification result, then the candidate page image may be determined to be the target page image.
[0097] The advantage of such a setting is that it can quickly determine whether the candidate page image is a target page image that needs to be subsequently optimized, which provides assistance for the optimization of the presentation.
[0098] Optionally, in this embodiment, determining whether the candidate page image is a target page image based on the page classification result of the candidate page image may include: if it is determined that the page classification result of the candidate page image meets a preset filtering condition, determining the candidate page image as a target page image; otherwise, determining the candidate page image as a non-target page image.
[0099] The screening condition may be a page category condition, such as a text page or a directory page, etc., which is not limited in this embodiment.
[0100] In an optional implementation of this embodiment, after determining the page classification result of the candidate page image, it can be further determined whether the page classification result matches the page category condition defined by the screening condition. If so, the candidate page image can be determined to be the target page image, otherwise it is a non-target page image.
[0101] In an example of the present embodiment, if the presentation to be optimized is parsed, a total of 3 candidate page images are obtained, and the page classification results of the 3 candidate page images are "cover page, text page, and last page", and the screening condition is "text page", then the candidate page image with the page classification result of the text page is the target page image to be optimized subsequently.
[0102] S240: When it is determined that the candidate page image is a target page image according to the page analysis result, optimize the target page image.
[0103] S250. In response to the instruction for generating page optimization results of all target page images of the presentation, determine a target recommendation template, and generate a target presentation based on the target recommendation template and each of the page optimization results.
[0104] The scheme of this embodiment, when it is determined that the page classification result of the candidate page image meets the preset screening conditions, the candidate page image can be determined as the target page image to be subsequently optimized, the candidate page images can be screened, and the optimization of candidate page images that do not need to be optimized can be prevented, thereby preventing the optimization efficiency of the presentation from being reduced.
[0105] Figure 3 1 is a flowchart of another method for optimizing a presentation provided according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 3 As shown, the optimization methods of presentation include the following:
[0106] S310: Determine candidate page images of the presentation to be optimized.
[0107] S320: Determine a page analysis result of the candidate page image.
[0108] In this embodiment, the page analysis result of the candidate page image may include, in addition to the page classification result involved in the above embodiment, also: a page parsing result.
[0109] Correspondingly, in the present embodiment, determining the page analysis result of the candidate page image may also include: understanding the image content of the candidate page image, and determining the page parsing result of the candidate page image based on the image understanding result; or, inputting the candidate page image into a pre-fine-tuned page image parsing large model to obtain the page parsing result of the target page image.
[0110] In an optional implementation of the present embodiment, after determining the page classification result of each candidate page image, the page parsing result of each candidate page image can be further determined. In the present embodiment, the image content of each candidate page image can be understood, and then the page parsing result of the candidate page image can be determined based on the image understanding result.
[0111] Exemplarily, in this embodiment, the text content of the candidate page image can be extracted and the text content can be semantically understood to obtain the page parsing results of the candidate page image; each candidate page image can also be input into an image recognition model, and then the page parsing results of each candidate page image can be output based on the image recognition model.
[0112] In another optional implementation of the present embodiment, after obtaining each candidate page image of the presentation to be optimized, each candidate page image can be further input into a pre-fine-tuned page parsing large model, and the page parsing results of each candidate page image can be determined based on the model capabilities of the page parsing large model. For example, the image features and text features of each candidate page image can be extracted based on the model capabilities of the page parsing large model, and the image features and text features can be fused, and then the page parsing results of the candidate page image can be output based on preset model prompt words.
[0113] The solution of this embodiment can obtain the page parsing results of the candidate page images in different ways, thereby improving the flexibility of the system and providing a basis for subsequent optimization of the screened target page images.
[0114] S330: Determine whether the candidate page image is a target page image according to the page classification result of the candidate page image.
[0115] S340: Input the target page image and the page parsing result of the target page image into the optimization big model, and optimize the target page image based on the optimization big model.
[0116] The optimized large model may be any large language model or a multimodal large model, which is not limited in this embodiment.
[0117] Optionally, in this embodiment, after each target page image is determined based on the page classification results of each candidate page image, the page parsing results of the target page image can be further input into the optimization large model to optimize the target page image based on the model capabilities of the optimization large model.
[0118] Exemplarily, in this embodiment, the text content in the target page image can be summarized to reduce redundant text content; the text content in the target page image can also be classified to obtain text content of different topics; the picture content in the target page image can also be replaced, for example, it can be replaced with a picture that is more consistent with the text content.
[0119] S350, in response to the instruction for generating page optimization results of all target page images of the presentation, determining a target recommendation template, and generating a target presentation based on the target recommendation template and each of the page optimization results.
[0120] The solution of this embodiment, after determining each target page image based on the page classification results of each candidate page image, can further input the page parsing results of the target page image into the optimization model, optimize the target page image based on the optimization model, and obtain the optimization result of the target page image. The optimization result of the target page image can be obtained quickly, thereby improving the execution efficiency of the algorithm.
[0121] Figure 4 1 is a flowchart of another method for optimizing a presentation provided according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 4 As shown, the optimization methods of presentation include the following:
[0122] S410: Determine candidate page images of the presentation to be optimized.
[0123] S420: Determine a page analysis result of the candidate page image, and when the candidate page image is determined to be a target page image according to the page analysis result, optimize the target page image.
[0124] S430. In response to the instruction for generating page optimization results of all target page images of the presentation, determine a target recommendation template, and generate a target presentation based on the target recommendation template and each of the page optimization results.
[0125] Optionally, in this embodiment, after obtaining the page optimization results of all target page images of the presentation to be optimized, one or more target recommendation templates can be further determined, and each page optimization result can be rendered based on the target recommendation template to obtain the target presentation.
[0126] In an optional implementation of this embodiment, after determining to obtain the target recommendation template, the page optimization results of each target page image can be further filled in the target recommendation template according to the order of the template page images, so as to obtain an optimized target presentation.
[0127] Optionally, in this embodiment, determining a target recommended template, and generating a target presentation based on the target recommended template and each of the page optimization results, may include: determining at least one recommended template that matches each of the page optimization results, and feeding back each of the recommended templates to a terminal device that uploads the presentation to be optimized; in response to a selection instruction of a target recommended template, rendering the optimization results of each of the target page images or non-target page images in the candidate page images based on the template information of the target recommended template to obtain the target presentation.
[0128] In an optional implementation of the present embodiment, after obtaining the page optimization results of all target page images of the presentation to be optimized, at least one recommended template matching each page optimization result can be further determined; illustratively, if each page optimization result is about new product promotion, then a plurality of product promotion presentation templates with higher scores (for example, scores greater than 8 points) can be selected from the database to be determined as the recommended templates.
[0129] Furthermore, each recommended template can be fed back to the terminal device that uploaded the presentation to be optimized, such as a computer or a smart phone; further, the user can select a recommended template from each recommended template as a target recommended template in the terminal device, and return relevant information of the selected target recommended template; further, the page optimization results of each target page image and the non-target page images in the candidate page images can be rendered based on the target recommended template to obtain the target presentation.
[0130] The non-target page images in the candidate page images are page images in the candidate page images that do not need to be optimized, such as a cover page, a chapter page, or a last page.
[0131] In an optional implementation of the present embodiment, the candidate page images can be sorted based on the specific internalization of the page optimization results of each target page image and the specific content of the non-target page images, and the page optimization results and non-target page images can be filled into the target recommendation template according to the sorting results to obtain the target presentation.
[0132] The advantage of such a setting is that the optimized target presentation can be quickly obtained based on the page optimization results of the target page image and the non-target page images, and the integrity of the target presentation can be guaranteed.
[0133] S440, feeding back the target presentation to the terminal device that uploaded the presentation to be optimized; in response to the modification instruction of the target presentation, determining the modification content, and adjusting the target presentation based on the modification content.
[0134] Optionally, in this embodiment, after obtaining the optimized target presentation, the target presentation can be further fed back to the terminal device that uploaded the presentation to be optimized so that the user can view the template presentation; if the user is not satisfied with the optimization effect of the target presentation, the user can provide feedback on the content to be modified.
[0135] Furthermore, after receiving the modification instruction fed back by the user, the electronic device may further determine the modification content based on the modification instruction, and adjust the target presentation based on the modification content until the user's needs are met.
[0136] The solution of this embodiment can also receive the user's modification instructions in real time after obtaining the optimized target presentation, and adjust the target presentation based on the modification instructions, so as to meet the user's diverse design needs.
[0137] Figure 5 This is a flow chart of a method for fine-tuning a large model of page image classification according to an embodiment of the present disclosure. This embodiment is applicable to the case of fine-tuning a large model of page image classification for determining the page classification result of a candidate page image. The method can be executed by a fine-tuning device for a large model of page image classification. The device can be implemented by software and / or hardware and integrated in an electronic device. The electronic device involved in the present disclosure can be a computer, a server, a tablet computer, or an intelligent robot, etc. Specifically, refer to Figure 5 , the method specifically includes the following:
[0138] S510: Obtain a first training data set.
[0139] The first training data set may include page images of a plurality of presentations and classification results of each of the page images.
[0140] Optionally, in this embodiment, multiple presentations may be obtained, for example, 500, 1,000 or 10,000, etc., and the page image of each presentation may be determined. For example, a slide of a presentation may be determined as a page image; further, each page image may be annotated separately to determine the page classification result of each page image.
[0141] In this embodiment, the page classification result of each page image may be determined based on annotators, or may be determined based on an automated annotation tool, which is not limited in this embodiment.
[0142] S520: Determine a first model prompt word, and input the first model prompt word and the first training data set into a first large model for iterative training.
[0143] Exemplarily, the first model prompt word can be: "If I want to classify this page image, the category options are ['table of contents', 'main text page', 'cover page', 'last page', 'chapter page'], please generate the corresponding category based on the image information, only one of the options, do not write other content"; it should be noted that, in this embodiment, the first model prompt word can also be a prompt word for other guiding models to uniquely output the page image category, and its content is not specifically limited in this embodiment.
[0144] In this embodiment, the first large model may be any multimodal large model, which is not specifically limited in this embodiment.
[0145] Optionally, in this embodiment, after obtaining the first training data set and the first model prompt words, the first training data set and the first large model can be further input into the first large model for iterative training. During this process, various parameters in the first large model can be continuously adjusted based on the first training data set and the first model prompt words, thereby achieving fine-tuning of the first large model.
[0146] S530: When the iteration stop condition is met, a large model for page image classification is obtained.
[0147] The page image classification model is used to determine the page analysis result of the candidate page image, especially to determine the page classification result of the candidate page image.
[0148] In this embodiment, the iteration stop condition may be a number condition or an accuracy condition, which is not limited in this embodiment.
[0149] Optionally, in this embodiment, after the first training data set and the first large model are input into the first large model for iterative training, if the number of iterations reaches a preset number of iterations, for example, 50,000 times or 1 million times, then the final page image classification large model can be output.
[0150] The solution of this embodiment is to obtain a first training data set; determine a first model prompt word, and input the first model prompt word and the first training data set into a first large model for iterative training; when the iteration stop condition is met, a page image classification large model is obtained; and various parameters in the first large model can be adjusted based on the first training data set and the first model prompt word to obtain a page image classification large model suitable for determining the page classification results of the candidate page images, thereby providing assistance for the optimization of the presentation.
[0151] Figure 6 1 is a flowchart of another fine-tuning method of a large model for page image classification provided according to an embodiment of the present disclosure; this embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 6 As shown, the fine-tuning method of the page image classification model includes the following:
[0152] S610: Obtain a first training data set.
[0153] S620: Determine the first model prompt word.
[0154] S630: Freeze the text encoder, input the first page image and the classification result of the first page image into the image encoder to extract image features, and adjust various parameters of the image encoder and the fusion module based on the feature extraction result.
[0155] In this embodiment, the first large model may include: an image encoder, a text encoder and a fusion module; it should be noted that, in this embodiment, only the main structure of the first large model is described, which is not the entire structure of the first large model, and it is not a limitation of this embodiment.
[0156] In this embodiment, after obtaining the first training data set and the first model prompt words, the first large model can be further fine-tuned based on the first training data set and the first model prompt words. The fine-tuning process can be divided into two stages. The first stage is to adjust the parameters of the image encoder and the fusion module. At this time, it is necessary to freeze the parameters of the text encoder, that is, do not adjust the parameters of the text encoder; after the first stage is completed, the parameters of the text encoder are unfrozen, and then the text encoder and the image encoder that have completed the parameter update of the first stage and the parameters of the fusion module are adjusted.
[0157] It can be understood that this step is the first stage of the first large model fine-tuning process, and freezing the text encoder means keeping the parameters of the text encoder unchanged and not updating them.
[0158] The first page image is any page image in the first training data set, which is not limited in this embodiment.
[0159] Optionally, in this embodiment, after the first page image and the classification result of the first page image are input into the first large model, feature extraction can be performed based on the image encoder of the first large model. Furthermore, various parameters of the image encoder and the fusion module can be adjusted based on the feature extraction results. For example, the weights, biases, learning rates or regularization terms can be adjusted, which are not limited in this embodiment.
[0160] In an optional implementation of the present embodiment, after acquiring each image feature through the image encoder, the gap between the predicted value and the true value (the classification result of each page image in the first training data set) can be measured by a loss function; further, the gradient can be calculated by a back-propagation algorithm, and the parameters in the image encoder and the fusion module responsible for combining image and text features can be updated based on the gradient.
[0161] S640, in response to the parameter adjustment completion instruction of the image encoder and the fusion module, unfreeze the text encoder, and input the first training data set and the first model prompt word into the image encoder, the text encoder and the fusion module respectively for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
[0162] Optionally, in this embodiment, after the parameter adjustment of the image encoder and the fusion module is completed, that is, after the parameter update of the image encoder and the fusion module is completed based on all page images of the first training data set and the classification results of each page image, the parameters of the text encoder can be unsealed, and the first training data set and the first model prompt word can be respectively input into the image encoder, the text encoder and the fusion module for iterative training, image features can be extracted based on the image encoder of the first large model, text features can be extracted based on the text encoder, and the extracted text features and image features can be input into the fusion module for feature fusion, and further, the parameters of the image encoder, the text encoder and the fusion module can be updated based on the feature extraction results and the feature fusion results, for example, the weights, biases, learning rates or regularization terms can be adjusted, which are not limited in this embodiment.
[0163] S650: When the iteration stop condition is met, a large model for page image classification is obtained.
[0164] Optionally, in this embodiment, when the iteration stop condition is met, a page image classification large model is obtained, which may include: when a preset number of iterations is reached, outputting the page image classification large model; or, when the output result meets a preset accuracy condition, outputting the page image classification large model.
[0165] The preset number of iterations may be 50,000, 100,000, 200,000 or 500,000, etc., which is not limited in this embodiment. The preset accuracy condition may be that the output accuracy result is 0.98 or 0.99, etc., which is not limited in this embodiment.
[0166] In an optional implementation of this embodiment, after the number of iterations reaches a preset number of iterations, the updating of the parameters of each module can be stopped to obtain the final page image classification model.
[0167] In another optional implementation of this embodiment, during the iterative training process, if the output accuracy reaches 0.999, which is greater than the preset accuracy condition 0.99, the updating of the parameters of each module can be stopped to obtain the final page image classification model.
[0168] The advantage of this setting is that it can prevent the output page image classification model from overfitting, which will lead to the phenomenon that the model has higher training accuracy but lower prediction accuracy.
[0169] In this embodiment, in order to verify the page image classification large model obtained by fine-tuning in this embodiment, quantitative analysis is performed based on the multimodal large model before fine-tuning and the page image classification large model, where Table 1 shows the processing results of the multimodal large model on different page images, and Table 2 shows the processing results of the page image classification large model on different page images.
[0170] Table 1
[0171]
[0172] Table 2
[0173]
[0174] The solution of this embodiment, through a two-stage fine-tuning strategy, makes the large model not only more accurate in visual understanding, but also significantly improves the classification accuracy of presentation page images, which can meet the requirements of high efficiency and high quality in actual application scenarios.
[0175] Figure 7This is a flow chart of a method for fine-tuning a large model of page image parsing according to an embodiment of the present disclosure. This embodiment is applicable to the case where a large model of page image parsing for determining a page parsing result of a candidate page image is obtained by fine-tuning the above embodiments. The method can be executed by a fine-tuning device for a large model of page image parsing, which can be implemented by software and / or hardware and integrated in an electronic device. The electronic device involved in the present disclosure can be a computer, a server, a tablet computer, or an intelligent robot, etc. Specifically, refer to Figure 7 , the method specifically includes the following:
[0176] S710: Obtain a second training data set.
[0177] The second training data set includes page images of multiple presentations and content analysis results of each of the page images.
[0178] Optionally, in this embodiment, multiple presentations may be obtained, for example, 500, 1000 or 10,000, etc., and the page image of each presentation may be determined. For example, a slide of a presentation may be determined as a page image; further, each page image may be parsed separately to determine the page parsing result of each page image.
[0179] In this embodiment, the page analysis results of each page image can be determined based on annotation personnel, or based on automated annotation tools, or can be obtained by identifying each page image through a pre-trained image recognition model, which is not limited in this embodiment.
[0180] S720: Determine a second model prompt word, and input the second model prompt word and the second training data set into the second largest model for iterative training.
[0181] For example, the second model prompt may be: "If I want to extract the text content of this page of PPT image and make a new page of PPT with the same content, please return the extracted content to me in md format; the structure is title\n-subtitle <colon>subcontent\n-subtitle <colon>subcontent\n, where title, subtitle and subcontent are empty strings if they do not have content". It should be noted that, in this embodiment, the second model prompt word can also be a prompt word for other guiding models to uniquely output the page image parsing result, and its content is not specifically limited in this embodiment.
[0182] In this embodiment, the second large model can be any multimodal large model, which is not specifically limited in this embodiment.
[0183] Optionally, in this embodiment, after obtaining the second training data set and the second model prompt words, the second training data set and the second largest model can be further input into the second largest model for iterative training. During this process, various parameters in the second largest model can be continuously adjusted based on the second training data set and the second model prompt words, thereby achieving fine-tuning of the second largest model.
[0184] S730: When the iteration stop condition is met, a large model of page image analysis is obtained.
[0185] The page image parsing macromodel is used to determine the page analysis results of the candidate page images, especially to determine the page parsing results of the candidate page images.
[0186] In this embodiment, the iteration stop condition may be a number condition or an accuracy condition, which is not limited in this embodiment.
[0187] Optionally, in this embodiment, after the second training data set and the second large model are input into the second large model for iterative training, if the number of iterations reaches a preset number of iterations, for example, 50,000 times or 1 million times, then the final page image parsing large model can be output.
[0188] The solution of this embodiment is to obtain a second training data set; determine a second model prompt word, and input the second model prompt word and the second training data set into the second large model for iterative training; when the iteration stop condition is met, obtain a page image parsing large model; and adjust each parameter in the second large model based on the second training data set and the second model prompt word to obtain a page image parsing large model suitable for determining the page parsing result of the candidate page image, thereby providing assistance for the optimization of the presentation.
[0189] Figure 8 1 is a flowchart of another method for fine-tuning a large model of page image parsing according to an embodiment of the present disclosure; this embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 8 As shown, the fine-tuning method of the page image parsing large model includes the following:
[0190] S810: Obtain a second training data set.
[0191] S820, determining the second model prompt word,
[0192] S830, freezing the text encoder, inputting the second page image and the content analysis result of the second page image into the image encoder to extract image features, and adjusting various parameters of the image encoder and the fusion module based on the feature extraction result.
[0193] In this embodiment, the second largest model may include: an image encoder, a text encoder and a fusion module; it should be noted that, in this embodiment, only the main structure of the second largest model is described, which is not the entire structure of the second largest model, and it is not a limitation of this embodiment.
[0194] In this embodiment, after obtaining the second training data set and the second model prompt words, the second large model can be further fine-tuned based on the second training data set and the second model prompt words. The fine-tuning process can be divided into two stages. The second stage is to adjust the parameters of the image encoder and the fusion module. At this time, it is necessary to freeze the parameters of the text encoder, that is, do not adjust the parameters of the text encoder; after the second stage is completed, the parameters of the text encoder are unfrozen, and then the text encoder and the image encoder that have completed the parameter update of the second stage and the parameters of the fusion module are adjusted.
[0195] It can be understood that this step is the second stage of the second largest model fine-tuning process, and freezing the text encoder means keeping the parameters of the text encoder unchanged and not updating them.
[0196] The second page image is any page image in the second training data set, which is not limited in this embodiment.
[0197] Optionally, in this embodiment, after the second page image and the analysis result of the second page image are input into the second large model, feature extraction can be performed based on the image encoder of the second large model. Furthermore, various parameters of the image encoder and the fusion module can be adjusted based on the feature extraction results. For example, weights, biases, learning rates or regularization terms can be adjusted, which are not limited in this embodiment.
[0198] In an optional implementation of the present embodiment, after acquiring each image feature through the image encoder, the gap between the predicted value and the true value (the parsing result of each page image in the second training data set) can be measured by a loss function; further, the gradient can be calculated by a back-propagation algorithm, and the parameters in the image encoder and the fusion module responsible for combining image and text features can be updated based on this gradient.
[0199] S840, in response to the parameter adjustment completion instruction of the image encoder and the fusion module, unfreeze the text encoder, and input the second training data set and the second model prompt words into the image encoder, the text encoder and the fusion module respectively for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
[0200] Optionally, in this embodiment, after the parameter adjustment of the image encoder and the fusion module is completed, that is, after the parameter update of the image encoder and the fusion module is completed based on all page images of the second training data set and the analysis results of each page image, the parameters of the text encoder can be unsealed, and the second training data set and the second model prompt words can be respectively input into the image encoder, the text encoder and the fusion module for iterative training, image features can be extracted based on the image encoder of the second largest model, text features can be extracted based on the text encoder, and the extracted text features and image features can be input into the fusion module for feature fusion, and further, the parameters of the image encoder, the text encoder and the fusion module can be updated based on the feature extraction results and the feature fusion results, for example, the weights, biases, learning rates or regularization terms can be adjusted, which are not limited in this embodiment.
[0201] S850: When the iteration stop condition is met, a large model of page image analysis is obtained.
[0202] Optionally, in this embodiment, when the iteration stop condition is met, the page image parsing large model is obtained, which may include: when a preset number of iterations is reached, the page image parsing large model is output; or, when the output result meets a preset accuracy condition, the page image parsing large model is output.
[0203] The preset number of iterations may be 50,000, 100,000, 200,000 or 500,000, etc., which is not limited in this embodiment. The preset accuracy condition may be that the output accuracy result is 0.98 or 0.99, etc., which is not limited in this embodiment.
[0204] In an optional implementation of this embodiment, after the number of iterations reaches a preset number of iterations, the updating of the parameters of each module can be stopped to obtain the final page image parsing large model.
[0205] In another optional implementation of this embodiment, during the iterative training process, if the output accuracy reaches 0.999, which is greater than the preset accuracy condition of 0.99, the updating of the parameters of each module can be stopped to obtain the final page image parsing large model.
[0206] The advantage of this setting is that it can prevent the output page image parsing large model from overfitting, which will lead to the phenomenon that the model has higher training accuracy but lower prediction accuracy.
[0207] The solution of this embodiment, through a two-stage fine-tuning strategy, makes the large model not only more accurate in visual understanding, but also significantly improves the parsing accuracy of the page images of the presentation, which can meet the requirements of high efficiency and high quality in actual application scenarios.
[0208] In order to better understand the optimization method of the presentation involved in this disclosure, Fig. 9 It is a flowchart of another method for optimizing a presentation provided according to an embodiment of the present disclosure, which mainly includes the following:
[0209] S910: Obtain a presentation to be optimized.
[0210] S920: Convert the presentation to be optimized into a target data format.
[0211] The target data may represent page results and element contents of each page image of the presentation.
[0212] S930: Analyze each page image of the presentation.
[0213] Among them, page analysis mainly includes page classification and page parsing; page classification mainly involves inputting each page image into a pre-fine-tuned page image classification model to obtain a page classification result; page parsing mainly involves inputting each page image into a pre-fine-tuned page image parsing model to obtain a page parsing result; further, based on the page classification results and page parsing results, the target page image that needs to be summarized and optimized can be determined, and the target page image can be input into the target large model for summary and optimization.
[0214] It should be noted that the fine-tuning process of the page image classification model and the page image analysis model has been introduced in the above embodiments and will not be repeated in this embodiment.
[0215] For example, Fig.10 It is a schematic diagram of the page optimization results provided according to an embodiment of the present disclosure, wherein a1-a3 are page images before optimization, and b1-b3 are page images after a1-a3 are optimized based on the summary of the large model.
[0216] S940: Generate a first presentation in a target format based on the target template.
[0217] S950, Target Presentation.
[0218] Optionally, the first presentation may be converted into a target presentation that is consistent with the format of the presentation to be optimized.
[0219] The solution of the disclosed embodiment reduces the reliance on complex manual annotation and manual intervention, thereby reducing the labor cost and resource consumption in the content parsing process. It directly processes page images and structured content through a large multimodal model, significantly improves the parsing speed and automation level, and reduces multiple conversions and processing steps in the traditional parsing process. It can automatically select the most appropriate solution, improve the adaptability to complex page structures, and ensure efficient and accurate performance in different scenarios.
[0220] Fig.11 is a schematic diagram of the structure of a presentation optimization device provided according to an embodiment of the present disclosure, and the device can execute the presentation optimization method involved in any embodiment of the present disclosure; Fig.11 The presentation optimization device includes: a candidate page image determination module 1101, a page analysis result determination module 1102 and a target presentation generation module 1103.
[0221] A candidate page image determination module 1101 is used to determine candidate page images of a presentation to be optimized;
[0222] A page analysis result determination module 1102 is used to determine a page analysis result of the candidate page image, and when the candidate page image is determined to be a target page image according to the page analysis result, optimize the target page image;
[0223] The target presentation generation module 1103 is used to determine a target recommendation template in response to a generation instruction of page optimization results of all target page images of the presentation, and to generate a target presentation based on the target recommendation template and each of the page optimization results.
[0224] The scheme of this embodiment is to determine the candidate page images of the presentation to be optimized through a candidate page image determination module; determine the page analysis results of the candidate page images through a page analysis result determination module, and when the candidate page image is determined to be the target page image according to the page analysis results, optimize the target page image; determine the target recommendation template through a target presentation generation module in response to the generation instruction of the page optimization results of all target page images of the presentation, and generate a target presentation based on the target recommendation template and each of the page optimization results, so as to quickly determine the target page images that need to be optimized and optimize the target page images, thereby improving the optimization efficiency and beautification effect of the presentation and providing help in meeting the diverse needs of different users.
[0225] In an optional implementation of this embodiment, the candidate page image determination module 1101 is specifically used to obtain a presentation to be optimized and parse the presentation;
[0226] Determine candidate page images based on the parsing results.
[0227] In an optional implementation of the present embodiment, the candidate page image determination module 1101 is further specifically used to obtain the presentation to be optimized and convert the presentation into a target data format; determine each page image based on the page structure and element content of the target data format, and screen each of the page images to obtain the candidate page images.
[0228] In an optional implementation of this embodiment, the page analysis result includes: a page classification result;
[0229] Correspondingly, the page analysis result determination module 1102 includes: a page classification result determination submodule, which is used to determine the page content of the candidate page image, and determine the page classification result of the candidate page image according to the page content of the candidate page image; or, input the candidate page image into a pre-fine-tuned page image classification model to obtain the page classification result of the target page image; the page classification result is: cover page, table of contents page, main text page, chapter page or last page.
[0230] In an optional implementation of this embodiment, the page analysis result includes: a page parsing result;
[0231] Correspondingly, the page analysis result determination module 1102 includes: a page parsing result determination sub-module, which is used to understand the image content of the candidate page image and determine the page parsing result of the candidate page image according to the image understanding result; or, input the candidate page image into a pre-fine-tuned page image parsing large model to obtain the page parsing result of the target page image.
[0232] In an optional implementation of this embodiment, the presentation optimization device further includes: a target page image determination module, which is used to determine whether the candidate page image is a target page image based on the page classification result of the candidate page image.
[0233] In an optional implementation of the present embodiment, the target page image determination module is specifically used to determine the candidate page image as the target page image if it is determined that the page classification result of the candidate page image meets a preset screening condition; otherwise, determine the candidate page image as a non-target page image.
[0234] In an optional implementation of this embodiment, the page analysis result determination module 1102 also includes an optimization submodule, which is used to input the target page image and the page parsing result of the target page image into the optimization model, and optimize the target page image based on the optimization model.
[0235] In an optional implementation of the present embodiment, the target presentation generation module 1103 is specifically used to determine at least one recommended template that matches the optimization results of each page, and feed back each of the recommended templates to the terminal device that uploads the presentation to be optimized; in response to a selection instruction of the target recommended template, each of the target page optimization results or the directly output candidate page image is rendered based on the template information of the target recommended template to obtain the target presentation.
[0236] In an optional implementation of the present embodiment, the presentation optimization device further includes: a target presentation adjustment module, which is used to feed back the target presentation to the terminal device that uploaded the presentation to be optimized; in response to the modification instruction of the target presentation, determine the modification content, and adjust the target presentation based on the modification content.
[0237] The above-mentioned presentation optimization device can execute the presentation optimization method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not described in detail in this embodiment, please refer to the presentation optimization method provided by any embodiment of the present disclosure.
[0238] Fig.12 is a schematic diagram of the structure of a fine-tuning device for a large model of page image classification provided according to an embodiment of the present disclosure, and the device can execute the optimization method of the presentation involved in any embodiment of the present disclosure; Fig.12 The fine-tuning device of the page image classification large model includes: a first training data set acquisition module 1201, an iterative training module 1202 and a large model generation module 1203.
[0239] The first training data set acquisition module 1201 is used to acquire a first training data set; the first training data set includes page images of a plurality of presentations and classification results of each of the page images;
[0240] Iterative training module 1202, used to determine a first model prompt word, and input the first model prompt word and the first training data set into a first large model for iterative training;
[0241] A large model generation module 1203 is used to obtain a large model for page image classification when an iteration stop condition is met;
[0242] The page image classification model is used to determine the page analysis results of the candidate page images.
[0243] The solution of this embodiment is to obtain a first training data set through a first training data set acquisition module; determine a first model prompt word through an iterative training module, and input the first model prompt word and the first training data set into a first large model for iterative training; and obtain a page image classification large model through a large model generation module when an iteration stop condition is met. A page image classification large model suitable for determining a page classification result of a candidate page image can be obtained, thereby providing assistance for optimizing a presentation.
[0244] In an optional implementation of this embodiment, the first large model includes: an image encoder, a text encoder, and a fusion module;
[0245] Correspondingly, the iterative training module 1202 is specifically used to freeze the text encoder, input the first page image and the classification result of the first page image into the image encoder for image feature extraction, and adjust the parameters of the image encoder and the fusion module based on the feature extraction result; in response to the parameter adjustment completion instruction of the image encoder and the fusion module, unfreeze the text encoder, and input the first training data set and the first model prompt word into the image encoder, the text encoder and the fusion module respectively for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
[0246] In an optional implementation of this embodiment, the large model generation module 1203 is specifically used to output the page image classification large model when a preset number of iterations is reached; or, when the output result meets a preset accuracy condition, output the page image classification large model.
[0247] The fine-tuning device for the page image classification large model can execute the fine-tuning method for the page image classification large model provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in this embodiment, please refer to the fine-tuning method for the page image classification large model provided in any embodiment of the present disclosure.
[0248] Fig.13 is a schematic diagram of a fine-tuning device for a large model of page image analysis provided according to an embodiment of the present disclosure, and the device can execute the optimization method of the presentation involved in any embodiment of the present disclosure; Fig.13 The fine-tuning device of the page image parsing large model includes: a second training data set acquisition module 1301, an iterative training module 1302 and a large model generation module 1303.
[0249] The second training data set acquisition module 1301 is used to acquire a second training data set; the second training data set includes page images of a plurality of presentations and content analysis results of each of the page images;
[0250] Iterative training module 1302, used to determine a second model prompt word, and input the second model prompt word and the second training data set into a second large model for iterative training;
[0251] A large model generation module 1303 is used to obtain a page image parsing large model when an iteration stop condition is met;
[0252] The page image parsing large model is used to determine the page analysis result of the candidate page image.
[0253] The solution of this embodiment is to obtain the second training data set through the second training data set acquisition module; determine the second model prompt words through the iterative training module, and input the second model prompt words and the second training data set into the second large model for iterative training; and obtain the page image parsing large model through the large model generation module when the iteration stop condition is met, so as to obtain the page image classification large model suitable for determining the page parsing results of the candidate page images, thereby providing assistance for the optimization of the presentation.
[0254] In an optional implementation of this embodiment, the second large model includes: an image encoder, a text encoder, and a fusion module;
[0255] Correspondingly, the iterative training module 1302 is specifically used to freeze the text encoder, input the second page image and the content analysis result of the second page image into the image encoder to extract image features, and adjust the parameters of the image encoder and the fusion module based on the feature extraction result;
[0256] In response to the parameter adjustment completion instruction of the image encoder and the fusion module, the text encoder is unfrozen, and the second training data set and the second model prompt words are respectively input into the image encoder, the text encoder and the fusion module for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
[0257] In an optional implementation of this embodiment, the large model generation module 1303 is specifically used to output the page image parsing large model when a preset number of iterations is reached; or, when the output result meets a preset accuracy condition, output the page image parsing large model.
[0258] The fine-tuning device for the large model of page image parsing can execute the fine-tuning method for the large model of page image parsing provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not described in detail in this embodiment, please refer to the fine-tuning method for the large model of page image parsing provided in any embodiment of the present disclosure.
[0259] In the technical solution disclosed in the present invention, the acquisition, storage and application of the presentations involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0260] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0261] Fig.14 A schematic block diagram of an example electronic device 1400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0262] like Fig.14 As shown, the device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. In the RAM 1403, various programs and data required for the operation of the device 1400 can also be stored. The computing unit 1401, the ROM 1402, and the RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0263] A number of components in the device 1400 are connected to the I / O interface 1405, including: an input unit 1406, such as a keyboard, a mouse, etc.; an output unit 1407, such as various types of displays, speakers, etc.; a storage unit 1408, such as a disk, an optical disk, etc.; and a communication unit 1409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1409 allows the device 1400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0264] The computing unit 1401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1401 performs the various methods and processes described above, such as a method for optimizing a presentation, a method for fine-tuning a large model for page image classification, or a method for fine-tuning a large model for page image parsing. For example, in some embodiments, a method for optimizing a presentation, a method for fine-tuning a large model for page image classification, or a method for fine-tuning a large model for page image parsing may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1400 via the ROM 1402 and / or the communication unit 1409. When the computer program is loaded into RAM 1403 and executed by the computing unit 1401, one or more steps of the above-described method for optimizing the presentation, the method for fine-tuning the large model for page image classification, or the method for fine-tuning the large model for page image parsing can be performed. Alternatively, in other embodiments, the computing unit 1401 can be configured to perform the method for optimizing the presentation, the method for fine-tuning the large model for page image classification, or the method for fine-tuning the large model for page image parsing by any other appropriate means (e.g., by means of firmware).
[0265] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0266] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0267] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0268] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0269] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0270] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0271] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0272] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.< / colon> < / colon>
Claims
1. A method for optimizing a presentation, comprising: determining candidate page images of the presentation to be optimized; Determining a page analysis result of the candidate page image, and optimizing the target page image when the candidate page image is determined to be a target page image according to the page analysis result; In response to the instruction to generate the page optimization results of all the target page images of the presentation, a target recommendation template is determined, and a target presentation is generated based on the target recommendation template and each of the page optimization results.
2. The method for optimizing presentation according to claim 1, wherein: The step of determining the candidate page images of the presentation to be optimized includes: Obtaining a presentation to be optimized, and parsing the presentation; Determine candidate page images based on the parsing results.
3. The method for optimizing presentation according to claim 1, wherein: The step of determining the candidate page images of the presentation to be optimized further includes: Acquire the presentation to be optimized, and convert the presentation into a target data format; Each page image is determined based on the page structure and element content of the target data format, and each page image is screened to obtain the candidate page image.
4. The method for optimizing presentation according to claim 1, wherein: The page analysis results include: page classification results; Accordingly, the determining of the page analysis result of the candidate page image includes: Determining the page content of the candidate page image, and determining the page classification result of the candidate page image according to the page content of the candidate page image; or, Inputting the candidate page image into a pre-fine-tuned page image classification model to obtain a page classification result of the candidate page image; The page classification results are: cover page, table of contents page, main text page, chapter page or last page.
5. The method for optimizing presentation according to claim 1, wherein: The page analysis results include: page parsing results; Accordingly, the determining of the page analysis result of the candidate page image includes: Understanding the image content of the candidate page image, and determining a page parsing result of the candidate page image according to the image understanding result; or, The candidate page image is input into a pre-fine-tuned page image parsing macro model to obtain a page parsing result of the target page image.
6. The method for optimizing presentation according to claim 4, wherein: After determining the page analysis result of the candidate page image, the method includes: It is determined whether the candidate page image is a target page image according to the page classification result of the candidate page image.
7. The method for optimizing presentation according to claim 6, wherein: The step of determining whether the candidate page image is a target page image according to the page classification result of the candidate page image comprises: In the case where it is determined that the page classification result of the candidate page image meets a preset screening condition, determining the candidate page image as a target page image; Otherwise, the candidate page image is determined as a non-target page image.
8. The method for optimizing presentation according to claim 4 or 5, wherein: The optimizing the target page image includes: The target page image and the page parsing result of the target page image are input into the optimization big model, and the target page image is optimized based on the optimization big model.
9. The method for optimizing presentation according to claim 1, wherein: The step of determining a target recommendation template and generating a target presentation based on the target recommendation template and each of the page optimization results includes: Determine at least one recommended template that matches the optimization result of each page, and feed back each of the recommended templates to the terminal device that uploaded the presentation to be optimized; In response to a selection instruction of a target recommended template, the optimization results of each target page image or non-target page images in the candidate page images are rendered based on the template information of the target recommended template to obtain the target presentation.
10. The method for optimizing presentation according to claim 1, wherein: After generating a target presentation based on the target recommendation template and each of the page optimization results, the method further includes: Feeding back the target presentation to the terminal device that uploaded the presentation to be optimized; In response to the modification instruction of the target presentation, the modification content is determined, and the target presentation is adjusted based on the modification content.
11. A method for fine-tuning a large model for page image classification, comprising: Obtain a first training data set; The first training data set includes page images of a plurality of presentations and classification results of each of the page images; Determine a first model prompt word, and input the first model prompt word and the first training data set into a first large model for iterative training; When the iteration stop condition is met, a large model for page image classification is obtained; The page image classification model is used to determine the page analysis results of the candidate page images.
12. The fine-tuning method for a large page image classification model according to claim 11, wherein: The first model includes: an image encoder, a text encoder, and a fusion module; Accordingly, the step of inputting the first model prompt word and the first training data set into the first large model for iterative training includes: Freeze the text encoder, input the first page image and the classification result of the first page image into the image encoder to extract image features, and adjust various parameters of the image encoder and the fusion module based on the feature extraction result; In response to the parameter adjustment completion instruction of the image encoder and the fusion module, the text encoder is unfrozen, and the first training data set and the first model prompt word are respectively input into the image encoder, the text encoder and the fusion module for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
13. The fine-tuning method for a large model for page image classification according to claim 11, wherein: When the iteration stop condition is met, a large model for page image classification is obtained, including: When the preset number of iterations is reached, the page image classification model is output; or, When the output result meets the preset accuracy condition, the page image classification model is output.
14. A method for fine-tuning a large model for page image parsing, comprising: Obtain a second training data set; The second training data set includes page images of a plurality of presentations and content parsing results of each of the page images; Determine a second model prompt word, and input the second model prompt word and the second training data set into the second largest model for iterative training; When the iteration stop condition is met, a large model of page image parsing is obtained; The page image parsing large model is used to determine the page analysis result of the candidate page image.
15. The fine-tuning method of the page image parsing large model according to claim 14, wherein: The second largest model includes: an image encoder, a text encoder, and a fusion module; Accordingly, the step of inputting the second model prompt word and the second training data set into the second large model for iterative training includes: Freeze the text encoder, input the second page image and the content analysis result of the second page image into the image encoder to extract image features, and adjust various parameters of the image encoder and the fusion module based on the feature extraction result; In response to the parameter adjustment completion instruction of the image encoder and the fusion module, the text encoder is unfrozen, and the second training data set and the second model prompt words are respectively input into the image encoder, the text encoder and the fusion module for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
16. The fine-tuning method of the page image parsing large model according to claim 14, wherein: When the iteration stop condition is met, a large model of page image parsing is obtained, including: When the preset number of iterations is reached, the page image parsing large model is output; or, When the output result meets the preset accuracy condition, the page image parsing large model is output.
17. A device for optimizing presentation, comprising: A candidate page image determination module, used to determine the candidate page images of the presentation to be optimized; a page analysis result determination module, configured to determine a page analysis result of the candidate page image, and optimize the target page image when the candidate page image is determined to be a target page image according to the page analysis result; The target presentation generation module is used to determine a target recommendation template in response to a generation instruction of page optimization results of all target page images of the presentation, and to generate a target presentation based on the target recommendation template and each of the page optimization results.
18. The presentation optimization device according to claim 17, wherein: The candidate page image determination module is specifically used to: Obtaining a presentation to be optimized, and parsing the presentation; Determine candidate page images based on the parsing results.
19. The presentation optimization device according to claim 17, wherein: The candidate page image determination module is further specifically used for: Acquire the presentation to be optimized, and convert the presentation into a target data format; Each page image is determined based on the page structure and element content of the target data format, and each page image is screened to obtain the candidate page image.
20. The presentation optimization device according to claim 17, wherein: The page analysis results include: page classification results; Accordingly, the page analysis result determination module includes: a page classification result determination submodule, which is used to: Determining the page content of the candidate page image, and determining the page classification result of the candidate page image according to the page content of the candidate page image; or, Inputting the candidate page image into a pre-fine-tuned page image classification model to obtain a page classification result of the target page image; The page classification results are: cover page, table of contents page, main text page, chapter page or last page.
21. The presentation optimization device according to claim 17, wherein: The page analysis results include: page parsing results; Correspondingly, the page analysis result determination module includes: a page parsing result determination submodule, which is used to: Understanding the image content of the candidate page image, and determining a page parsing result of the candidate page image according to the image understanding result; or, The candidate page image is input into a pre-fine-tuned page image parsing macro model to obtain a page parsing result of the target page image.
22. The presentation optimization device according to claim 20, wherein: The presentation optimization device further includes: a target page image determination module for It is determined whether the candidate page image is a target page image according to the page classification result of the candidate page image.
23. The presentation optimization device according to claim 22, wherein: The target page image determination module is specifically used for: In the case where it is determined that the page classification result of the candidate page image meets a preset screening condition, determining the candidate page image as a target page image; Otherwise, the candidate page image is determined as a non-target page image.
24. The presentation optimization device according to claim 20 or 21, wherein: The page analysis result determination module also includes an optimization submodule for: The target page image and the page parsing result of the target page image are input into the optimization big model, and the target page image is optimized based on the optimization big model.
25. The presentation optimization device according to claim 17, wherein: The target presentation document generation module is specifically used for: Determine at least one recommended template that matches the optimization result of each page, and feed back each of the recommended templates to the terminal device that uploaded the presentation to be optimized; In response to a selection instruction of a target recommended template, each target page optimization result or directly output candidate page image is rendered based on the template information of the target recommended template to obtain the target presentation.
26. The presentation optimization device according to claim 17, wherein: The presentation optimization device further includes: a target presentation adjustment module, which is used to: Feeding back the target presentation to the terminal device that uploaded the presentation to be optimized; In response to the modification instruction of the target presentation, the modification content is determined, and the target presentation is adjusted based on the modification content.
27. A fine-tuning device for a large model of page image classification, comprising: A first training data set acquisition module, used to acquire a first training data set; The first training data set includes page images of a plurality of presentations and classification results of each of the page images; An iterative training module, used for determining a first model prompt word, and inputting the first model prompt word and the first training data set into a first large model for iterative training; A large model generation module is used to obtain a large model for page image classification when an iteration stop condition is met; The page image classification model is used to determine the page analysis results of the candidate page images.
28. The fine-tuning device for the large model of page image classification according to claim 27, wherein: The first model includes: an image encoder, a text encoder, and a fusion module; Correspondingly, the iterative training module is specifically used for: Freeze the text encoder, input the first page image and the classification result of the first page image into the image encoder to extract image features, and adjust various parameters of the image encoder and the fusion module based on the feature extraction result; In response to the parameter adjustment completion instruction of the image encoder and the fusion module, the text encoder is unfrozen, and the first training data set and the first model prompt word are respectively input into the image encoder, the text encoder and the fusion module for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
29. The fine-tuning device for the large model of page image classification according to claim 27, wherein: The large model generation module is specifically used for: When the preset number of iterations is reached, the page image classification model is output; or, When the output result meets the preset accuracy condition, the page image classification model is output.
30. A fine-tuning device for a large model of page image analysis, comprising: A second training data set acquisition module, used to acquire a second training data set; The second training data set includes page images of a plurality of presentations and content parsing results of each of the page images; An iterative training module, used for determining a second model prompt word, and inputting the second model prompt word and the second training data set into a second large model for iterative training; A large model generation module is used to obtain a large model of page image analysis when an iteration stop condition is met; The page image parsing large model is used to determine the page analysis result of the candidate page image.
31. The fine-tuning device for the page image parsing large model according to claim 30, wherein: The second largest model includes: an image encoder, a text encoder, and a fusion module; Correspondingly, the iterative training module is specifically used for: Freeze the text encoder, input the second page image and the content analysis result of the second page image into the image encoder to extract image features, and adjust various parameters of the image encoder and the fusion module based on the feature extraction result; In response to the parameter adjustment completion instruction of the image encoder and the fusion module, the text encoder is unfrozen, and the second training data set and the second model prompt words are respectively input into the image encoder, the text encoder and the fusion module for iterative training to adjust the parameters of the image encoder, the text encoder and the fusion module.
32. The fine-tuning device for the page image parsing model according to claim 30, wherein: The large model generation module is specifically used for When the preset number of iterations is reached, the page image parsing large model is output; or, When the output result meets the preset accuracy condition, the page image parsing large model is output.
33. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10, 11-13, or 14-16.
34. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10, 11-13 or 14-16.
35. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-10, 11-13 or 14-16.