Generation method and electronic equipment
By combining the target layout template to process sub-information of document data, the problem of content mismatch in slide document synthesis is solved, and the generation process is efficient and accurate.
Patent Information
- Application Number
- CN202510572061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-19
AI Technical Summary
The existing slide document synthesis technology has problems with the controllability of content generation, and it is easy to cause the synthetic content to not match the actual needs.
By in response to the generation request of document data, sub-information types are obtained and prompt information is generated in combination with the target layout template, and the generation model is used for processing to generate target display documents to ensure that the sub-information matches the target layout template.
Improves the flexibility and efficiency of the generation process, ensures that the generated presentation documents match actual needs, and reduces generation bias and layout misalignment.
Smart Images

Figure CN120508541A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a generation method and electronic device. Background Art
[0002] Automated synthesis of slide documents has garnered widespread attention in recent years. Its core goal is to rapidly generate clearly structured and visually appealing presentations through intelligent algorithms, meeting the demands for efficient content production in areas such as corporate reporting, academic exchanges, and education and training. However, current synthesis methods still face challenges in the controllability of content generation, making it prone to mismatches between synthesized content and actual needs. Summary of the Invention
[0003] The technical solutions provided in this application are as follows:
[0004] The first aspect of the present application provides a generation method, comprising:
[0005] In response to a request for generating document data, obtaining at least one type of sub-information in the document data, the type comprising: at least one of a text type, an image type, an audio type, a video type, and a table type;
[0006] Obtaining a target layout template; the target layout template is used to define the display layout of the sub-information;
[0007] The prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain a target presentation document; various types of sub-information in the target presentation document are matched with the target layout template.
[0008] The prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain a target presentation document, including:
[0009] Inputting the prompt information generated based on the sub-information and the target layout template into the generation model, and processing the prompt information by the generation model to generate the target presentation document;
[0010] or,
[0011] Inputting the prompt information generated based on the sub-information and the target layout template into a generation model, and generating various types of sub-information by the generation model;
[0012] Various types of sub-information generated by the generation model are embedded in the target layout template to form a target presentation document.
[0013] The generating method further comprises:
[0014] Obtaining display feature information corresponding to the document data; the display feature information represents a display style of the document data;
[0015] The prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain a target presentation document, including:
[0016] Inputting the prompt information generated based on the display feature information, the sub-information and the target layout template into a generation model, and generating various types of sub-information by the generation model; the various types of sub-information generated by the generation model are matched with the target layout template and the display style;
[0017] Various types of sub-information generated by the generation model are embedded in the target layout template to form a target presentation document.
[0018] The prompt information generated based on the display feature information, the sub-information and the target layout template is input into a generation model, and the generation model generates various types of sub-information, including at least one of the following:
[0019] inputting the first prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and generating sub-information of the same type but different content by the generation model;
[0020] inputting the second prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and generating sub-information different from the one type by the generation model;
[0021] The third prompt information generated based on the display feature information, the at least two types of sub-information and the target layout template is input into a generation model, and the generation model generates sub-information of one type belonging to the at least two types.
[0022] The one type includes a text type,
[0023] The first prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template is input into a generation model, and the generation model generates sub-information of the same type but different content as the one type, including:
[0024] Inputting first prompt information generated based on the display feature information, the sub-information of the text type, and the target layout template into a first model, and generating the sub-information of the text type that matches the display style and the target layout template by the first model;
[0025] The second prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template is input into a generation model, and the generation model generates sub-information different from the one type, including:
[0026] The second prompt information generated based on the display feature information, the sub-information of the text type and the target layout template is input into the second model, and the second model generates sub-information of the image type, audio type, video type or table type that matches the display style and the target layout template.
[0027] The one type includes an image type,
[0028] The first prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template is input into a generation model, and the generation model generates sub-information of the same type but different content as the one type, including:
[0029] Inputting the first prompt information generated based on the display feature information, the sub-information of the image type and the target layout template into a third model, and allowing the third model to generate the sub-information of the image type that matches the display style and the target layout template;
[0030] The second prompt information generated based on the display feature information, the sub-information of one type, and the target layout template is input into a generation model, and the generation model generates sub-information different from the one type, including:
[0031] The second prompt information generated based on the display feature information, the sub-information of the image type and the target layout template is input into the fourth model, and the fourth model generates the sub-information of the video type or table type that matches the display style and the target layout template.
[0032] The at least two types include a text type and an image type,
[0033] The step of inputting the third prompt information generated based on the display feature information, the at least two types of sub-information, and the target layout template into a generation model, and generating sub-information belonging to one of the at least two types by the generation model, includes:
[0034] The third prompt information generated based on the display feature information, the sub-information of the text type, the sub-information of the image type and the target layout template is input into the fifth model, and the fifth model generates the sub-information of the image type that matches the display style and the target layout template.
[0035] The step of obtaining the target layout template includes:
[0036] In response to the selection request, determining one from a pre-built layout template set as the target layout template;
[0037] The layout template set is constructed in the following way:
[0038] Obtaining a plurality of candidate presentation documents; wherein the plurality of candidate presentation documents have different presentation layouts and / or content areas;
[0039] Obtaining a set of vector representations corresponding to each candidate display document; the set of vector representations includes a plurality of multidimensional vectors, each of the multidimensional vectors being used to represent a layout frame in the candidate display document;
[0040] Clustering multiple groups of vector representations corresponding to the multiple candidate display documents to obtain at least one cluster;
[0041] Based on each cluster in the at least one cluster, the layout template corresponding to each cluster is determined to construct the layout template set.
[0042] The dimensions of the multidimensional vector include:
[0043] The first type of dimension is used to represent the position of the layout box in the candidate presentation document;
[0044] The second type of dimension is used to indicate the size of the layout frame;
[0045] The third type dimension is used to indicate the type of layout box.
[0046] In another aspect of the present application, an electronic device is provided, comprising a memory, at least one processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following method steps:
[0047] In response to a request for generating document data, obtaining at least one type of sub-information in the document data, the type comprising: at least one of a text type, an image type, an audio type, a video type, and a table type;
[0048] Obtaining a target layout template; the target layout template is used to define the display layout of the sub-information;
[0049] The prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain a target presentation document; various types of sub-information in the target presentation document are matched with the target layout template. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0051] Figure 1 A schematic flow chart of a generation method provided in the first embodiment of the present application;
[0052] Figure 2 A schematic flow chart of a generation method provided in the second embodiment of the present application;
[0053] Figure 3 A schematic flow chart of a generation method provided in the third embodiment of the present application;
[0054] Figure 4 A schematic flow chart of a generation method provided in the fourth embodiment of the present application;
[0055] Figure 5 A schematic flow chart of a generation method provided in the twelfth embodiment of the present application;
[0056] Figure 6 A schematic flow chart of a generation method provided for the fourteenth embodiment of the present application;
[0057] Figure 7 This is a schematic diagram of the structure of a generating device provided in this application. DETAILED DESCRIPTION
[0058] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0059] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0060] The terms "first", "second" etc. in this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of identical properties when describing them in the embodiments of the present application. In addition, the terms "comprise" and "have" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0061] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0062] Reference Figure 1 , which is a flow chart of a generation method provided in the first embodiment of the present application, such as Figure 1 As shown, the method may include but is not limited to the following steps:
[0063] Step S101: In response to a request for generating document data, obtain at least one type of sub-information in the document data.
[0064] In this embodiment, the user may input a request for generating document data through a Web application, a desktop client, or a mobile APP.
[0065] The document data may include, but is not limited to, at least one of a text document (eg, Word, PDF), an image collection, a video file, an audio file, and a table file.
[0066] The at least one type of document data may include, but is not limited to, at least one of a text type, an image type, an audio type, a video type, and a table type.
[0067] Obtaining sub-information of the text type in the document data may include but is not limited to:
[0068] Based on natural language processing technology (such as large language model (LLM)), the key information of the text in the document data is extracted as sub-information of the text type.
[0069] Obtaining sub-information of the image type in the document data may include but is not limited to:
[0070] The image recognition algorithm is used to identify objects (such as objects, logos, faces, etc.) in the image of the document data, the position of the object in the image, and the size of the object as sub-information of the image type.
[0071] Obtaining sub-information of the audio type in the document data may include but is not limited to:
[0072] Extract audio files and audio metadata (such as duration, sampling rate, encoding format, etc.) from document data; convert the audio files into text; and use at least one of the audio files and text and the audio metadata as sub-information of the audio type.
[0073] Obtaining sub-information of the video type in the document data may include but is not limited to:
[0074] Extract video files and video metadata (e.g., duration, resolution, frame rate, etc.) from document data; extract key frames from video files;
[0075] At least one of a video file and a key frame and video metadata are used as sub-information of the video type.
[0076] Obtaining sub-information of the table type in the document data may include but is not limited to:
[0077] Extract the row and column structure, cell content, and header information of the table in the document data as sub-information of the table type.
[0078] Step S102: Obtain a target layout template; the target layout template is used to define the display layout of the sub-information.
[0079] In this embodiment, there is no restriction on the method for obtaining the target layout template. However, the target layout template can be obtained to meet the actual display requirements of the document data. For example, the target layout template can adapt to the display requirements of the document data, such as the display position, display size, and arrangement of each information in the layout; and / or, the target layout template can adapt to the content area that the document data needs to display (for example, corporate reports, academic conferences, education and training, etc.).
[0080] In this embodiment, the target layout template may include a set of text instructions, for example, the title position is centered at the top, the body paragraphs are left aligned, and the image occupies 50% of the width of the right side of the page.
[0081] The target layout template can also be represented by a set of vectors. A set of vectors can include: multiple target multi-dimensional vectors, each target multi-dimensional vector is used to identify a layout box in the target layout template.
[0082] The layout frame may include but is not limited to: a text type layout frame, an image type layout frame and a placeholder type layout frame. A placeholder type layout frame can be understood as an area reserved in the target layout template with no specified content type.
[0083] Step S103: Input the prompt information generated based at least on the sub-information and the target layout template into the generation model for processing to obtain a target presentation document.
[0084] In this embodiment, the prompt information is generated based at least on the sub-information and the target layout template, and may include but is not limited to:
[0085] According to the sub-information and the target layout template, natural language or structured instructions are generated as prompt information.
[0086] In this embodiment, the target presentation document can be understood as a visual medium used to present the content of document data (e.g., at least one type of sub-information in the document data) in a structured manner. For example, the types of target presentation documents may include, but are not limited to, slides, interactive reports, and e-books.
[0087] Various types of sub-information in the target presentation document may be matched with the target layout template.
[0088] In this embodiment, by responding to a generation request for document data, obtaining at least one type of sub-information in the document data, obtaining a target layout template, and inputting prompt information generated at least based on the sub-information and the target layout template into the generation model for processing, a target display document is obtained. During the generation process, the generation model can be guided by the sub-information and constrained by the target layout template, thereby reducing the generation deviation of the generation model during the generation process (such as the sub-information deviating from the content of the document data, layout misalignment, etc.), and ensuring that the target display document finally obtained matches the actual needs.
[0089] As another optional embodiment of the present application, refer to Figure 2 , is a flow chart of a generation method provided in the second embodiment of the present application. This embodiment is mainly an implementation of the above step S103. Figure 2 As shown, step S103 may include but is not limited to:
[0090] Step S1031: input the prompt information generated based on the sub-information and the target layout template into a generation model, and the generation model processes the prompt information to generate a target presentation document.
[0091] In this embodiment, the process of generating prompt information based on the sub-information and the target layout template is similar to the principles of the above steps S11-S13, but the generated prompt information is specifically used for the scenario of directly generating the target presentation document based on the generation model.
[0092] For example, text-type sub-information in document data may include "annual revenue increased by 20%", "sales of product A is 10 million", and image-type sub-information in document data may include "product A schematic diagram".
[0093] The prompt information may include: "Based on text-type sub-information such as 'Company Annual Report', 'Annual Revenue Increased by 20%...Product A Sales of 10 Million...' and image-type sub-information 'Product A Schematic Diagram', generate a presentation document that conforms to this target layout template. This target layout template is..."
[0094] In this embodiment, the prompt information generated based on the sub-information and the target layout template is input into the generation model, and the generation model can generate a target display document that is compatible with the target layout template at one time, thereby improving the generation efficiency of the target display document.
[0095] As another optional embodiment of the present application, refer to Figure 3 , is a flow chart of a generation method provided in the third embodiment of the present application. This embodiment is mainly an implementation of the above step S103. Figure 3 As shown, step S103 may include but is not limited to:
[0096] Step S1032: input the prompt information generated based on the sub-information and the target layout template into a generation model, and use the generation model to generate various types of sub-information.
[0097] The type of the sub-information generated by the generation model may partially or completely overlap with at least one type already existing in the document data. It should be noted that even if the types are the same, the specific content of the generated sub-information and the sub-information in the document data may be different.
[0098] The type of the sub-information generated by the generation model may also be completely non-overlapping with all types already existing in the document data. For example, the document data only contains text-type sub-information and table-type sub-information, and the generation model may generate image-type sub-information. For example, the text-type sub-information in the document data may include "annual revenue growth of 20%", "product A sales of 10 million", and the image-type sub-information in the document data may include "product A schematic diagram". The prompt information may include: based on "annual revenue growth of 20%", "product A sales of 10 million", generate an image that meets this target layout template, and this target layout template is..."
[0099] In this embodiment, the sub-information generated by the generation model is adapted to the target layout template. This can be understood as follows: the size (e.g., length, width, height) of the generated sub-information matches the size of the corresponding type of layout frame in the target layout template. For example, the number of characters or paragraph lines in the generated text-type sub-information can be consistent with the preset character capacity or line count threshold of the text-type layout frame in the target layout template. Alternatively, the width and height of the generated image-type sub-information can match the size of the image-type layout frame in the target layout template.
[0100] Step S1033: embed the various types of sub-information generated by the generation model into the target layout template to form a target presentation document.
[0101] In this embodiment, the position of the generated sub-information in the target layout template can be determined based on the correspondence between the size of the generated sub-information and the size of each layout frame in the target layout template, and embedded according to the position.
[0102] If the sub-information type of the generated model does not overlap with any existing types in the document data (e.g., only an image is generated), then only the generated sub-information (e.g., the image) is embedded into the target layout template according to the aforementioned size correspondence, without requiring the entire target layout template to be covered. Unused areas of the target layout template (e.g., title and body text) can remain blank or be processed by other processes, and this is not a limitation in this embodiment.
[0103] In this embodiment, the prompt information generated based on the sub-information and the target layout template is input into the generation model, and various types of sub-information are generated by the generation model. During the generation process, the generation model can be guided by the sub-information and constrained by the target layout template, ensuring that the various types of sub-information generated by the generation model do not deviate from the content of the document data, and that the various types of sub-information generated are compatible with the target layout template.
[0104] On this basis, the various types of sub-information generated by the generation model are embedded into the target layout template to form the target presentation document. This ensures that the resulting target presentation document matches the actual presentation layout requirements. Furthermore, this step-by-step generation approach, which first generates the sub-information and then embeds it into the target layout template, allows for independent optimization and adjustment of each type of sub-information without unnecessarily impacting other parts. This way, even if local errors occur during the generation process, only the corresponding sub-information needs to be corrected, without having to regenerate the entire presentation document. This improves the flexibility and efficiency of the entire generation process.
[0105] As another optional embodiment of the present application, refer to Figure 4 , is a flow chart of a generation method provided in the fourth embodiment of the present application, such as Figure 4 As shown, the method may include but is not limited to the following steps:
[0106] Step S201: In response to a request for generating document data, obtain at least one type of sub-information in the document data, where the sub-information type includes at least one of a text type, an image type, an audio type, a video type, and a table type.
[0107] Step S202: Obtain a target layout template; the target layout template is used to define the display layout of the sub-information.
[0108] The detailed process of steps S201-S202 can be found in the relevant introduction of the above steps S101-S102, which will not be repeated here.
[0109] Step S203: Obtain display feature information corresponding to the document data; the display feature information represents the display style of the document data.
[0110] In this embodiment, the display feature information may include but is not limited to: text feature information (such as font type, font size, line spacing, etc.), image feature information (such as theme color, auxiliary color, background color), chart feature information (such as bar charts, line charts, scatter plots, etc.), table feature information (such as border style and thickness) and interactive behavior feature information (such as automatic playback and manual playback of video and / or audio).
[0111] For example, if the display style of the document data may include business style, the display characteristic information may include: the main color is dark blue, the auxiliary color is light color, the chart is a bar chart, and the font type is bold or Arial.
[0112] If the display style of the document data may include academic style, the display characteristic information may include: the main body color is dark green, the background color is off-white, the chart is a scatter plot, and the font type is Times New Roman or Songti.
[0113] Step S204: input the prompt information generated based on the display feature information, the sub-information and the target layout template into the generation model, and the generation model generates various types of sub-information; the various types of sub-information generated by the generation model match the target layout template and the display style.
[0114] In this embodiment, the prompt information generated based on the display feature information, the sub-information and the target layout template may include but is not limited to:
[0115] According to the display feature information, the sub-information and the target layout template, natural language or structured instructions are generated as prompt information.
[0116] For example, in an implementation where the presentation style is academic, generating natural language based on the presentation characteristic information, the sub-information, and the target layout template may include: based on the sub-information "Experimental Group A mean = 75, standard deviation = 5; Experimental Group B mean = 82, standard deviation = 4," generating an experimental data comparison chart that conforms to the presentation characteristic information and the target layout template. Presentation characteristic information: scatter plot; dark green theme color, off-white background; must include axis labels, legend, and error range. Target layout template: b_1 = (x1, y1, h1, w1, c1), b_2 = (x2 y2, h2, w2, c2), b_3 = (x3 y3, h3, w3, c3).
[0117] In the above prompt information, b_1 may represent a text type layout frame, b_2 may represent a chart type (ie, image and table) layout frame, and b_3 may represent a chart type (ie, image and table) layout frame.
[0118] In this embodiment, the understanding of the type of the sub-information generated by the generation model can refer to the relevant introduction in the above step S1032, which will not be repeated here.
[0119] Step S205: embed the various types of sub-information generated by the generation model into the target layout template to form a target presentation document.
[0120] Steps S204-S205 are an implementation of the above-mentioned step S103.
[0121] In this embodiment, the prompt information generated based on the display feature information, the sub-information and the target layout template is input into the generation model, and various types of sub-information are generated by the generation model. During the generation process, the generation model can be guided by the sub-information and constrained by the target layout template and the display style, ensuring that the various types of sub-information generated by the generation model do not deviate from the content of the document data, and that the various types of sub-information generated match the target layout template and the display style.
[0122] On this basis, the various types of sub-information generated by the generation model are embedded into the target layout template to form the target presentation document. This ensures that the resulting target presentation document matches the actual presentation layout and presentation style requirements. Furthermore, this step-by-step generation method, which first generates the sub-information and then embeds it into the target layout template, allows for independent optimization and adjustment of each type of sub-information without unnecessarily impacting other parts. This way, even if local errors occur during the generation process, only the corresponding sub-information needs to be corrected, without having to regenerate the entire presentation document. This improves the flexibility and efficiency of the entire generation process.
[0123] As another optional embodiment of the present application, a generation method is provided in the fifth embodiment of the present application. This embodiment is mainly an implementation of the above step S203. Step S203 may include but is not limited to any of the following:
[0124] Step S2031: Obtain the customized display feature information input by the user.
[0125] In this embodiment, the style keyword input by the user (eg, business style dark blue theme) may be used as the customized display feature information.
[0126] In this embodiment, by obtaining the customized display feature information input by the user, the user can customize the display style more flexibly.
[0127] Step S2032: extracting display feature information from the candidate display document input by the user.
[0128] The candidate presentation document may be, but is not limited to, a public presentation document downloaded from the Internet or a private presentation document uploaded by a user.
[0129] In this embodiment, the display style information can be extracted from the candidate display documents input by the user through a multimodal AI model and organized into style prompt words as display feature information.
[0130] In this embodiment, by extracting presentation feature information from the candidate presentation documents input by the user, the user does not need to manually summarize the presentation feature information, thereby reducing human errors and ensuring that the presentation feature information meets the user's requirements for presentation style.
[0131] As another optional embodiment of the present application, a generation method is provided in the sixth embodiment of the present application. This embodiment is mainly an implementation of the above step S204. Step S204 may include but is not limited to at least one of the following:
[0132] Step S2041: input the first prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and the generation model generates sub-information of the same type as the one type but different content.
[0133] In this embodiment, the sub-information generated by the generative model and having the same type (e.g., any one of a text type, an image type, an audio type, a video type, and a table type) but different content can be understood from at least one of the following implementations:
[0134] Stay consistent with one of the above types, but adjust the presentation of the sub-information to match the presentation style and target layout template;
[0135] Under the premise of maintaining consistency with the one type and the content field of the document data, the semantics of the sub-information is expanded or simplified.
[0136] For example, one type of sub-information is a product main body picture. The sub-information generated by the generation model that is of the same type but different in content may include: adding text annotations to the same product main body picture (i.e., an implementation method of semantic extension), and the product main body picture with added text annotations matches the display style and target layout template.
[0137] In this embodiment, the first prompt information generated based on the display feature information, the sub-information of one type and the target layout template is input into the generation model, and the generation model generates sub-information of the same type but different content as the one type. This can achieve rewriting of the sub-information of one type so that its type remains unchanged and its content matches the display style of the document data and the target layout template.
[0138] Step S2042: input the second prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and use the generation model to generate sub-information different from the one type.
[0139] The generation model generates sub-information that is different from the one type. Although the presentation method is different from that of the one type of sub-information, its semantics still matches the semantics of the one type of sub-information (i.e., logically associated, and there is no contradiction between the semantics).
[0140] In this embodiment, the second prompt information generated based on the display feature information, the sub-information of one type and the target layout template is input into the generation model, and the sub-information different from the one type is generated by the generation model. It can be ensured that the sub-information different from the one type and the sub-information of one type are semantically matched (that is, logically associated, and there is no contradiction between the semantics), and the generated sub-information matches the display style and the target layout template, thereby avoiding the display style and semantic conflicts in the target display document.
[0141] Step S2043: inputting the third prompt information generated based on the display feature information, the at least two types of sub-information and the target layout template into a generation model, and allowing the generation model to generate sub-information of one type belonging to the at least two types.
[0142] In this embodiment, the sub-information of one type belonging to the at least two types generated by the generation model may match the semantics of the sub-information of the at least two types (ie, logically associated, and without contradiction between the semantics).
[0143] In this embodiment, the third prompt information generated based on the display feature information, at least two types of sub-information and the target layout template is input into the generation model, and the generation model generates sub-information belonging to one type of the at least two types, so that the generated sub-information can carry more comprehensive content and enhance the expression effect of the generated sub-information.
[0144] As another optional embodiment of the present application, a generation method is provided for the seventh embodiment of the present application. This embodiment is mainly an implementation of the above step S2041. In this implementation, the one type may include a text type, and step S2041 may include but is not limited to:
[0145] Step S20411: input the first prompt information generated based on the display feature information, the sub-information of the text type and the target layout template into the first model, and the first model generates the sub-information of the text type that matches the display style and the target layout template.
[0146] The first model is an implementation of the generative model. In this application, the specific structure of the first model is not limited. For example, the first model may include but is not limited to: a large language model (LLM).
[0147] In this embodiment, the first prompt information may allow the first model to retain the original content of the text-type sub-information in the document data, but may adjust the presentation style of the text-type sub-information and require it to match the target layout template. Accordingly, the text-type sub-information generated by the first model that matches the presentation style and the target layout template may maintain the same content as the text-type sub-information in the document data.
[0148] For example, the sub-information of text type in the document data may include: "This product adopts a new generation of energy-saving technology, with standby power consumption as low as 0.5W, and power consumption in daily use mode is 2.1W, which saves about 30% power compared with similar products", and the font type of this sub-information is Kaiti. If the display style corresponding to the document data is a business style, the sub-information of text type generated by the first model that matches the display style and the target layout template may include: "This product adopts a new generation of energy-saving technology, with standby power consumption as low as 0.5W, and power consumption in daily use mode is 2.1W, which saves about 30% power compared with similar products", and the font type of this sub-information is bold, Arial, and the sub-information matches the target layout template.
[0149] Of course, the first prompt information may also allow the first model to adjust the original content of the text-type sub-information in the document data, and simultaneously adjust the presentation style of the text-type sub-information, requiring it to match the target layout template. Accordingly, the text-type sub-information generated by the first model that matches the presentation style and the target layout template may not be completely identical in content to the text-type sub-information in the document data, but may maintain semantic consistency.
[0150] In this embodiment, the first prompt information generated based on the display feature information, the sub-information of the text type and the target layout template is input into the first model, and the first model generates the sub-information of the text type that matches the display style and the target layout template. The sub-information of the text type can be rewritten so that its type remains unchanged and the content matches the display style and target layout template of the document data.
[0151] As another optional embodiment of the present application, a generation method is provided for the eighth embodiment of the present application. This embodiment is mainly an implementation of the above step S2042. In this implementation, the one type may include a text type, and step S2042 may include but is not limited to:
[0152] Step S20421: input the second prompt information generated based on the display feature information, the sub-information of the text type and the target layout template into the second model, and the second model generates sub-information of the image type, audio type, video type or table type that matches the display style and the target layout template.
[0153] The second model is an implementation of the generation model. In this application, the specific structure of the second model is not limited. For example, the second model may include but is not limited to: a Wensheng graph model, a Wensheng table model, etc.
[0154] In this embodiment, the sub-information of the text type in the document data has the characteristics of high structuredness and concentrated semantic density. By inputting the second prompt information generated based on the display feature information, the sub-information of the text type, and the target layout template into the second model, the second model can quickly understand and process the sub-information of the text type, thereby improving generation efficiency. In addition, the sub-information of the text type in the document data often has strong descriptive and expressive capabilities. By inputting the second prompt information generated based on the display feature information, the sub-information of the text type, and the target layout template into the second model, the generation process of the second model can be accurately constrained by the sub-information of the text type, so that the sub-information of the image type, audio type, video type, or table type generated by the second model is more closely matched with the display style and the target layout template.
[0155] As another optional embodiment of the present application, a generation method is provided for the ninth embodiment of the present application. This embodiment is mainly an implementation of the above step S2041. In this implementation, the one type may include an image type, and step S2041 may include but is not limited to:
[0156] Step S20412: input the first prompt information generated based on the display feature information, the sub-information of the image type and the target layout template into the third model, and the third model generates the sub-information of the image type that matches the display style and the target layout template.
[0157] The third model is an implementation of the generative model.
[0158] In this embodiment, the first prompt information may allow the third model to retain the original content of the image-type sub-information in the document data, but may adjust the presentation style of the image-type sub-information and require it to match the target layout template. Accordingly, the image-type sub-information generated by the third model that matches the presentation style and the target layout template may maintain the same content as the image-type sub-information in the document data.
[0159] Of course, the first prompt information may also allow the third model to adjust the original content of the image-type sub-information in the document data, while also adjusting the presentation style of the text-type sub-information and requiring it to match the target layout template. Accordingly, the image-type sub-information generated by the third model that matches the presentation style and the target layout template may not be completely identical in content to the image-type sub-information in the document data, but may maintain semantic consistency.
[0160] In this embodiment, the first prompt information generated based on the display feature information, the sub-information of the image type and the target layout template is input into the third model, and the third model generates the sub-information of the image type that matches the display style and the target layout template. The sub-information of the image type can be rewritten so that its type remains unchanged and the content matches the display style and target layout template of the document data.
[0161] As another optional embodiment of the present application, a generation method is provided for the tenth embodiment of the present application. This embodiment is mainly an implementation of the above step S2042. In this implementation, the one type may include an image type, and step S2042 may include but is not limited to:
[0162] Step S20422: input the second prompt information generated based on the display feature information, the sub-information of the image type and the target layout template into the fourth model, and the fourth model generates the sub-information of the video type or table type that matches the display style and the target layout template.
[0163] The fourth model is an implementation of the generative model.
[0164] In this embodiment, the second prompt information allows the fourth model to retain the original content of the image-type sub-information in the document data, but requires adjusting the presentation style of the image-type sub-information and matching it with the target layout template. Accordingly, the video-type or table-type sub-information generated by the fourth model that matches the presentation style and the target layout template can maintain the same content as the image-type sub-information in the document data.
[0165] Of course, the second prompt information may also allow the fourth model to adjust the original content of the image-type sub-information in the document data, while also adjusting the presentation style of the text-type sub-information and requiring it to match the target layout template. Accordingly, the video-type or table-type sub-information generated by the fourth model that matches the presentation style and the target layout template may not be completely identical in content to the image-type sub-information in the document data, but the semantics may remain consistent.
[0166] In this embodiment, the second prompt information generated based on the display feature information, the sub-information of the image type and the target layout template is input into the fourth model, so that the fourth model can directly generate sub-information of the video type or table type based on the pixel-level information, thereby ensuring that the generated sub-information matches the display style and the target layout template while having a higher degree of match with the details of the sub-information of the image type in the document data.
[0167] As another optional embodiment of the present application, a generation method is provided for the eleventh embodiment of the present application. This embodiment is mainly an implementation of the above step S2043. In this implementation, the at least two types may include a text type and an image type. Step S2043 may include but is not limited to:
[0168] Step S20431: input the third prompt information generated based on the display feature information, the sub-information of the text type, the sub-information of the image type and the target layout template into the fifth model, and the fifth model generates the sub-information of the image type that matches the display style and the target layout template.
[0169] The fifth model is an implementation of the generative model.
[0170] In this embodiment, the third prompt information can allow the fifth model to retain the semantics of the text type sub-information and the image type sub-information in the document data, but new image type sub-information needs to be generated and needs to match the display style and target layout template.
[0171] The image type sub-information generated by the fifth model that matches the display style and target layout template can match the semantics of the text type sub-information and image type sub-information in the document data (i.e., logically associated and without contradiction between the semantics).
[0172] In this embodiment, the third prompt information generated based on the display feature information, the sub-information of the text type, the sub-information of the image type and the target layout template is input into the fifth model, so that the fifth model can combine the accuracy of the sub-information of the text type and the intuitiveness of the sub-information of the image type, and ensure that the sub-information of the image type generated by the fifth model not only matches the display style and the target layout template, but also can fully express the content of the document data.
[0173] As another optional embodiment of the present application, refer to Figure 5 , is a flow chart of a generation method provided in the twelfth embodiment of the present application. This embodiment is mainly an implementation of the above step S102. Figure 5 As shown, step S102 may include but is not limited to:
[0174] Step S1021: In response to the selection request, determine one from a pre-built layout template set as the target layout template.
[0175] In this embodiment, each layout template in the pre-built layout template set may belong to the same content domain. Of course, the content domains to which each layout template belongs may also be different from each other.
[0176] In this embodiment, the layout template set can be constructed by, but is not limited to, the following methods:
[0177] Step S10211: Obtain multiple candidate display documents; the multiple candidate display documents have different display layouts and / or content areas.
[0178] In this embodiment, the plurality of candidate presentation documents may be obtained by downloading a public presentation document from the network, uploading a private presentation document by a user, or the like.
[0179] In this embodiment, there must be differences between the presentation layouts of the multiple candidate presentation documents.
[0180] The content domains of the candidate presentation documents may be the same or different. For example, the candidate presentation documents may all be from the domains of corporate reports, academic conferences, or education and training. Alternatively, the candidate presentation documents may all be from at least two of the content domains of corporate reports, academic conferences, or education and training.
[0181] Step S10212: Obtain a set of vector representations corresponding to each candidate display document.
[0182] In this embodiment, candidate presentation documents may include, but are not limited to, candidate slides. In this embodiment, candidate slides may be parsed based on a slide layout parsing function to obtain a set of vector representations corresponding to the candidate slides. For slides in PowerPoint format, the slide layout parsing function may include Microsoft Office components. For slides in non-PowerPoint format, the non-PowerPoint slides may be converted to image format using a format conversion tool. The slide layout parsing function may include a layout detection model. The layout detection model may identify layout boxes in images.
[0183] The set of vector representations includes a plurality of multi-dimensional vectors, each of which is used to represent a layout frame in the candidate presentation document.
[0184] Step S10213: Cluster the multiple groups of vector representations corresponding to the multiple candidate display documents to obtain at least one cluster.
[0185] In this embodiment, each group of vector representations can be normalized separately, and its numerical range can be scaled to the interval [0, 1]. This normalization method can eliminate the dimensional differences between different groups of vector representations and avoid deviations in the subsequent clustering process caused by uneven numerical scales.
[0186] After normalization, the K-means clustering method can be used to cluster multiple vector representations. The clustering process using the K-means clustering method can be formally described by the following relationship:
[0187]
[0188] Wherein, t_i represents the i-th cluster, l can represent a set of vector representations, l is one of N sets of vector representations, and N represents the number of multiple candidate display documents. l = {b_i|i∈[1,M]}, where M represents the number of layout boxes in a candidate display document; bi_i represents a multidimensional vector, representing a layout box. K can represent the number of layout templates (user-settable, K can be greater than 1 and K is not greater than the above N); L_μi can represent the cluster center within the i-th cluster, which can be the average representation of all group vector representations within the i-th cluster; D(l,L_μi) can represent the distance between a set of vector representations and the cluster center within the i-th cluster; arg T min can mean assigning multiple groups of vector representations to K clusters so that the sum of the distances between each group of vector representations and the cluster center within the cluster to which they belong is minimized.
[0189] Step S10214: Based on each cluster in the at least one cluster, determine the layout template corresponding to each cluster to construct the layout template set.
[0190] In this embodiment, the cluster centers within the K clusters obtained by clustering the above relationship expressions can be used as layout templates corresponding to the clusters.
[0191] In this embodiment, by obtaining multiple candidate display documents, each having different display layouts and / or content domains, obtaining a set of vector representations corresponding to each of the candidate display documents, clustering the multiple sets of vector representations corresponding to the multiple candidate display documents to obtain at least one cluster, and determining, based on each cluster in the at least one cluster, a layout template corresponding to each cluster to construct the layout template set, the representativeness and diversity of the layout templates in the layout template set can be ensured. When a user initiates a selection request, a layout template that meets both the content domain requirements and the document data display layout requirements can be quickly determined from the pre-constructed layout template set in response to the selection request as the target layout template.
[0192] As another optional embodiment of the present application, a generation method is provided in the thirteenth embodiment of the present application. This embodiment is mainly an implementation method of the above-mentioned multidimensional vector. The dimensions of the multidimensional vector may include but are not limited to:
[0193] The first type of dimension is used to indicate the position of the layout box in the candidate presentation document.
[0194] The first type of dimension may include a first dimension and a second dimension, where the first dimension may represent the horizontal coordinate of the upper left corner of the layout box (which may be represented as x), and the second dimension may represent the vertical coordinate of the upper left corner of the layout box (which may be represented as y).
[0195] The second type of dimension is used to represent the size of the layout frame. The second type of dimension may include a third dimension and a fourth dimension. The third dimension may represent the height of the layout frame (which may be represented as h), and the fourth dimension may represent the width of the layout frame (which may be represented as w).
[0196] The third dimension type is used to indicate the type of layout frame. The third dimension type may include a fifth dimension, which may indicate the layout frame type code (which may be represented as c). For example, 0, 1, 2, 0 indicates a placeholder type layout frame, 1 indicates a chart type layout frame, and 2 indicates a text type layout frame.
[0197] For example, corresponding to the method of clustering multiple groups of vector representations using the K-means clustering method, multiple clusters can be expressed as T = {t_i|i∈[1,K]}, t_i represents the i-th cluster, K can represent the number of layout templates, and the cluster center within the i-th cluster can contain multiple 5-dimensional vectors (which can be expressed as L_μi = {b_j|j∈[1,M]}), b_j (i.e., 5-dimensional vector) can include: (x, y, h, w, 2), (x, y, h, w, 1) or (0.5, 0.5, 0, 0, 0).
[0198] A 5-dimensional vector with c > 0.5 is considered a valid layout frame, whereas a 5-dimensional vector with c > 0.5 is considered an invalid layout frame. 0.5 is a set value that can be set as needed.
[0199] As another optional embodiment of the present application, a generation method is provided in the fourteenth embodiment of the present application. This embodiment is mainly an implementation of the above step S102. Step S102 may include but is not limited to:
[0200] Step S1022: Obtain the user-defined layout template input as the target layout template.
[0201] The customized layout template may match the content of the document data.
[0202] In this embodiment, users can break through the limitations of pre-built layout templates by customizing the layout templates to meet the customization needs of the display layout.
[0203] Next, combine Figure 6 , the generation method of this application is described in detail. For example, Figure 6As shown, multiple candidate display documents are obtained, and layout parsing is performed on each candidate display document to obtain a set of vector representations corresponding to each candidate display document. The multiple sets of vector representations corresponding to the multiple candidate display documents are clustered to obtain at least one cluster. The cluster center within the cluster is used as the layout template corresponding to the cluster to construct a layout template set. In response to a selection request, a target layout template is determined from the pre-constructed layout template set.
[0204] Extract key text information from document data based on a large language model.
[0205] The presentation style information in the candidate presentation documents input by the user is extracted through a multimodal AI model and organized into style prompt words (i.e., an implementation method of presentation feature information).
[0206] The prompt information generated based on the style prompt words, target layout template and text key information is input into the large language model, and the large language model generates sub-information of the text type that matches the display style represented by the style prompt words and the target layout template.
[0207] The prompt information generated based on the style prompt words, the target layout template and the text key information is input into the text-graph model, and the text-graph model generates sub-information of the image type that matches the display style represented by the style prompt words and the target layout template.
[0208] The text-type sub-information and the image-type sub-information are embedded in the target layout template to form a target display document.
[0209] Next, the generation device provided in this application is introduced. The generation device introduced below and the generation method introduced above can be referenced to each other.
[0210] Reference Figure 7 The generating device includes: a first obtaining module 100, a second obtaining module 200 and a processing module 300.
[0211] The first obtaining module 100 is configured to obtain at least one type of sub-information in the document data in response to a request for generating document data, wherein the sub-information includes at least one of text type, image type, audio type, video type and table type.
[0212] The second obtaining module 200 is used to obtain a target layout template; the target layout template is used to define the display layout of the sub-information.
[0213] The processing module 300 is used to input the prompt information generated based on at least the sub-information and the target layout template into the generation model for processing to obtain a target display document; various types of sub-information in the target display document are matched with the target layout template.
[0214] The processing module 300 may be specifically configured to:
[0215] The prompt information generated based on the sub-information and the target layout template is input into the generation model, and the target presentation document is generated after being processed by the generation model.
[0216] or,
[0217] Inputting the prompt information generated based on the sub-information and the target layout template into a generation model, and generating various types of sub-information by the generation model;
[0218] Various types of sub-information generated by the generation model are embedded in the target layout template to form a target presentation document.
[0219] The generating device may further include:
[0220] The third obtaining module is configured to obtain display feature information corresponding to the document data; the display feature information represents a display style of the document data.
[0221] The processing module 300 may be specifically configured to:
[0222] Inputting the prompt information generated based on the display feature information, the sub-information and the target layout template into a generation model, and generating various types of sub-information by the generation model; the various types of sub-information generated by the generation model are matched with the target layout template and the display style;
[0223] Various types of sub-information generated by the generation model are embedded in the target layout template to form a target presentation document.
[0224] The processing module 300 inputs the prompt information generated based on the display feature information, the sub-information, and the target layout template into a generation model, and the generation model generates various types of sub-information, which may include at least one of the following:
[0225] inputting the first prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and generating sub-information of the same type but different content by the generation model;
[0226] inputting the second prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and generating sub-information different from the one type by the generation model;
[0227] The third prompt information generated based on the display feature information, the at least two types of sub-information and the target layout template is input into a generation model, and the generation model generates sub-information of one type belonging to the at least two types.
[0228] The one type may include a text type. The processing module 300 inputs the first prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template into a generation model. The generation model generates sub-information of the same type but different content, which may include:
[0229] The first prompt information generated based on the display feature information, the sub-information of the text type and the target layout template is input into the first model, and the first model generates the sub-information of the text type that matches the display style and the target layout template.
[0230] The processing module 300 inputs the second prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template into a generation model, and the generation model generates sub-information different from the one type, which may include:
[0231] The second prompt information generated based on the display feature information, the sub-information of the text type and the target layout template is input into the second model, and the second model generates sub-information of the image type, audio type, video type or table type that matches the display style and the target layout template.
[0232] The one type may include an image type. The processing module 300 inputs the first prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template into the generation model. The generation model generates sub-information of the same type but different content, which may include:
[0233] The first prompt information generated based on the display feature information, the sub-information of the image type and the target layout template is input into the third model, and the third model generates the sub-information of the image type that matches the display style and the target layout template.
[0234] The processing module 300 inputs the second prompt information generated based on the display feature information, the sub-information of one type, and the target layout template into a generation model, and the generation model generates sub-information different from the one type, which may include:
[0235] The second prompt information generated based on the display feature information, the sub-information of the image type and the target layout template is input into the fourth model, and the fourth model generates the sub-information of the video type or table type that matches the display style and the target layout template.
[0236] The at least two types may include a text type and an image type. The processing module 300 inputs the third prompt information generated based on the display feature information, the at least two types of sub-information, and the target layout template into the generation model, and the generation model generates sub-information of one type belonging to the at least two types, which may include:
[0237] The third prompt information generated based on the display feature information, the sub-information of the text type, the sub-information of the image type and the target layout template is input into the fifth model, and the fifth model generates the sub-information of the image type that matches the display style and the target layout template.
[0238] The second obtaining module 200 may be specifically used to:
[0239] In response to the selection request, one is determined from a set of pre-built layout templates as the target layout template.
[0240] The layout template set can be constructed in the following way:
[0241] Obtaining a plurality of candidate presentation documents; wherein the plurality of candidate presentation documents have different presentation layouts and / or content areas;
[0242] Obtaining a set of vector representations corresponding to each candidate display document; the set of vector representations includes a plurality of multidimensional vectors, each of the multidimensional vectors being used to represent a layout frame in the candidate display document;
[0243] Clustering multiple groups of vector representations corresponding to the multiple candidate display documents to obtain at least one cluster;
[0244] Based on each cluster in the at least one cluster, the layout template corresponding to each cluster is determined to construct the layout template set.
[0245] The dimensions of the multidimensional vector may include:
[0246] The first type of dimension is used to represent the position of the layout box in the candidate presentation document;
[0247] The second type of dimension is used to indicate the size of the layout frame;
[0248] The third type dimension is used to indicate the type of layout box.
[0249] In another embodiment of the application, an electronic device is provided, including a memory, at least one processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following method steps:
[0250] In response to a request for generating document data, obtaining at least one type of sub-information in the document data, the type comprising: at least one of a text type, an image type, an audio type, a video type, and a table type;
[0251] Obtaining a target layout template; the target layout template is used to define the display layout of the sub-information;
[0252] The prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain a target presentation document; various types of sub-information in the target presentation document are matched with the target layout template.
[0253] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0254] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0255] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0256] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A generation method comprising: In response to a request for generating document data, obtaining at least one type of sub-information in the document data, the type comprising: at least one of a text type, an image type, an audio type, a video type, and a table type; Obtaining a target layout template; the target layout template is used to define the display layout of the sub-information; The prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain a target presentation document; various types of sub-information in the target presentation document are matched with the target layout template.
2. The generation method according to claim 1, wherein the prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain the target presentation document, comprising: Inputting the prompt information generated based on the sub-information and the target layout template into the generation model, and processing the prompt information by the generation model to generate the target presentation document; or, Inputting the prompt information generated based on the sub-information and the target layout template into a generation model, and generating various types of sub-information by the generation model; Various types of sub-information generated by the generation model are embedded in the target layout template to form a target presentation document.
3. The generation method according to claim 1, further comprising: Obtaining display feature information corresponding to the document data; The display feature information represents the display style of the document data; The prompt information generated based on at least the sub-information and the target layout template is input into the generation model for processing to obtain a target presentation document, including: Inputting the prompt information generated based on the display feature information, the sub-information and the target layout template into a generation model, and generating various types of sub-information by the generation model; the various types of sub-information generated by the generation model are matched with the target layout template and the display style; Various types of sub-information generated by the generation model are embedded in the target layout template to form a target presentation document.
4. The generation method according to claim 3, wherein the prompt information generated based on the display feature information, the sub-information, and the target layout template is input into a generation model, and the generation model generates various types of sub-information, including at least one of the following: inputting the first prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and generating sub-information of the same type but different content by the generation model; inputting the second prompt information generated based on the display feature information, the sub-information of one type and the target layout template into a generation model, and generating sub-information different from the one type by the generation model; The third prompt information generated based on the display feature information, the at least two types of sub-information and the target layout template is input into a generation model, and the generation model generates sub-information of one type belonging to the at least two types.
5. The generation method according to claim 4, wherein the one type comprises a text type, The first prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template is input into a generation model, and the generation model generates sub-information of the same type but different content as the one type, including: Inputting first prompt information generated based on the display feature information, the sub-information of the text type, and the target layout template into a first model, and generating the sub-information of the text type that matches the display style and the target layout template by the first model; The second prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template is input into a generation model, and the generation model generates sub-information different from the one type, including: The second prompt information generated based on the display feature information, the sub-information of the text type and the target layout template is input into the second model, and the second model generates sub-information of the image type, audio type, video type or table type that matches the display style and the target layout template.
6. The generating method according to claim 4, wherein the one type comprises an image type, The first prompt information generated based on the display feature information, the sub-information of the one type, and the target layout template is input into a generation model, and the generation model generates sub-information of the same type but different content as the one type, including: Inputting the first prompt information generated based on the display feature information, the sub-information of the image type and the target layout template into a third model, and allowing the third model to generate the sub-information of the image type that matches the display style and the target layout template; The second prompt information generated based on the display feature information, the sub-information of one type, and the target layout template is input into a generation model, and the generation model generates sub-information different from the one type, including: The second prompt information generated based on the display feature information, the sub-information of the image type and the target layout template is input into the fourth model, and the fourth model generates the sub-information of the video type or table type that matches the display style and the target layout template.
7. The generating method according to claim 4, wherein the at least two types include a text type and an image type. The step of inputting the third prompt information generated based on the display feature information, the at least two types of sub-information, and the target layout template into a generation model, and generating sub-information belonging to one of the at least two types by the generation model, includes: The third prompt information generated based on the display feature information, the sub-information of the text type, the sub-information of the image type and the target layout template is input into the fifth model, and the fifth model generates the sub-information of the image type that matches the display style and the target layout template.
8. The generation method according to claim 1, wherein obtaining the target layout template comprises: In response to the selection request, determining one from a pre-built layout template set as the target layout template; The layout template set is constructed in the following way: Obtaining a plurality of candidate presentation documents; wherein the plurality of candidate presentation documents have different presentation layouts and / or content areas; Obtaining a set of vector representations corresponding to each candidate display document; the set of vector representations includes a plurality of multidimensional vectors, each of the multidimensional vectors being used to represent a layout frame in the candidate display document; Clustering multiple groups of vector representations corresponding to the multiple candidate display documents to obtain at least one cluster; Based on each cluster in the at least one cluster, the layout template corresponding to each cluster is determined to construct the layout template set.
9. The method according to claim 8, wherein the dimensions of the multidimensional vector include: The first type of dimension is used to represent the position of the layout box in the candidate presentation document; The second type of dimension is used to indicate the size of the layout frame; The third type dimension is used to indicate the type of layout box.
10. An electronic device comprising a memory, at least one processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following method steps: In response to a request for generating document data, obtaining at least one type of sub-information in the document data, the type comprising: At least one of a text type, an image type, an audio type, a video type, and a table type; Get the target layout template; The target layout template is used to define the display layout of the sub-information; Inputting the prompt information generated based on at least the sub-information and the target layout template into the generation model for processing to obtain a target presentation document; Various types of sub-information in the target presentation document match the target layout template.