Document generation method and device, electronic equipment and computer readable storage medium

Through multimodal large model and natural language processing technology, the visual characteristics and content structure of the presentation are identified and high-quality documents that meet user needs are generated, which solves the problems of cumbersome production and low intelligence in traditional presentations, and realizes efficient and personalized document generation.

CN120493892APending Publication Date: 2025-08-15BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510754358.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional presentations are cumbersome, ordinary users lack professional design capabilities, and it is difficult to produce high-quality attractive presentations. The existing tools support dynamic content is not perfect, and the degree of intelligence is low, so they cannot automatically optimize the layout and style according to user needs.

Method used

The multimodal big model uses a multimodal big model to identify the reference document set, combine it with the first big model to identify the content semantics and content relationships of the document description information, generate document planning information, and adjust the target reference document based on this to generate documents that meet user needs.

Benefits of technology

It improves the efficiency and quality of document generation and meets user personalized needs. The generated documents can achieve high quality at both visual and content levels, saving labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493892A_ABST
    Figure CN120493892A_ABST
Patent Text Reader

Abstract

The invention provides a document generation method and device, and relates to the technical fields of document processing, artificial intelligence, natural language processing, computer vision and the like. According to the specific implementation scheme, a reference document set and document description information are determined based on document demand information of a user; performing visual feature recognition on the reference document set by adopting a multi-modal large model to obtain a visual feature set; performing content semantic and content relationship identification on the document description information by adopting a first large model to obtain data structure information; determining a target reference document from the reference document set based on the visual feature set and the data structure information, and generating document planning information; and based on the document planning information, adjusting the target reference document to obtain a target demand document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the field of data processing technology, and in particular to technical fields such as document processing, artificial intelligence, natural language processing, and computer vision, and particularly to a document generation method and apparatus, an electronic device, and a computer-readable storage medium. Background Art

[0002] Presentations are a widely used information display and communication tool in modern work and learning. Creating high-quality presentations is a common but time-consuming task. To create new presentations with unique presentation styles, the traditional approach involves manually adjusting text, images, layout, and other details, which consumes considerable time and effort. Furthermore, ordinary users lack professional design skills, making it difficult to create engaging presentations. Existing tools lack comprehensive support for dynamic content and are unable to automatically optimize layout and style based on user needs. Summary of the Invention

[0003] The present disclosure provides a document generation method and apparatus, an electronic device, and a computer-readable storage medium.

[0004] According to a first aspect, a document generation method is provided, which includes: determining a reference document set and document description information based on a user's document demand information; using a multimodal large model to perform visual feature recognition on the reference document set to obtain a visual feature set; using a first large model to perform content semantics and content relationship recognition on the document description information to obtain data structure information; based on the visual feature set and the data structure information, determining a target reference document from the reference document set and generating document planning information; based on the document planning information, adjusting the target reference document to obtain a target demand document.

[0005] According to the second aspect, a document generation device is provided, which includes: a determination unit, configured to determine a reference document set and document description information based on a user's document demand information; a visual acquisition unit, configured to use a multimodal large model to perform visual feature recognition on the reference document set to obtain a visual feature set; a structure acquisition unit, configured to use a first large model to perform content semantics and content relationship recognition on the document description information to obtain data structure information; a generation unit, configured to determine a target reference document from the reference document set based on the visual feature set and the data structure information, and generate document planning information; a document acquisition unit, configured to adjust the target reference document based on the document planning information to obtain a target demand document.

[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method as described in any implementation of the first aspect.

[0008] The document generation method and device provided by the embodiments of the present disclosure first determine a reference document set and document description information based on the user's document requirement information; secondly, use a multimodal large model to perform visual feature recognition on the reference document set to obtain a visual feature set; thirdly, use the first large model to perform content semantics and content relationship recognition on the document description information to obtain data structure information; then, based on the visual feature set and data structure information, determine the target reference document from the reference document set and generate document planning information; finally, based on the document planning information, adjust the target reference document to obtain the target requirement document. Therefore, this method can not only meet the personalized needs of users, but also achieve high-quality document generation at the visual and content levels, significantly improving the efficiency and quality of document generation and saving labor costs.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0011] Figure 1 is a flow chart of an embodiment of a document generation method according to the present disclosure;

[0012] Figure 2 It is a system architecture diagram of the document generation method disclosed herein;

[0013] Figure 3 It is a structural diagram of an embodiment of the document generating device disclosed herein;

[0014] Figure 4 It is a block diagram of an electronic device used to implement the document generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] Unless expressly stated otherwise, throughout the specification and claims, the term "comprise" or variations such as "include" or "comprising", etc., will be understood to include the stated elements or components but not to exclude other elements or other components.

[0016] The technical solutions of the present disclosure are described below through specific examples. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.

[0017] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.

[0018] Traditional presentation creation methods suffer from the following problems: Manual operations are cumbersome, requiring users to manually adjust details like text, images, and layout, consuming significant time and effort. Design capabilities are insufficient, as ordinary users lack professional design skills and struggle to create engaging presentations. Dynamic content processing is difficult, with existing tools offering inadequate support for dynamic content (such as animations and interactive effects). Traditional tools also lack intelligence, preventing them from automatically optimizing layout and style based on user needs.

[0019] This year, the development of AI technology and multimodal large models has provided new ideas for solving the above problems. For example, existing technologies generate presentations through natural language processing, but these technologies mainly focus on text generation in presentations, and support for functions such as image recognition and typesetting analysis is still insufficient.

[0020] In response to the defects in traditional technologies, the present disclosure proposes a document generation method that converts the document to be processed into the target document by utilizing the capabilities of the graph structure, thereby improving the efficiency of target document conversion. Figure 1 A process 100 according to an embodiment of a document generation method of the present disclosure is shown. The document generation method includes the following steps:

[0021] Step 101: Determine a reference document set and document description information based on the user's document requirement information.

[0022] In this embodiment, the document requirement information is the requirement information of the target requirement document to be generated as described by the user, wherein the target requirement document is the document that the user expects to generate and is also the document generated by the document generation method disclosed herein. The document requirement information may include: a directory of the reference document set, content to be filled, layout requirements, etc. The reference document set generates an initial document set for the target requirement document. The reference document set includes at least one reference document. The style and layout of any two reference documents in the reference document set may be different. The document description information is the description information of the target requirement document to be generated. The document description information may include: a style description of the target requirement document, the type of the target requirement document, and the content to be recorded in the target requirement document.

[0023] In this embodiment, the system first receives document requirement information input by the user. This information may include key elements such as the document's subject, purpose, style requirements, and word count range. Subsequently, the system uses built-in algorithms and databases to filter out a set of reference documents that are highly relevant to the user's needs from a vast amount of document resources. These reference documents will serve as an important source of material for the target requirement document to be generated. At the same time, the system will also conduct an in-depth analysis of the user's requirement information and extract document description information including the document's core content points and descriptive content. This document description information will be used to guide content organization and style matching in the subsequent document generation process, ensuring that the generated document can accurately meet the user's personalized needs.

[0024] Optionally, the above step 101 includes: based on the user's document requirement information, extracting the style description information, document type and document content information to be filled in the target requirement document to be generated; based on the document type and style description information, searching for all reference documents with the same document type on the Internet, and extracting reference documents that meet the style description information from the searched reference documents to obtain a reference document set; based on the document content information to be filled in, generating document description information.

[0025] In this embodiment, the target requirement document can be a presentation, and the document requirement information may include: a PPTX file related to the reference document set and a Markdown file related to the document description information, wherein the PPTX file is the file format of the presentation, the reference document set can be obtained through the PPTX file, and the document description information can be obtained through the Markdown file.

[0026] Step 102: Use a multimodal large model to perform visual feature recognition on a reference document set to obtain a visual feature set.

[0027] In this embodiment, an advanced multimodal large model is used to perform in-depth visual feature recognition on a reference document set. The multimodal large model is capable of processing multiple types of data, including text, images, charts, etc., thereby comprehensively analyzing the visual elements in the reference documents. By identifying and extracting elements from each reference document in the reference document set, a detailed visual feature set can be generated, which contains key information such as the reference document's layout style, color matching, font selection, chart style, etc. These visual features will provide an important style reference for subsequent document generation, ensuring that the generated document is visually consistent with the user's desired style, while also better meeting the user's personalized needs for document appearance and layout.

[0028] In this embodiment, the visual feature set can be used as follows Figure 2 The Json format shown is stored, and step 102 can be implemented using an intelligent agent, such as using Figure 2 The ppt_analysis_agent agent in the implementation is mainly responsible for generating page-by-page images of the reference document set and calling the multimodal large model analysis layout, generating a detailed visual feature set, and saving the visual feature set in Json data format.

[0029] In this embodiment, ppt_analysis_agent is a key component responsible for analyzing PPT template files, extracting visual features from a set of visual features such as layout, style, and theme features, rendering the PPT as images, processing images in batches to handle large presentations, and using a large visual model to identify the style, layout, and design elements of each slide, extracting a set of visual features and design suggestions. Specifically, in code form, the data structure provided by ppt_analysis_agent is as follows:

[0030]

[0031]

[0032]

[0033] The above data structure consists of the following main components: templateName: template name; slideCount: total number of slides; layouts: list of available layouts, including each layout's name, placeholder information, and inferred purpose; theme: theme information, including colors and fonts; visualFeatures: visual features, including style, color scheme, and layout complexity; slideLayouts: specific layout information for each slide; recommendations: content recommendations based on template features. Step 103: Use the first large model to identify content semantics and content relationships in the document description information to obtain data structure information.

[0034] In this embodiment, the first large model is used to conduct an in-depth analysis of the document description information. This first large model uses advanced natural language processing technology to accurately identify key semantic information in the document description and understand the chapters, vocabulary meanings, sentence structure, and paragraph themes. At the same time, the first large model can also keenly capture the logical relationships between the content, such as cause and effect, parallelism, and progression, thereby constructing clear data structure information. This data structure information provides a solid foundation for the generation of subsequent documents, ensuring that the generated document content is logically rigorous and well-structured, and can accurately convey the document theme and structure expected by the user, making the generation process of the target demand document more efficient and in line with user needs.

[0035] In this embodiment, the above step 103 includes: determining a document editing template according to the required format of the target required document, obtaining a document template file based on the document description information and the document editing template; inputting the document template file into the first large model to obtain data structure information.

[0036] In this embodiment, the data structure information can be used as follows Figure 2 The Json format shown is stored, and step 103 can be implemented by an intelligent agent, such as using Figure 2 The markdown_agent in the markdown_agent is mainly responsible for analyzing and parsing the document template file, obtaining data structure information including title, content, and the relationship between the contents, and saving the data structure information in Json data format.

[0037] At the code level, the data structure information generated by markdown_agent is as follows:

[0038]

[0039]

[0040]

[0041]

[0042] The above data structure information includes the following main parts: title: document main title; subtitle: document subtitle; sections: document chapter array, each chapter contains: title: chapter title; content: chapter content, which can be text or structured content (such as a list); semantic_type: semantic type of content (concept, process, comparison, etc.); relation_type: relationship type with other chapters; visualization_suggestion: visualization suggestions suitable for the content; subsections: subsection array, the structure is the same as sections.

[0043] Step 104 : Based on the visual feature set and the data structure information, determine the target reference document from the reference document set and generate document planning information.

[0044] In this embodiment, the document planning information is information for planning the overall content of the target requirement document, such as the pages of the target requirement document, the number of pages, the chapters contained in each page, the content of the chapters in each page, etc.

[0045] In this embodiment, the visual feature set and data structure information are comprehensively analyzed to screen out the target reference document from the reference document set that best matches the document to be generated in terms of visual style and content structure. Then, based on the style, layout, and content organization of the target reference document, and in combination with the specific needs of the user, detailed and accurate document planning information is formulated to provide clear guidance and framework for the subsequent document generation, ensuring that the generated document meets the user's expectations in terms of visual presentation and content logic.

[0046] In this embodiment, the document planning information can be used as follows Figure 2 The Json format shown is stored, and step 104 can be implemented using an intelligent agent, such as using Figure 2 The intelligent agent Content_plannning_agent in the implementation is mainly responsible for planning each page of the PPT slide by using the visual feature set provided by ppt_analysis_agent and the data structure information provided by markdown_agent to obtain document planning information. At this time, the document planning information includes: the order of the slides, the title of the slides, the content of the slides, and the form of the content plan presentation.

[0047] The chapter planning data structure generated by content_plannning_agent is as follows:

[0048]

[0049]

[0050]

[0051] The above chapter planning data structure contains the following main parts: slides: slide planning array, each slide contains: slide_index: slide index; slide_id: slide unique identifier; slide_type: slide type; content: slide content, including title, text, list, etc.; template: template information used, including layout name, index and purpose; slide_count: total number of slides.

[0052] Step 105: Based on the document planning information, adjust the target reference document to obtain the target requirement document.

[0053] In this embodiment, the target reference document is adjusted according to the document planning information, thereby generating a target document that meets the user's needs. This process mainly involves fine-tuning the content, structure, format, and visual effects of the target reference document. For example, according to the requirements for content depth and breadth in the document planning information, paragraphs in the target reference document are added, deleted, or modified; the chapters and paragraph order of the document are reorganized according to the planned logical structure; the font, font size, line spacing and other layout details are adjusted according to the planned format requirements; and according to the visual feature set, the layout and style of charts and pictures are optimized to ensure that the final generated document is accurate and complete in content, clear and reasonable in structure, and visually beautiful and coordinated, fully meeting the target requirement document originally proposed by the user.

[0054] In this embodiment, the target requirement document may be a presentation document, and the JSON structure of the target requirement document is as follows:

[0055]

[0056]

[0057]

[0058] The JSON structure of the above target requirements document contains the following main parts: name: presentation name; slide_count: total number of slides; theme: theme information, including color and font; slides: slide array, each slide contains: slide_id: slide ID; real_index: the actual index of the slide in the presentation; layout_name: the name of the layout used; elements: slide element array, each element contains: element_id: element ID, which is a globally unique ID generated by the uuid4() method; type: element type (text, picture, shape, etc.); content: element content; position: element position and size; style: element style (font, color, size, etc.).

[0059] The document generation method provided by the embodiment of the present disclosure first determines a reference document set and document description information based on the user's document demand information; secondly, uses a multimodal large model to perform visual feature recognition on the reference document set to obtain a visual feature set; thirdly, uses the first large model to perform content semantics and content relationship recognition on the document description information to obtain data structure information; then, based on the visual feature set and data structure information, determines the target reference document from the reference document set and generates document planning information; finally, based on the document planning information, adjusts the target reference document to obtain the target demand document. Thus, by performing visual feature recognition on the reference document through the multimodal large model, it is possible to comprehensively consider the visual elements and textual content of the document, and the generated document can better meet the user's needs in terms of visual effects and content. Content semantics and content relationship recognition are performed on the document description information to generate data structure information, so that the generated document structure is clear and logically rigorous, avoiding confusion and disorder in the document content. Based on the user's specific needs and the reference document set, it is possible to generate documents that meet the user's personalized needs, thereby improving the pertinence and practicality of document generation. The entire process involves multiple units working together and utilizing the powerful processing capabilities of large models to complete document generation in a shorter time, thereby improving the efficiency of document generation.

[0060] In some optional implementations of the present disclosure, the above-mentioned determination of the reference document set and document description information based on the user's document demand information includes: obtaining the user's document demand information; extracting at least one document from the document demand information to obtain a reference document set; performing semantic recognition on the document demand information to determine the document description information.

[0061] In this optional implementation, the document requirement information includes a user-provided document collection. After obtaining the user's document requirement information, at least one document can be directly extracted from the document requirement information to obtain a reference document set. The document requirement information also includes content requirement information to be populated into the target requirement document to be generated. The aforementioned semantic recognition of the document requirement information to determine the document description information includes: performing semantic recognition on the content requirement information to determine the content requirement semantic information; and inputting the content requirement semantic information into a pre-trained model to obtain the document description information. Thus, the document description information allows for a comprehensive and in-depth description of the content of the target requirement document to be generated.

[0062] The method for determining the reference document set and document description information provided by this optional implementation directly extracts the reference document set from the document requirement information; performs semantic recognition on the document requirement information and determines the document description information, providing a reliable implementation method for obtaining the reference document set and document description information.

[0063] In some optional implementations of the present disclosure, the above-mentioned use of a multimodal large model to perform visual feature recognition on a reference document set to obtain a visual feature set includes: converting reference documents in the reference document set into images to obtain converted images of each reference document in the reference document set; inputting the converted image of each reference document into the multimodal large model to obtain a visual feature set including the visual features of each reference document, and the multimodal large model is used to analyze the layout of text in the image.

[0064] In this optional implementation, the visual feature set includes at least one visual feature. Visual features are features derived from visually identifying a reference document. Visual features are salient attributes or information in an image or video that can be perceived, extracted, and used for recognition, classification, or understanding by a computer or human visual system. They are fundamental to fields such as computer vision, image processing, and machine learning, and are used to describe the core characteristics of visual data.

[0065] In this optional implementation, visual features include layout features, style features, pattern features, and theme features. Layout features are features reflected in the layout of text in the reference document, and style features are features reflected in the overall style of the reference document, such as a classical style. A visual feature set is a collection of multiple visual features.

[0066] The method for obtaining a visual feature set provided by this optional implementation first converts the reference documents in the reference document set into images, and then performs visual feature recognition on the images through a multimodal large model to obtain a visual feature set, which provides a reliable implementation method for obtaining the visual feature set of the reference document set and improves the reliability of obtaining the visual feature set.

[0067] Optionally, the above-mentioned document generation method also includes: using a multimodal large model to identify the design elements of each reference document in the reference document set, and providing design suggestion information based on the relationship between the identified design elements, wherein the design suggestion information is a reference for the subsequent adjustment of the target requirement document. For example, the design suggestion information is: this reference document is suitable for medium content density, and it is recommended that no more than seven key points per page; pictures can be placed in the picture placeholder on the right, and it is recommended to use high-contrast images; please use the provided theme color to ensure sufficient contrast between text and background.

[0068] In some embodiments of the present disclosure, the above-mentioned use of the first large model to perform content semantics and content relationship recognition on document description information to obtain data structure information includes: inputting document description information and content semantic analysis prompt words into the first large model to obtain content semantic information output by the first large model; inputting document description information and chapter relationship analysis prompt words into the first large model to obtain content relationship information output by the first large model; and constructing structured data structure information based on the content semantic information and content relationship information.

[0069] In this optional implementation, the content semantic analysis prompt word is used to prompt the first large model to perform semantic analysis on the document description information. Under the prompt of the content semantic analysis prompt word, the first large model performs semantic analysis on the document description information to determine the content semantic information of the document description information. The content semantic information includes: chapter titles, title hierarchical relationships and content semantic types, such as content semantic types include: concepts, processes, comparisons, etc.

[0070] In this optional implementation, the chapter relationship analysis prompt words are used to prompt the first large model to perform chapter and content relationship analysis on the document description information. Under the prompt of the chapter relationship analysis prompt words, the first large model performs semantic analysis on the document description information to determine the content relationship information in the document description information. The content relationship information is used to characterize the attribution and association relationship between different chapters and paragraphs in the content. For example, the content relationship information includes: chapters in the content, paragraphs in each chapter, the relationship between multiple paragraphs, etc., wherein the association relationship includes: hierarchy, parallelism, cause and effect, etc.

[0071] In this optional implementation, the structured data structure information is information that characterizes the structural relationship between the contents in the document description information. The data structure information is represented in a structured form. For example, the data structure information includes: document title, document subtitle, document chapter array, each chapter contains chapter titles, and under each chapter title, includes: chapter content, content semantic type, and relationship type with other chapters. In one example, the data structure information generated by the first large model is: title "Application of Artificial Intelligence in Enterprises", subtitle "Transition from Theory to Practice", and content "AI technology has been widely used in various industries, helping enterprises improve efficiency, reduce costs and create value."

[0072] The method for obtaining data structure information provided by this optional implementation method inputs document description information and content semantic analysis prompt words into the first large model to obtain content semantic information output by the first large model; inputs document description information and chapter relationship analysis prompt words into the first large model to obtain content relationship information output by the first large model; and constructs structured data structure information based on the content semantic information and content relationship information. Thus, a structured content representation is constructed through the large model, thereby improving the reliability of obtaining data structure information.

[0073] Optionally, the document generation method further includes inputting the document description information and visualization suggestion prompts into the first large model to obtain visualization suggestion information output by the first large model, wherein the visualization suggestion prompts the large model to provide visualization content that can be realized by the document description information. The visualization suggestion information can be used to generate a target requirement document that conforms to the visualization suggestion information when generating the target requirement document. In this embodiment, the generated data structure information can also include visualization suggestions suitable for the content.

[0074] In some embodiments of the present disclosure, the above-mentioned determining the target reference document from the reference document set based on the visual feature set and the data structure information and generating the document planning information includes: matching the data structure information with the visual feature set to determine the matching visual features; selecting a reference document as the target reference document from the reference document set based on the matching visual features; and generating the document planning information using the second largest model based on the data structure information, the matching visual features and the target reference document.

[0075] In this optional implementation, the matching visual feature is a visual feature in the visual feature set whose similarity with the data structure information is greater than a similarity threshold; the above-mentioned matching visual feature, selecting a reference document from the reference document set as the target reference document includes: selecting a reference document with matching visual features from the reference document set to obtain the target reference document.

[0076] In this optional implementation, the document planning information is information for planning the overall content of the target requirement document, such as the pages of the target requirement document, the number of pages, the chapters contained in each page, the content of the chapters in each page, etc. The above-mentioned method of generating document planning information based on the second largest model based on data structure information, matching visual features and target reference documents includes: determining the number of chapters based on data structure information; determining the layout features of each chapter based on matching visual features; inputting the target reference document, the number of chapters and the layout features of each chapter into the second largest model to obtain the document planning information output by the second largest model.

[0077] Optionally, a second large model can be used to calibrate the initially generated document planning information to generate more accurate document planning information. The above-mentioned method of generating document planning information using the second large model based on data structure information, matching visual features, and target reference documents includes: determining the layout information of the target requirement document based on matching visual features, where the layout information refers to the content arrangement information of the target requirement document; extracting the document style of the target reference document as the style of the target requirement document; filling in corresponding content in each part of the layout information based on the data structure information and style to obtain initial planning information that conforms to the style; inputting the initial planning information and the target reference document into the second large model to obtain the document planning information after the second large model adjusts the initial planning information.

[0078] This optional implementation provides a method for generating document planning information. The method matches data structure information with a set of visual features to determine matching visual features. Based on the matching visual features, a reference document is selected from a reference document set as a target reference document. Finally, based on the data structure information, the matching visual features, and the target reference document, the document planning information is generated using a second model. This provides a reliable method for obtaining document planning information.

[0079] Optionally, the above-mentioned determining the target reference document from the reference document set based on the visual feature set and data structure information and generating document planning information includes: matching the visual feature set with the data structure information to determine the matching visual features, and extracting the target reference document from the reference document set based on the matching visual features; extracting the document title, subtitle and chapter content based on the data structure information; generating document layout information through a large model based on the document title, subtitle and chapter content, filling the chapter content into each part corresponding to the document layout information, and obtaining document planning information.

[0080] In some embodiments of the present disclosure, the above-mentioned use of the second largest model to generate document planning information based on data structure information, matching visual features and target reference documents includes: determining the content distribution in the chapter based on the content semantic information in the data structure information; extracting the hierarchical relationship of the content based on the content relationship in the data structure information; determining the page layout information of the page based on matching visual features; determining each page in the document to be generated and the order and hierarchy of each page based on the hierarchical relationship and page layout information; using a pre-set document component library to collect information on the target reference document to obtain the original document information; inputting the content distribution, hierarchical relationship, order, hierarchy and original document information into the second largest model to obtain the document planning information output by the second largest model.

[0081] In this optional implementation, the document planning information includes: the overall layout of the target requirement document to be generated, the content distribution style and other information. A complete document can be generated through the document planning information. In this optional implementation, the second model (such as Figure 2 As shown in the figure, the most appropriate layout position can be arranged for each chapter, and the corresponding content can be filled in for each chapter to obtain document planning information.

[0082] Optionally, the above-mentioned use of the second largest model to generate document planning information based on the data structure information, matching visual features, and target reference document includes: matching a suitable layout distribution based on the chapter content features (text, lists, tables, images, etc.) in the data structure information; selecting a target layout from the layout distribution based on the semantic type and visualization suggestions in the data structure information; sorting the order and hierarchy of the document pages in the document corresponding to the target layout based on the hierarchical relationship of the content in the data structure information; and inputting the page order, hierarchy, matching visual features, and target reference document into the second largest model to obtain document planning information output by the second largest model. The pages in the document planning information include: an opening page, a content page, and an end page.

[0083] In some embodiments of the present disclosure, based on the document planning information, the target reference document is adjusted to obtain the target requirement document, including: inputting the document planning information into the third largest model to obtain the document operation instructions output by the third largest model; based on the document operation instructions and a pre-set document component library, modifying the target reference document to obtain a modified reference document; converting the modified reference document into a document image, using a multimodal large model to perform layout, content, and style detection on the document image to obtain a detection result; in response to detecting that the detection result is unqualified, modifying the document operation instructions, and continuing to modify the target reference document based on the modified document operation instructions and the document component library until the detection result is qualified, thereby obtaining a final reference document; based on the final reference document, generating the target requirement document.

[0084] In this optional implementation, the document operation instruction is an instruction for operating a document component group. The above-mentioned document operation instruction and the pre-set document component library modify the target reference document to obtain a modified reference document, including: adding components in the document component library to the target reference document through the document operation instruction, or changing components in the target reference document through the document operation instruction, or deleting components in the target reference document through the document operation instruction to obtain a modified reference document.

[0085] In this optional implementation, the above-mentioned multimodal large model is used to detect the layout, content, and style of the document image, and the detection results include: using the multimodal large model to detect whether the layout of the document image matches the layout of the target reference document; using the multimodal large model to detect whether the similarity between the content of the document image and the document planning information is within the similarity threshold (90%); using the multimodal large model to detect whether the style of the document graphic matches the sample of the target reference document; in response to detecting that the layout and sample match the target reference document, and the similarity between the content of the document image and the document planning information is greater than the similarity threshold, determining that the detection result is qualified, and determining that the document currently modified by the document operation instruction is the target requirement document.

[0086] In this optional implementation, when it is detected that the layout or any sample items do not match those of the target reference document, or the similarity between the content of the document image and the document planning information is less than the similarity threshold, the detection result is determined to be unqualified, the document operation instructions are modified, and the modified document operation instructions are used to control the document components in the document component library. The target reference document continues to be modified until the detection result obtained is qualified. At this time, the modified document is the final reference document.

[0087] Optionally, the above-mentioned use of a multimodal large model to detect the layout, content, and style of the document image, and the detection results include: extracting semantic information from the document image, comparing the similarity between the semantic information and the document planning information, and in response to the similarity between the semantic information and the document planning information being greater than a similarity threshold, inputting the document image and the image of the target reference document into the multimodal large model, so that the multimodal large model detects whether the document image matches the layout, content, and style described in the document planning information. If they match, the detection result is qualified.

[0088] In this optional implementation, the document operation instruction includes: the element to be operated, the operation type, the updated element content, and the specific content of the operation. The specific content of the operation varies depending on the operation type. Updating element content includes: replacing images, adjusting text font size, adjusting element position and size, etc.

[0089] In this optional implementation, generating the target requirements document based on the final reference document includes checking whether the final reference document contains any information that does not match the document planning information, and if so, deleting the mismatching information to obtain the target requirements document. For example, if the document planning information includes the number of document pages, if the number of pages in the final reference document exceeds the number of pages in the document, deleting a corresponding number of pages from the final reference document to obtain the target requirements document.

[0090] In this embodiment, the target requirement document can be a presentation. The above-mentioned step of adjusting the target reference document based on the document planning information to obtain the target requirement document can be implemented by an intelligent agent, such as using Figure 2 The intelligent agents slide_generator_agent and PPTFinalizerAgent in it are implemented. Among them, slide_generator_agent is the PPT generation intelligent agent. The main responsibilities of slide_generator_agent are: based on the planning content provided by Content_plannning_agent and the Dom content structure of the PPT file, it uses the big model to analyze the Dom node content that needs to be modified on each page of the target reference document, and generates corresponding operation instructions to provide to slide_generator_agent; modify the corresponding slides in the target reference document, generate images of the modified slides, and use the multimodal big model to check the layout, content, and style, and then provide modification instructions until the requirements are met or the maximum modification number threshold is reached.

[0091] PPTFinalizerAgent is a PPT archiving agent. It is mainly responsible for adjusting the order of slides in the PPT based on the content and order planned by Content_plannning_agent, removing redundant slides, and finally saving the PPT file.

[0092] The operation instruction protocol used by slide_generator_agent is as follows:

[0093]

[0094]

[0095] The above operation instructions contain the following main fields: element_id: the element ID to be operated; operation: operation type, which includes: update_element_content: update element content; replace_image: replace image; adjust_text_font_size: adjust text font size; adjust_element_position: adjust element position and size; content: the specific content of the operation, which varies according to the operation type.

[0096] The method for generating a target requirement document provided by this optional implementation method inputs document planning information into the third largest model to obtain document operation instructions output by the third largest model; based on the document operation instructions and a pre-set document component library, the target reference document is modified to obtain a modified reference document; the modified reference document is converted into a document image, and the document image is subjected to typesetting, content, and style detection using a multimodal large model to obtain a detection result; in response to detecting that the detection result is unqualified, the document operation instructions are modified, and the target reference document is continued to be modified based on the modified document operation instructions and the document component library until the detection result is qualified, thereby obtaining a final reference document; based on the final reference document, a target requirement document is generated, document operation instructions are generated through the third largest model, and the target reference document is modified through the document operation instructions and the document component library, so that the target requirement document can be effectively obtained, thereby improving the reliability of obtaining the target requirement document.

[0097] Further references Figure 3 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a document generation device, which is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0098] like Figure 3As shown, the document generation device 300 provided in this embodiment includes: a determination unit 301, a visual acquisition unit 302, a structure acquisition unit 303, a generation unit 304, and a document acquisition unit 305. The determination unit 301 can be configured to determine the reference document set and document description information based on the user's document demand information. The visual acquisition unit 302 can be configured to use a multimodal large model to perform visual feature recognition on the reference document set to obtain a visual feature set. The structure acquisition unit 303 can be configured to use the first large model to perform content semantics and content relationship recognition on the document description information to obtain data structure information. The generation unit 304 can be configured to determine the target reference document from the reference document set based on the visual feature set and data structure information, and generate document planning information. The document acquisition unit 305 can be configured to adjust the target reference document based on the document planning information to obtain a target demand document.

[0099] In this embodiment, the specific processing of the determination unit 301, the visual obtaining unit 302, the structure obtaining unit 303, the generation unit 304, and the document obtaining unit 305 and the technical effects thereof can be referred to in detail. Figure 1 The relevant descriptions of step 101, step 102, step 103, step 104 and step 105 in the corresponding embodiment are not repeated here.

[0100] In some embodiments of the present disclosure, the above-mentioned determination unit 301 is configured to: obtain the user's document demand information; extract at least one document from the document demand information to obtain a reference document set; perform semantic recognition on the document demand information to determine document description information.

[0101] In some embodiments of the present disclosure, the above-mentioned visual acquisition unit 302 is configured to: convert the reference documents in the reference document set into images to obtain the converted images of each reference document in the reference document set; input the converted images of each reference document into a multimodal large model to obtain a visual feature set including the visual features of each reference document, and the multimodal large model is used to analyze the layout of the text in the image.

[0102] In some embodiments of the present disclosure, the above-mentioned structure obtaining unit 303 is configured to: input document description information and content semantic analysis prompt words into the first large model to obtain content semantic information output by the first large model; input document description information and chapter relationship analysis prompt words into the first large model to obtain content relationship information output by the first large model; and construct structured data structure information based on the content semantic information and content relationship information.

[0103] In some embodiments of the present disclosure, the above-mentioned generation unit 304 is configured to: match the data structure information with the visual feature set to determine the matching visual features; based on the matching visual features, select a reference document from the reference document set as the target reference document; based on the data structure information, the matching visual features and the target reference document, use the second largest model to generate document planning information.

[0104] In some embodiments of the present disclosure, the above-mentioned generation unit 304 is further configured to: determine the content distribution in the chapter based on the content semantic information in the data structure information; extract the hierarchical relationship of the content based on the content relationship in the data structure information; determine the page layout information of the page based on matching visual features; determine the individual pages in the document to be generated and the order and hierarchy of each page based on the hierarchical relationship and page layout information; use a pre-set document component library to collect information on the target reference document to obtain the original document information; input the content distribution, hierarchical relationship, order, hierarchy and original document information into the second largest model to obtain the document planning information output by the second largest model.

[0105] In some embodiments of the present disclosure, the above-mentioned document obtaining unit 305 is configured to: input document planning information into the third largest model to obtain document operation instructions output by the third largest model; based on the document operation instructions and a pre-set document component library, modify the target reference document to obtain a modified reference document; convert the modified reference document into a document image, and use a multimodal large model to perform typesetting, content, and style detection on the document image to obtain a detection result; in response to detecting that the detection result is unqualified, modify the document operation instructions, and continue to modify the target reference document based on the modified document operation instructions and the document component library until the detection result is qualified, thereby obtaining a final reference document; based on the final reference document, generate a target requirement document.

[0106] The document generation device provided by the embodiment of the present disclosure first determines the reference document set and document description information based on the user's document demand information; secondly, the visual acquisition unit 302 uses a multimodal large model to perform visual feature recognition on the reference document set to obtain a visual feature set; thirdly, the structure acquisition unit 303 uses the first large model to perform content semantics and content relationship recognition on the document description information to obtain data structure information; then, the generation unit 304 determines the target reference document from the reference document set based on the visual feature set and data structure information, and generates document planning information; finally, the document acquisition unit 305 adjusts the target reference document based on the document planning information to obtain the target demand document. Thus, by performing visual feature recognition on the reference document through the multimodal large model, the visual elements and text content of the document can be comprehensively considered, and the generated document can better meet the user's needs in terms of visual effects and content. Content semantics and content relationship recognition are performed on the document description information to generate data structure information, so that the generated document structure is clear and logically rigorous, avoiding confusion and disorder in the document content. Based on the user's specific needs and reference document sets, documents can be generated that meet the user's personalized needs, improving the pertinence and practicality of document generation. The entire process, through the collaborative work of multiple units and the powerful processing capabilities of large models, can complete document generation in a relatively short time, improving document generation efficiency.

[0107] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0108] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are provided for example only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0109] like Figure 4As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0110] Various components in device 400 are connected to I / O interface 405, including an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0111] The computing unit 401 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the document generation method. For example, in some embodiments, the document generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the document generation method described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the document generation method by any other appropriate means (e.g., by means of firmware).

[0112] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0113] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable document generation device so that when the program code is executed by the processor or controller, the modes / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0114] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0115] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0116] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0117] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0118] The foregoing descriptions of specific exemplary embodiments of the present disclosure are for purposes of illustration and description. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many variations and modifications are possible in light of the foregoing teachings. The exemplary embodiments have been selected and described for the purpose of explaining the specific principles of the present disclosure and their practical application, thereby enabling those skilled in the art to realize and utilize a variety of exemplary embodiments of the present disclosure and various options and modifications. The scope of the present disclosure is intended to be defined by the claims and their equivalents.

Claims

1. A document generation method, comprising: Determine the reference document set and document description information based on the user's document requirement information; Using a multimodal large model to perform visual feature recognition on the reference document set to obtain a visual feature set; Using the first model to identify content semantics and content relationships of the document description information to obtain data structure information; Based on the visual feature set and the data structure information, determining a target reference document from the reference document set, and generating document planning information; Based on the document planning information, the target reference document is adjusted to obtain a target requirement document.

2. The method according to claim 1, wherein Determining the reference document set and document description information based on the user's document requirement information includes: Obtain user's document requirement information; extracting at least one document from the document requirement information to obtain a reference document set; Perform semantic recognition on the document requirement information to determine document description information.

3. The method according to claim 1, wherein The multimodal large model is used to perform visual feature recognition on the reference document set, and the obtained visual feature set includes: Converting the reference documents in the reference document set into images to obtain converted images of the reference documents in the reference document set; The transformed images of the respective reference documents are input into the multimodal large model to obtain a visual feature set including the visual features of the respective reference documents. The multimodal large model is used to analyze the layout of the text in the image.

4. The method according to claim 1, wherein The first model is used to identify the content semantics and content relationships of the document description information, and the data structure information obtained includes: Inputting the document description information and content semantic analysis prompt words into the first large model to obtain content semantic information output by the first large model; Inputting the document description information and chapter relationship analysis prompt words into the first model to obtain content relationship information output by the first model; Based on the content semantic information and the content relationship information, structured data structure information is constructed.

5. The method according to claim 1, wherein The determining of a target reference document from the reference document set based on the visual feature set and the data structure information and generating document planning information includes: Matching the data structure information with the visual feature set to determine matching visual features; selecting a reference document from the reference document set as a target reference document based on the matching visual features; Based on the data structure information, the matching visual features and the target reference document, a second large model is used to generate document planning information.

6. The method according to claim 5, wherein: The generating of document planning information using the second largest model based on the data structure information, the matching visual features, and the target reference document includes: Determining content distribution in chapters based on content semantic information in the data structure information; Extracting a hierarchical relationship of content based on the content relationship in the data structure information; determining page layout information of the page based on the matching visual features; Determining the pages and the order and hierarchy of the pages in the document to be generated based on the hierarchical relationship and the page layout information; Using a pre-set document component library to collect information from the target reference document to obtain original document information; The content distribution, the hierarchical relationship, the sequence, the level and the original document information are input into the second largest model to obtain document planning information output by the second largest model.

7. The method according to any one of claims 1 to 6, wherein: The adjusting the target reference document based on the document planning information to obtain the target requirement document includes: Inputting the document planning information into a third model to obtain a document operation instruction output by the third model; Based on the document operation instruction and a preset document component library, modify the target reference document to obtain a modified reference document; Converting the modified reference document into a document image, and using the multimodal large model to perform layout, content, and style detection on the document image to obtain a detection result; In response to detecting that the test result is unqualified, modifying the document operation instruction, and continuing to modify the target reference document using the modified document operation instruction and document component group until the test result is qualified, thereby obtaining a final reference document; A target requirement document is generated based on the final reference document.

8. A document generation device, comprising: a determining unit configured to determine a reference document set and document description information based on the user's document requirement information; a visual obtaining unit configured to perform visual feature recognition on the reference document set using a multimodal large model to obtain a visual feature set; a structure obtaining unit configured to use the first large model to perform content semantics and content relationship recognition on the document description information to obtain data structure information; a generating unit configured to determine a target reference document from the reference document set based on the visual feature set and the data structure information, and generate document planning information; The document obtaining unit is configured to adjust the target reference document based on the document planning information to obtain a target requirement document.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Document processing method and device, computer equipment and storage medium

    CN116776850A

  • Design document generation method and device, electronic equipment and storage medium

    CN118295698A

  • Generation method and device of presentation file

    CN118821749A

  • Online webpage generation method and device based on document processing

    CN119003906A

  • Document generation method and device, equipment and storage medium

    CN119474363A

Cited By

  • Multi-modal document content cross-platform analysis system

    CN120726658A

  • Battery operation file generation method, device, equipment and program product

    CN121562566A