Document generation method and device, electronic equipment and computer readable storage medium
By constructing a graph structure to parse the attributes and relationships of presentation elements, the inefficiency problem in traditional methods is solved, and efficient cross-format document conversion and generation is achieved.
Patent Information
- Application Number
- CN202510589603.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Traditional presentations are inefficient in conversion, making it difficult to effectively express complex relationships between elements in the document, resulting in error-prone manual operations and difficulty in processing across formats.
By building a graph structure, using artificial intelligence models to parse page element attributes and relationships of the to be processed documents, build an initial graph structure, and operate it based on user needs to generate a target document.
It realizes cross-format content integration and conversion, improves the efficiency of target document generation, and accurately represents the content, attributes and relationships of document elements.
Smart Images

Figure CN120493904A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of computer technology, and in particular to the fields of natural language processing, operations research, decision science, etc. In particular, it provides a document generation method and apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] Presentations are widely used information display and communication tools in modern work and learning. However, to create new presentations with different presentation styles, the traditional approach is to first parse the presentation document, then manually edit the parsed results based on the document parsing results and convert them into new presentations. Document parsing methods include parsing the presentation's underlying XML (Extensible Markup Language) structure or simply mapping the presentation's content to common data structures such as JSON (JavaScript Object Notation). While this document parsing approach can extract basic content, it struggles to effectively express the complex structure and relationships within the PPT and lacks a comprehensive understanding of the presentation's content semantics and visual layout. While this manual editing approach can convert the parsed content into presentations of other styles, it requires manual effort and is prone to errors, resulting in low efficiency in generating new presentations. Summary of the Invention
[0003] The present disclosure provides a document generation method and apparatus, an electronic device, and a computer-readable storage medium.
[0004] According to a first aspect, a document generation method is provided, which includes: obtaining a document to be processed; parsing the page element attributes and page element relationships of the document to be processed to obtain a page element parsing result; constructing an initial graph structure based on the page element parsing result; operating the initial graph structure based on the document generation requirements input by the user to obtain an operation graph structure; and obtaining a target document based on the operation graph structure.
[0005] According to the second aspect, a document generation device is provided, which includes: an acquisition unit, configured to acquire a document to be processed; a parsing unit, configured to parse page element attributes and page element relationships of the document to be processed to obtain a page element parsing result; a construction unit, configured to construct an initial graph structure based on the page element parsing result; an operation unit, configured to operate the initial graph structure based on a document generation requirement input by a user to obtain an operation graph structure; and an acquisition unit, configured to obtain a target document based on the operation graph structure.
[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.
[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method as described in any implementation of the first aspect.
[0008] The document generation method and device provided by the embodiments of the present disclosure first obtain the document to be processed; secondly, parse the page element attributes and page element relationships of the document to be processed to obtain the page element parsing results; thirdly, construct an initial graph structure based on the page element parsing results; thirdly, operate on the initial graph structure based on the document generation requirements input by the user to obtain an operation graph structure; finally, obtain the target document based on the operation graph structure. Thus, by performing page element and relationship parsing on the document to be processed by an artificial intelligence model and constructing a graph structure, the content, attributes and complex spatial, logical and semantic relationships between page elements in the document to be processed can be accurately and richly represented based on the capabilities of the graph structure; by operating on the initial graph structure and converting it into the operation graph structure of the target document, and converting it into the target document based on the operation graph structure, the graph structure can be automatically used as a universal intermediate representation to achieve cross-format content integration and conversion, thereby improving the efficiency of target document generation compared to manual operation of other data structures.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 is a flow chart of an embodiment of a document generation method according to the present disclosure;
[0012] Figure 2 is a structural diagram of the initial graph structure in the present disclosure;
[0013] Figure 3 It is a structural diagram of an embodiment of the document generating device disclosed herein;
[0014] Figure 4It is a block diagram of an electronic device used to implement the document generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0015] Unless expressly stated otherwise, throughout the specification and claims, the term "comprise" or variations such as "include" or "comprising", etc., will be understood to include the stated elements or components but not to exclude other elements or other components.
[0016] The technical solutions of the present disclosure are described below through specific examples. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.
[0017] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.
[0018] Traditional document parsing methods, such as directly parsing the underlying XML structure of presentation documents, simply mapping their content to general data structures like JSON, or relying on specialized API libraries that provide access to basic document structures (e.g., pages, shapes, and text boxes), can extract basic content but struggle to effectively represent the complex structure and relationships within presentation documents. Specifically, JSON, as a lightweight data exchange format, inherently uses a tree-like or key / value pair structure, which is limited in its ability to represent the rich spatial layout relationships (e.g., overlap, alignment, and combination), logical relationships (e.g., animation sequences, hyperlink jumps), semantic associations, and style inheritance between presentation document page elements. This simple data structure cannot support deep, relationship-based querying, intelligent modification, or complex generation tasks of presentation document content. For example, it is difficult to precisely answer the question, "Find all pages containing 'market analysis' with a growth trend chart next to them," or to uniformly replace the font style of all titles with the style of the template page.
[0019] Furthermore, existing tools often lack a comprehensive understanding of the semantics and visual layout of presentation content, making it difficult to implement intelligent content query, modification, style transfer, and automated generation. Users often need to perform manual operations, which is inefficient and prone to errors.
[0020] Traditional technologies for parsing presentation documents lack the expressive power of data structures, making it difficult to effectively and intuitively represent complex relationships between elements. For example, spatial layout relationships (relative positions (up and down, left and right, overlap, alignment), and distances between elements; visual style relationships (font, color, fill, and other style attributes shared by multiple elements); content semantic relationships (logical connections between different text boxes and charts); and interaction / animation relationships (jump logic, animation triggering sequence, and dependencies between elements). Inefficient queries and modifications: Based on simple data structures, complex queries (such as "find all text boxes containing the word 'core technology' and located on the right side of the page") or structural modifications (such as "change the font size of all titles and adjust the position of the first image below them") often require complex traversal and judgment logic, which is inefficient and prone to errors. Cross-format processing is difficult: Parsing results from different document formats vary in their representation methods, making it difficult to form a unified intermediate representation, hindering cross-format content migration and style application.
[0021] For this reason, traditional technologies (especially those that rely on simple data structures such as JSON) cannot effectively represent the complex relationships between elements within a document (such as spatial, stylistic, semantic, and interactive relationships), making it difficult to implement advanced applications such as deep content query, intelligent editing, style transfer, and cross-format processing.
[0022] In response to the defects in traditional technologies, the present disclosure proposes a document generation method that converts the document to be processed into the target document by utilizing the capabilities of the graph structure, thereby improving the efficiency of target document conversion. Figure 1 A process 100 according to an embodiment of a document generation method of the present disclosure is shown. The document generation method includes the following steps:
[0023] Step 101: Obtain the document to be processed.
[0024] In this embodiment, the to-be-processed document is a document to be processed. The to-be-processed document may include at least one page arranged in a certain order, and each page may have corresponding page elements. By dividing the to-be-processed document into two levels, namely, the page as a whole and the page elements, the to-be-processed document may be converted into a graph structure. The document to be processed may be of various types, such as a presentation, a spreadsheet software document, a word processing software document, or a Portable Document Format (PDF). In practical applications, the to-be-processed document may be any of PPT, Word, PDF, and Excel.
[0025] In this embodiment, the document to be processed can be obtained in various ways. For example, the document to be processed can be a document obtained from a server by the execution entity on which the document generation method is running through real-time communication with the server; the document to be processed can also be a document extracted from a database by the execution entity. Alternatively, the document to be processed can be a document uploaded by a user. By processing the document to be processed using the method disclosed herein, the target document required by the user can be obtained.
[0026] Step 102: parse the page element attributes and page element relationships of the document to be processed to obtain a page element parsing result.
[0027] In this embodiment, the document to be processed has multiple pages, each page has corresponding page elements. The page elements are the basic units that make up the page, and can be buttons, display content, images, and other units. The page elements of each page or between pages have corresponding relationships. For example, the button of the first page is used to control the display of the second page. The document to be processed can be effectively identified by parsing the page elements and the relationship between page elements.
[0028] In this embodiment, step 102 includes: using an artificial intelligence model that integrates natural language processing and computer vision to deeply analyze the document to be processed, obtaining page element analysis results, wherein the page element analysis results include: all pages of the document to be processed, page order, master pages, and layouts; page elements of each page, including but not limited to text boxes (TextBox), images (Picture), shapes (Shape), tables (Table), charts (Chart), SmartArt graphics, multimedia objects (video / audio), group objects (Group), etc.; attributes of each page element, wherein the attributes of each page element include: content attributes, including text content, image content (or its description / label), table data, chart data series and category, etc.; position and geometric attributes, including the precise coordinates (x, y), dimensions (width, height), and rotation angle of the page element on the page; and visual style attributes, including font (name, size, color, bold, italic, etc.), fill color, border style, shape type, transparency, shadow effect, etc. Structural and relationship properties include the page to which the page element belongs, its z-order, group relationships (whether it belongs to a group), hyperlink targets, animation effects (type, order, triggering method), and notes. Page element relationships include implicit relationships between page elements, such as spatial proximity, alignment, content relevance (for example, a text description is related to the adjacent chart data), and style similarity / inheritance.
[0029] In this embodiment, when the document to be processed is a document in PPT, Word, PDF, or Excel, the initial page nodes of PPT, Word, and PDF are generated and connected in sequence according to the physical order when the tool is designed and laid out, and Excel generates nodes and connects them in sequence according to the read sheet (the table displayed in the workbook window).
[0030] Step 103: construct an initial graph structure based on the page element parsing results.
[0031] In this embodiment, the execution entity on which the document generation method runs can put the summary, tags and physical serial number of the page summarized in the page content analysis results into the virtual graph node associated with the page node in the initial graph structure, so as to facilitate the user to perform in-depth semantic retrieval and operation.
[0032] In this embodiment, based on the page element parsing results, a graph data structure is constructed - the initial graph structure, which is used to represent the entire document to be processed. The initial graph structure is the core intermediate representation of this disclosure. Figure 2 As shown in the figure, A to F are nodes representing the pages or page elements of the document to be processed, and the edges between A to F are used to represent the relationship between the nodes. Figure 2 In the example, node E has corresponding relationships with nodes B, D, and F. Node A only has relationships with nodes B and F, and node C only has relationships with nodes B and D.
[0033] In this embodiment, the page element parsing results include: all pages and page attributes (such as page order, layout, etc.), page elements, page element attributes, and page element relationships. The page element attributes include the element's own attributes as well as structure and relationship attributes. The structure and relationship attributes are used to characterize the relationship between the element and the page, such as which page the page element belongs to. Step 103 includes: determining nodes based on all pages and page elements in the page element parsing results; determining edges based on the page element relationships and structure and relationship attributes in the page element parsing results; connecting each node using edges, and adding corresponding attribute information to the edges and nodes based on the attributes of the page elements in the page element parsing results to obtain an initial graph structure.
[0034] In this embodiment, the graph structure can naturally and richly represent the complex relationships (space, style, semantics, interaction, etc.) between elements in the document, overcoming the limitations of simple structures such as JSON. When the documents to be processed are of various types, step 103 can realize the conversion of supported input document formats (such as PPT, Word, Excel, PDF, etc.) to a unified intermediate graph structure, and can generate any supported output format from the graph structure, realizing seamless migration of content and format conversion.
[0035] Step 104 : Based on the document generation requirement input by the user, the initial graph structure is operated to obtain an operation graph structure.
[0036] In this embodiment, the document generation requirement is a design requirement for the target document. Since the initial graph structure supports efficient add, delete, modify, and query operations, the document generation requirement may include: adding information about adding pages or page elements of the document to be processed, deleting information about deleting pages or page elements of the document to be processed, changing information about changing pages or page elements in the document to be processed, and querying information about querying pages or page elements in the document to be processed.
[0037] In this embodiment, the operations performed on the initial graph structure include: query operations on the initial graph structure, modification operations on the initial graph structure, addition operations on the initial graph structure, and deletion operations on the initial graph structure.
[0038] Query operations on the initial graph structure include: attribute-based query operations, such as finding all text box nodes with the font "Microsoft YaHei". Content-based query operations, such as finding text nodes whose content contains specific keywords / synonyms (text matching or vector similarity can be used). Relationship-based query operations, such as finding text nodes that are spatially adjacent to a certain image node, or finding all element nodes that link to page 5. Complex graph pattern queries, such as finding page nodes that "contain title A and chart B, and chart B is below title A". Natural language query interface: Users can ask questions in natural language (such as "What is on page 3?", "Which pages mention 'user growth'?"), and the system converts them into graph query statements for execution.
[0039] Modifications to the initial graph structure include: modifying node attributes, such as changing the text content, font size, or color of a text node; modifying the position of an image node; and modifying edges, such as changing the page order (modifying the NextPage edge) and changing the hyperlink destination (modifying the HyperlinkTo edge).
[0040] Operations that add to the initial graph structure include: adding new nodes, such as adding a new text box node or image node under a specified page node; adding new edges, and establishing relationship edges between new elements and pages or other elements.
[0041] Deletion operations on the initial graph structure include: node deletion operations, such as deleting a page node or an element node on a page (and processing the associated edges at the same time), and edge deletion operations, such as removing a relationship between elements.
[0042] Optionally, for some complex requirements of users, the document generation requirement may be a combination of the above-mentioned multiple aspects of information, and for each aspect of the document generation requirement, the operation mode of that aspect may be determined.
[0043] In this embodiment, the above step 104 includes: determining modification information of the document to be processed based on the document generation requirements input by the user; determining the modification content of the initial graph structure based on the modification information, and modifying the initial graph structure according to the modification content to obtain the operation graph structure.
[0044] Step 105: Obtain the target document based on the operation graph structure.
[0045] In this embodiment, after operating on the initial graph structure, an operation graph structure is obtained. Based on the operation graph structure or new content generated based on the operation graph structure, the new content can be re-rendered into a document in a target format. The target document can be of the same type as the document to be processed, or of a different type than the document to be processed. When the target document and the document to be processed are of different types, the target document generated from the document to be processed is a document that has undergone style conversion.
[0046] In a specific example, the document to be processed is a PPT, and the graph structure information is converted back into a target document in a PPT file format, preserving the content, layout, style, and relationship.
[0047] In this embodiment, the target document can be the document obtained after performing any of the following steps: text replacement, content query and filtering, or style transfer on the existing document to be processed. In a specific example of text replacement, the user provides a document generation requirement: a replacement rule for an existing document to be processed. The system parses the PPT into a graph structure, searches for all text nodes containing "Company A," modifies their content attributes to "Company B," and then renders the modified graph structure back into a new PPT file.
[0048] In a specific example of content query and filtering, the user's document generation requirement is: asking "What is the main content of page 5?" or "Which pages discuss 'marketing strategy'?" The system parses the PPT into a graph structure, executes the corresponding graph query (queries the main content nodes under the page 5 node, or queries the nodes containing content related to "marketing strategy" and traces back to the page node), and returns the results (text description or page list and sorting).
[0049] In a specific example of style transfer, a user uploads a PPT (as a style template) and new content (e.g., Markdown text). The system parses the template PPT, extracting the layout pattern (relative position relationships), font style (node attributes), color scheme, and other information from its graph structure. It then creates a preliminary graph structure based on the new content and applies the extracted style information to the corresponding nodes and relationships in the new graph, ultimately generating a new PPT that matches the template style.
[0050] The document generation method provided by the embodiment of the present disclosure first obtains the document to be processed; secondly, parses the page element attributes and page element relationships of the document to be processed to obtain the page element parsing result; thirdly, constructs an initial graph structure based on the page element parsing result; thirdly, operates on the initial graph structure based on the document generation requirements input by the user to obtain the operation graph structure; finally, obtains the target document based on the operation graph structure. Thus, by performing page element and relationship parsing on the document to be processed by the artificial intelligence model and constructing a graph structure, the content, attributes and complex spatial, logical and semantic relationships between page elements in the document to be processed can be accurately and richly represented based on the ability of the graph structure; by operating the initial graph structure and converting it into the operation graph structure of the target document, and converting it into the target document based on the operation graph structure, the graph structure can be automatically used as a universal intermediate representation to realize cross-format content integration and conversion, and improve the efficiency of target document generation compared to manual operation of other data structures.
[0051] In some optional implementations of the present disclosure, the above-mentioned parsing of page element attributes and page element relationships of the document to be processed to obtain page element parsing results includes: using an artificial intelligence model to perform page layout parsing on the document to be processed to obtain element parsing results and relationship parsing results; based on the element parsing results, identifying the page elements in each page and the key attributes of the page elements; based on the relationship parsing results, determining the association relationships of all page elements in the document to be processed; and using the page elements, key attributes and association relationships as the page element parsing results.
[0052] In this optional implementation, the AI model can integrate natural language processing and computer vision. Specifically, the AI model can identify both text and image content within the document being processed, thereby comprehensively recognizing the document. Combined with AI technology, this enables a more accurate and comprehensive understanding of the document's visual layout, content semantics, and implicit structure.
[0053] In this optional implementation, the above-mentioned local page parsing of the document to be processed to obtain element parsing results and relationship parsing results includes: parsing the page information of all pages of the document to be processed and the page relationship between any two pages, and for each page among all pages, parsing the layout information of the page elements in the page; based on the layout information, determining the element information of the page elements and the element relationship between any two page elements, taking the page information of all pages and the element information of the page elements on each page as the element parsing result, and taking the page relationship and the element relationship as the relationship parsing result.
[0054] In this optional implementation, based on the element parsing results, identifying the page elements in each page and the key attributes of the page elements includes: extracting the page information of each page from the element parsing results, and determining the page elements belonging to the page for each page information; based on the element information of the page elements in the element parsing results, extracting key information with other pages and page elements as key attributes.
[0055] In this optional implementation, the above-mentioned determination of the association relationship of all page elements in the document to be processed based on the relationship analysis results includes: clustering page relationships belonging to the relationship of the same page, and obtaining the clustering type of the page; based on the clustering type, determining the related pages in the page, and establishing a first association relationship related to the clustering type for the page and the related pages; clustering element relationships belonging to the relationship of the same element, and obtaining the clustering type of the element; based on the clustering type, determining the related elements in the element, and establishing a second association relationship related to the clustering type for the element and the related elements, and using the first association relationship and the second association relationship as the association relationship of all page elements.
[0056] The method for obtaining page element parsing results provided by this optional implementation method performs page layout parsing on the document to be processed, and can deeply analyze the document to be processed from the content and external layout, thereby improving the accuracy of the parsing of the document to be processed; based on the element parsing results and relationship parsing results parsed by the artificial intelligence model, key attributes and association relationships are determined respectively, thereby improving the reliability of the page element parsing results.
[0057] In some optional implementations of the present disclosure, the above-mentioned parsing of page element attributes and page element relationships of the document to be processed to obtain the page element parsing results includes: determining the task information of different intelligent agents among all intelligent agents based on the document to be processed; sending the task information to the corresponding intelligent agent among all intelligent agents, so that each intelligent agent performs the parsing of its own page elements; obtaining and merging the parsing results of all intelligent agents to obtain the page element parsing results.
[0058] In this optional implementation, the page element parsing result can be the information parsed directly after the actual content and relationship identification of the document to be processed. The page element parsing result can also be the summary information obtained after intelligent summarization by the intelligent agent. This summary information can facilitate the generation of virtual graph nodes in the initial graph structure.
[0059] In this optional implementation, an agent is an agent capable of perceiving its environment and taking actions to achieve specific goals. Agents possess autonomy, adaptability, and interaction, enabling them to perceive their environment, make decisions, and execute tasks. They perceive changes in the environment, make judgments and decisions based on learned knowledge and algorithms, and then execute actions to influence the environment or achieve predetermined goals.
[0060] In this optional implementation, for large target document generation tasks, the tasks can be broken down, task information can be obtained, and assigned to different agents. For example, Agent A is responsible for generating text content based on the outline (manipulating text nodes), Agent B is responsible for generating accompanying images based on the text description (manipulating image nodes), and Agent C is responsible for the overall layout and style (manipulating position, style attributes, and relationships). These agents can operate on different parts of the graph structure or page ranges in parallel and finally merge the results, significantly improving generation efficiency and quality while leveraging the expertise of different agents.
[0061] The method for obtaining the page element parsing results provided by this optional implementation implements the parsing of the page elements of the document to be processed through multiple intelligent agents, giving full play to the expertise of each intelligent agent and improving the reliability of the page element parsing results.
[0062] Optionally, the operation graph structure can also be a graph structure generated by multiple agents. In a specific example, the user provides a PPT outline containing 100 pages of content. The system assigns tasks to multiple agents based on the outline structure and content type (such as the text agent processes pages 1-30, the graphic agent processes pages 31-70, and the data visualization agent processes the parts involving charts on pages 71-100). Each agent generates and edits the content of the initial graph structure in parallel within the scope of the graph structure it is responsible for. Finally, the coordinating agent merges the various parts of the graph structure and combines the graph structures of all agents to obtain the operation graph structure.
[0063] In some embodiments of the present disclosure, the above-mentioned construction of the initial graph structure based on the page element parsing results includes: determining nodes based on the page elements in the page element parsing results; determining edges based on the association relationships in the page element parsing results; adding key attributes in the page element parsing results to the attributes of the nodes, and connecting the nodes according to the association relationships to obtain the initial graph structure.
[0064] In this optional implementation, for a document to be processed with only one page, the page element parsing results may include all page elements in the page, and the page elements are used as nodes of the initial graph structure; when there is only one page of a document to be processed, the relationship between any two page elements on the page is an association relationship, and the nodes with the association relationship are connected together to determine the edge.
[0065] In this optional implementation, the key attributes are attributes of the page or page elements. Adding the attributes to the attributes of the nodes in the initial graph structure can effectively interpret the nodes in the initial graph structure.
[0066] The method for constructing a graph structure provided by this optional implementation method determines nodes based on page elements in the page element parsing results; determines edges based on the association relationships in the page element parsing results; adds key attributes in the page element parsing results to the attributes of the nodes, and connects the nodes according to the association relationships to obtain the initial graph structure, providing a reliable implementation method for obtaining the initial graph structure.
[0067] Optionally, for a document to be processed with multiple pages, the page element parsing results include: multiple pages, page attributes of the pages, page elements on each page, element attributes of each page element, and association relationships; wherein the association relationships in the page element parsing results include: the relationship between any two pages in the document to be processed and the relationship between any two page elements, and the page attributes and element attributes constitute key attributes. The above-mentioned construction of the initial graph structure based on the page element parsing results includes: determining nodes based on the pages and page elements in the page element parsing results; determining edges corresponding to each page and page element based on the association relationships in the page element parsing results; adding the key attributes in the page element parsing results to the attributes of the nodes, and connecting the nodes according to the association relationships to obtain the initial graph structure.
[0068] In some embodiments of the present disclosure, the above-mentioned document generation requirements based on user input are used to operate the initial graph structure to obtain the operation graph structure, including: determining the target elements of the target document, the association relationships of the target elements, and the attributes of the target elements based on the document generation requirements input by the user; modifying the initial graph structure based on the target elements of the target document, the association relationships of the target elements, and the attributes of the target elements to obtain the operation graph structure.
[0069] In this optional implementation, the target document is the document to be converted into, and the target element is the page element that needs to be generated in the target document as expected by the user's document generation requirements. The above-mentioned determination of the target elements of the target document, the association relationship of the target elements, and the attributes of the target elements based on the document generation requirements input by the user includes: performing keyword recognition on the document generation requirements input by the user to determine the target elements in the document generation requirements; and using a deep learning model to perform keyword recognition on the requirements of each target element in the document generation requirements to determine the association relationship and attributes of each target element.
[0070] In this optional implementation, the above-mentioned modification of the initial graph structure based on the target elements of the target document, the association relationships of the target elements, and the attributes of the target elements to obtain the operation graph structure includes: querying the mid-nodes of the initial graph structure based on the target elements of the target document, the association relationships of the target elements, and the attributes of the target elements to obtain the node query results; generating new node information or new node attributes of the target node based on the node query results; replacing the node information of the target node in the initial graph structure with the new node information, and replacing the node attributes of the target node in the initial graph structure with the new node attributes to obtain the operation graph structure.
[0071] The method for obtaining the operation graph structure provided by this optional implementation first determines the target elements of the target document, the association relationships of the target elements, and the attributes of the target elements based on the document generation requirements input by the user; then, based on the target elements of the target document, the association relationships of the target elements, and the attributes of the target elements, the initial graph structure is modified to obtain the operation graph structure. Thus, based on the user's needs, the initial graph structure is modified to obtain the operation graph structure, which can make the operation graph structure more in line with user needs and improve the accuracy of the generation of the operation graph structure.
[0072] Optionally, the above-mentioned document generation requirements based on user input operate on the initial graph structure to obtain an operation graph structure including: determining the requirement nodes and requirement edges based on the document generation requirements input by the user; detecting whether the initial graph structure has the requirement nodes and requirement edges; in response to detecting that the initial graph structure does not have the requirement nodes and requirement edges, modifying the initial graph structure so that the modified initial graph structure has the requirement nodes and requirement edges, and using the modified initial graph structure as the operation graph structure.
[0073] In some embodiments of the present disclosure, the above method also includes: obtaining newly added translation node information; determining the node to be operated in the initial graph structure based on the newly added translation node information; using a translation module to translate the node to be operated to obtain an updated node, and adding the updated node to the initial graph structure.
[0074] In this optional implementation, the newly added translation node information is information prompting to add a translation node. The newly added translation node can be used to determine the node that needs to be translated, that is, the node to be operated.
[0075] In this optional implementation, the translation module is a module that optimizes or interprets the nodes to be operated. The updated nodes obtained by the translation module can better interpret the nodes to be operated. The translation module can be a tool or an API (Application Programming Interface). The nodes in the graph structure can serve as anchor points for interaction with external tools / APIs. For example, a text node is selected, and the translation API is called to translate and then update the node content. The above-mentioned tool can be a built-in or dynamically connected tool. The tool acts as a function acting on the graph nodes / edges to realize arbitrary modification, enhancement and generation of the graph content.
[0076] The document generation method provided by this optional implementation method obtains newly added translation node information; based on the newly added translation node information, determines the nodes to be operated in the initial graph structure; uses a translation module to translate the nodes to be operated to obtain updated nodes, and adds the updated nodes to the initial graph structure, providing a reliable implementation method for improving and enhancing the initial graph structure, and improving the reliability and accuracy of the initial graph structure.
[0077] In some embodiments of the present disclosure, the above method also includes: obtaining newly added picture node information; determining the target node area in the initial graph structure based on the newly added picture node information; using the picture generation module to generate picture nodes, and inserting the picture nodes into the target node area.
[0078] In this optional implementation, the newly added picture node information is information prompting the addition of picture nodes. The newly added picture node information can be used to determine the area where additional picture nodes need to be added, that is, the target node area. The target node area can include one or more nodes related to the newly added picture node.
[0079] In this optional implementation, the image generation module directly generates new image nodes. The image generation module can be a tool or API. The image nodes generated by the image generation module can better interpret the nodes in the target node area. Selecting a region node, calling the image generation API to generate an image, and inserting it as a new image node can effectively expand the initial graph structure. This tool can be built-in or dynamically accessed. This tool acts as a function on graph nodes / edges to enable arbitrary modification, enhancement, and generation of graph content.
[0080] The document generation method provided by this optional implementation method obtains the newly added image node information; based on the newly added image node information, determines the target node area in the initial graph structure; uses the image generation module to generate image nodes, and inserts the image nodes into the target node area, providing another reliable implementation method for improving and enhancing the initial graph structure, thereby improving the reliability and accuracy of the initial graph structure.
[0081] Optionally, the above method also includes: receiving content enhancement and processing content, processing the operation graph structure based on the content enhancement and processing content to obtain a new operation graph structure, and obtaining a target document based on the new operation graph structure. The content enhancement and processing content can be information obtained from an external interface based on user operations. The specific example is as follows: the user uploads PPT (Input.pptx), and the system generates an operation graph structure Graph_A. The content enhancement and processing content is: a text box node N on page 3 selected by the user, and an external "text summary" API service is selected. The system extracts the content attributes (text) of node N. Call the external API and send the text to the summary service. Receive the summary text returned by the API. Choose to update the content of node N to summary text, or add a new text node whose content is a summary, and establish its relationship with node N or the page it is on in the operation graph structure.
[0082] Further references Figure 3 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a document generation device, which is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0083] like Figure 3 As shown, the document generation device 300 provided in this embodiment includes: an acquisition unit 301, a parsing unit 302, a construction unit 303, an operation unit 304, and a obtaining unit 305. Among them, the above-mentioned acquisition unit 301 can be configured to obtain the document to be processed. The above-mentioned parsing unit 302 can be configured to parse the page element attributes and page element relationships of the document to be processed to obtain the page element parsing results. The above-mentioned construction unit 303 can be configured to construct an initial graph structure based on the page element parsing results. The above-mentioned operation unit 304 can be configured to operate on the initial graph structure based on the document generation requirements input by the user to obtain an operation graph structure. The above-mentioned obtaining unit can be configured to obtain the target document based on the operation graph structure.
[0084] In this embodiment, the document generation device 300 includes the acquisition unit 301, the parsing unit 302, the construction unit 303, the operation unit 304, and the obtaining unit 305. The specific processing and technical effects thereof can be referred to in the respective Figure 1The relevant descriptions of step 101, step 102, step 103, step 104 and step 105 in the corresponding embodiment are not repeated here.
[0085] In some embodiments of the present disclosure, the above-mentioned parsing unit 302 is configured to: use an artificial intelligence model to perform page layout analysis on the document to be processed to obtain element parsing results and relationship parsing results; based on the element parsing results, identify the page elements in each page and the key attributes of the page elements; based on the relationship parsing results, determine the association relationship of all page elements in the document to be processed; and use the page elements, key attributes and association relationships as the page element parsing results.
[0086] In some embodiments of the present disclosure, the above-mentioned parsing unit 302 is configured to: determine the task information of different agents among all agents based on the document to be processed; send the task information to the corresponding agent among all agents so that each agent can parse its own page elements; obtain and merge the parsing results of all agents to obtain the page element parsing results.
[0087] In some embodiments of the present disclosure, the above-mentioned construction unit 303 is configured to: determine the nodes based on the page elements in the page element parsing results; determine the edges based on the association relationships in the page element parsing results; add the key attributes in the page element parsing results to the attributes of the nodes, and connect the nodes according to the association relationships to obtain the initial graph structure.
[0088] In some embodiments of the present disclosure, the above-mentioned operation unit 304 is configured to: determine the target elements, the association relationships of the target elements, and the attributes of the target elements of the target document based on the document generation requirements input by the user; modify the initial graph structure based on the target elements, the association relationships of the target elements, and the attributes of the target elements of the target document to obtain the operation graph structure.
[0089] In some embodiments of the present disclosure, the above-mentioned device 300 also includes: a translation unit (not shown in the figure), which is configured to: obtain newly added translation node information; determine the nodes to be operated in the initial graph structure based on the newly added translation node information; use the translation module to translate the nodes to be operated to obtain updated nodes, and add the updated nodes to the initial graph structure.
[0090] In some embodiments of the present disclosure, the above-mentioned device 300 also includes: a picture adding unit (not shown in the figure), and the above-mentioned picture adding unit is configured to: obtain new picture node information; determine the target node area in the initial graph structure based on the new picture node information; use the picture generation module to generate picture nodes, and insert the picture nodes into the target node area.
[0091] The document generation device provided by the embodiment of the present disclosure is as follows: first, the acquisition unit 301 acquires the document to be processed; second, the parsing unit 302 performs page element attribute and page element relationship parsing on the document to be processed to obtain the page element parsing result; third, the construction unit 303 constructs an initial graph structure based on the page element parsing result; third, the operation unit 304 operates on the initial graph structure based on the document generation requirement input by the user to obtain the operation graph structure; finally, the acquisition unit 305 obtains the target document based on the operation graph structure. Thus, by performing page element and relationship parsing on the document to be processed by the artificial intelligence model and constructing the graph structure, the content, attributes and complex spatial, logical and semantic relationships between page elements in the document to be processed can be accurately and richly represented based on the ability of the graph structure; by operating the initial graph structure and converting it into the operation graph structure of the target document, and converting it into the target document based on the operation graph structure, the graph structure can be automatically used as a universal intermediate representation to realize cross-format content integration and conversion, thereby improving the efficiency of target document generation compared to manual operation of other data structures.
[0092] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0093] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are provided for example only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0094] like Figure 4 As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0095] Various components in device 400 are connected to I / O interface 405, including an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0096] The computing unit 401 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the document generation method. For example, in some embodiments, the document generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the document generation method described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the document generation method by any other appropriate means (e.g., by means of firmware).
[0097] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0098] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable document generation device so that when the program code is executed by the processor or controller, the modes / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0099] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0100] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0101] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0102] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0103] The foregoing descriptions of specific exemplary embodiments of the present disclosure are for purposes of illustration and description. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many variations and modifications are possible in light of the foregoing teachings. The exemplary embodiments have been selected and described for the purpose of explaining the specific principles of the present disclosure and their practical application, thereby enabling those skilled in the art to realize and utilize a variety of exemplary embodiments of the present disclosure and various options and modifications. The scope of the present disclosure is intended to be defined by the claims and their equivalents.
Claims
1. A document generation method, comprising: Get documents to be processed; Analyzing page element attributes and page element relationships of the document to be processed to obtain a page element analysis result; Constructing an initial graph structure based on the page element parsing results; Based on the document generation requirements input by the user, the initial graph structure is operated to obtain an operation graph structure; Based on the operation graph structure, a target document is obtained.
2. The method according to claim 1, wherein The performing of page element attribute and page element relationship analysis on the document to be processed to obtain the page element analysis result includes: Utilizing an artificial intelligence model to perform page layout analysis on the document to be processed, and obtaining element analysis results and relationship analysis results; Based on the element parsing results, identifying page elements and key attributes of the page elements in each page; Based on the relationship analysis result, determining the association relationship of all page elements in the document to be processed; The page element, the key attribute and the association relationship are used as the page element parsing result.
3. The method according to claim 1, wherein The performing of page element attribute and page element relationship analysis on the document to be processed to obtain the page element analysis result includes: Determining task information of different agents among all agents based on the document to be processed; Sending the task information to corresponding agents among all agents, so that each agent can parse its own page elements; Obtain and merge the parsing results of all agents to obtain the page element parsing results.
4. The method according to claim 1, wherein The constructing of the initial graph structure based on the page element parsing result includes: Determine the node based on the page element in the page element parsing result; Determining edges based on association relationships in the page element parsing results; The key attributes in the page element parsing result are added to the attributes of the node, and the nodes are connected according to the association relationship to obtain an initial graph structure.
5. The method according to claim 1, wherein The document generation requirement input by the user and the operation of the initial graph structure to obtain the operation graph structure include: Based on the document generation requirements input by the user, determine the target elements of the target document, the association relationships between the target elements, and the attributes of the target elements; Based on the target elements of the target document, the association relationships between the target elements, and the attributes of the target elements, the initial graph structure is modified to obtain an operation graph structure.
6. The method according to any one of claims 1 to 5, further comprising: Get the newly added translation node information; Determining a node to be operated in the initial graph structure based on the newly added translation node information; A translation module is used to translate the node to be operated to obtain an updated node, and the updated node is added to the initial graph structure.
7. The method according to any one of claims 1 to 5, further comprising: Get the newly added image node information; Determining a target node region in the initial graph structure based on the newly added image node information; An image node is generated by using an image generation module, and the image node is inserted into the target node area.
8. A document generation device, comprising: an acquiring unit, configured to acquire a document to be processed; a parsing unit configured to parse page element attributes and page element relationships of the document to be processed to obtain a page element parsing result; A construction unit configured to construct an initial graph structure based on the page element parsing result; an operating unit configured to operate the initial graph structure based on a document generation requirement input by a user to obtain an operation graph structure; The obtaining unit is configured to obtain a target document based on the operation graph structure.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for abstracting document structure
CN102982010A
Page layout method and device, equipment and medium
CN114579912A
Method and device for judging page display state, electronic equipment and storage medium
CN117112952A
Demonstration document generation method and device, electronic equipment and storage medium
CN119441157A
Method and apparatus relating to webpages and real estate information
WO2008040046A1