Active enhanced document understanding system, method and equipment and medium

By designing an actively enhanced document understanding system, using the intelligent collaboration technology to deeply analyze and establish an association network for documents, it solves the problem that traditional document processing systems cannot deeply understand complex document content, and achieves efficient and accurate document processing and information correlation.

CN119990095APending Publication Date: 2025-05-13BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510051401.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When facing complex and changing document content, traditional document processing systems cannot deeply understand the implicit information and deep meanings in the document, and fail to establish a network of correlations between document elements, affecting the accuracy and efficiency of document processing.

Method used

An actively enhanced document understanding system was designed, and the PDF files were initially parsed through the online document understanding module to generate lightweight md files. The document actively enhances the intellectual leader, intellectual member and knowledge integration agent in the understanding module, decomposes the document deep understanding task, conducts preliminary analysis, special understanding and knowledge integration, and establishes a correlation network between document elements.

Benefits of technology

By deeply analyzing document content, the system can more accurately capture key information in the document, improve the accuracy and efficiency of document processing, realize deep fusion and semantic correlation of information, and is suitable for the processing and display of complex information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990095A_ABST
    Figure CN119990095A_ABST
Patent Text Reader

Abstract

The invention discloses an actively enhanced document understanding system, method and device and a medium. And the md file output by the online document understanding module is deeply analyzed through the document active enhancement understanding module of the system, so that the understanding and processing quality of the document content is further improved. According to the document active enhancement understanding module, a document deep understanding task is decomposed into three stages of preliminary analysis, special understanding task processing and result integration. The agent group leader is responsible for preliminary analysis and task allocation of the document, and more accurate and deep analysis is carried out on different content elements in the document content. A plurality of agent group members respectively complete special understanding tasks under the scheduling of the agent group leader, efficient document understanding and information processing are realized, the knowledge integration agent is responsible for result integration, summarization and integration of special document understanding results, deep fusion and semantic association of information are realized, and the method is suitable for processing and display of complex information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing and deep learning technology, and in particular to an actively enhanced document understanding system, method, device and medium. Background Art

[0002] In today's information society, with the rapid development of information technology, document processing has become an indispensable part of people's daily life and work. As an important carrier of information transmission and storage, the improvement of document understanding and processing capabilities is of great significance to improving work efficiency and optimizing decision-making processes. However, traditional document processing systems often seem to be unable to cope with complex and changing document content.

[0003] At present, the existing technology mainly converts PDF documents into md format files that can be understood by large language models through algorithm models. Although this conversion process solves the problems of document readability and parsability to a certain extent, making it easier for computers to process document content, it still has many shortcomings.

[0004] First, the conversion result is only to convert the document from one format to another, without further understanding of the document content and its internal elements. In other words, the existing conversion technology is still at the surface level and cannot dig deep into the implicit information and deep meaning of the document. This results in the system often failing to accurately capture key information when processing some highly professional and complex documents, thus affecting the accuracy and efficiency of document processing.

[0005] Secondly, the relationship between document elements is not fully reflected in the final MD file. In a document, there are often close connections and logical relationships between various parts. These relationships are crucial to understanding the overall structure and intent of the document. However, existing conversion technologies often ignore this point and simply arrange and display the document content in a certain format without establishing a network of associations between the various elements. This results in the system often being unable to give satisfactory answers when dealing with some tasks that require understanding the overall structure and intent of the document. Summary of the invention

[0006] The present application provides an actively enhanced document understanding system, method, device and medium for solving the problem that traditional document processing systems simply arrange and display document contents in a certain format without establishing an association network between various elements, thereby affecting subsequent document understanding.

[0007] In a first aspect, the present application provides an actively enhanced document understanding system, the system comprising:

[0008] An online document understanding module, used for receiving a PDF file uploaded by a user, and performing online document understanding on the PDF file to obtain a lightweight markup language md file of the PDF file;

[0009] The document active enhancement understanding module is connected to the online document understanding module, and is used to receive unprocessed md files and perform document deep understanding tasks on them; wherein, the document active enhancement understanding module includes: an agent group leader, multiple agent group members and a knowledge integration agent;

[0010] The agent group leader is connected to the plurality of agent group members and the knowledge integration agent respectively, and is used to receive an unprocessed md file input, and perform a preliminary parsing task on the unprocessed md file to parse and obtain a document object corresponding to the unprocessed md file;

[0011] The plurality of intelligent agent members are used to receive the input of the document object and perform their own corresponding special understanding tasks on the document object to obtain special document understanding results;

[0012] The knowledge integration agent is used to receive the input of the understanding results of each of the special documents of the unprocessed md file, and perform the knowledge integration task on each of the special document understanding results to obtain the document in-depth understanding results;

[0013] The result merging and displaying module is used to display the md file generated by the online document understanding module and the document deep understanding results generated by the document active enhanced understanding module.

[0014] In a second aspect, the present application provides a method for actively enhancing document understanding, the method comprising:

[0015] S1: receiving a PDF file uploaded by a user through an online document understanding module, and performing online document understanding on the PDF file to obtain a lightweight markup language md file of the PDF file;

[0016] S2: The md file is transferred to the document active enhancement understanding module, and the following steps are performed in sequence:

[0017] S21: The intelligent agent group leader in the active enhanced understanding module receives the input of the md file and performs a preliminary parsing task on the md file to parse and obtain the document object corresponding to the md file;

[0018] S22: Distribute the document object to multiple agent members in the active enhanced understanding module, and the agent members respectively perform their own special understanding tasks on the received document objects to generate special document understanding results;

[0019] S23: the knowledge integration agent in the active enhanced understanding module receives the input of the understanding results of each of the special documents of the unprocessed md file, and performs a knowledge integration task on each of the special document understanding results to obtain a document in-depth understanding result;

[0020] S3: The md file generated by the online document understanding module and the document deep understanding result generated by the document active enhancement understanding module are displayed through the result merging and displaying module.

[0021] In a third aspect, the present application provides a computer device, the computer device comprising a processor, the processor being configured to implement the steps of the active enhanced document understanding method as described above when executing a computer program stored in a memory.

[0022] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the active enhanced document understanding method as described above.

[0023] The beneficial effects of this application are as follows:

[0024] 1. The online document understanding module performs preliminary analysis on the PDF files uploaded by users and quickly generates lightweight md files. This lightweight processing reduces the complexity of subsequent tasks and improves processing efficiency. At the same time, the document active enhancement understanding module further improves the understanding and processing quality of document content by deeply analyzing the md files output by the online document understanding module.

[0025] 2. Since the document active enhanced understanding module decomposes the document deep understanding task into three stages: preliminary analysis, special understanding task processing and result integration. The agent leader of the document active enhanced understanding module is responsible for the preliminary analysis and task allocation of the document, and multiple agent members of the document active enhanced understanding module complete special understanding tasks respectively. The knowledge integration agent of the document active enhanced understanding module is used to summarize and integrate special document understanding results. Among them, multiple agent members independently perform their respective special understanding tasks, which significantly improves the task processing speed of the system. And through the special understanding task processing of the agent members, the system can perform more accurate and in-depth analysis of different content elements in the document content. Through the task scheduling of the agent leader and the result integration of the knowledge integration agent, efficient document understanding and information processing are achieved, and deep fusion and semantic association of information are achieved, which is suitable for the processing and display of complex information. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0027] Figure 1 A schematic diagram of the structure of an active enhanced document understanding system provided in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of the internal processing flow of a document active enhancement comprehension module provided in an embodiment of the present application;

[0029] Figure 3 A schematic diagram of the structure of a document in-depth understanding result provided in an embodiment of the present application;

[0030] Figure 4 A schematic diagram of the workflow of a specific document active enhancement comprehension module provided in an embodiment of the present application;

[0031] Figure 5 A schematic diagram of the structure of an intelligent agent provided in an embodiment of the present application;

[0032] Figure 6 A schematic diagram of a process of actively enhancing document understanding provided by an embodiment of the present application;

[0033] Figure 7 It is a structural schematic diagram of a computer device provided in an optional embodiment of the present application. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0035] In order to deeply understand the content of a document and establish associations between content elements in the document, the present application provides an actively enhanced document understanding system, method, device and medium.

[0036] Embodiment 1:

[0037] This application provides an actively enhanced document understanding system. Figure 1 A schematic diagram of the structure of an active enhanced document understanding system provided in an embodiment of the present application, the system comprising:

[0038] The online document understanding module 11 is used to receive the PDF file uploaded by the user and perform online document understanding on the PDF file to obtain a lightweight markup language md file of the PDF file;

[0039] The document active enhancement understanding module 12 is connected to the online document understanding module 11, and is used to receive the unprocessed md file and perform the document deep understanding task on it; wherein, the document active enhancement understanding module 12 includes: an agent group leader 121, multiple agent group members 122 and a knowledge integration agent 123;

[0040] The agent group leader 121 is connected to the plurality of agent group members 122 and the knowledge integration agent 123 respectively, and is used to receive the input of the unprocessed md file, and perform a preliminary parsing task on the unprocessed md file to parse and obtain the document object corresponding to the unprocessed md file;

[0041] The plurality of intelligent agent members 122 are used to receive the input of the document object and perform their own corresponding special understanding tasks on the document object to obtain special document understanding results;

[0042] The knowledge integration agent 123 is used to receive the input of the understanding results of each of the special documents of the unprocessed md file, and perform the knowledge integration task on each of the special document understanding results to obtain the document in-depth understanding results;

[0043] The result merging and displaying module 13 is used to display the md file generated by the online document understanding module 11 and the document in-depth understanding result generated by the document active enhanced understanding module 12.

[0044] The present application provides an active enhanced document understanding system, which is mainly composed of an online document understanding module 11, a document active enhanced understanding module 12 and a result merging and display module 13, and each module works together to achieve in-depth understanding of the document and result display. Among them, the online document understanding module 11 is connected to the document active enhanced understanding module 12 and the result merging and display module 13 respectively, and the document active enhanced understanding module 12 is connected to the result merging and display module 13. The following is a detailed introduction to each module included in the active enhanced document understanding system:

[0045] 1. Online document understanding module 11.

[0046] The online document understanding module 11 is responsible for receiving the PDF file uploaded by the user, and extracting text, pictures and other elements by parsing the structure of the PDF file, and converting it into an md file (Markdown file) in a lightweight markup language format for subsequent processing.

[0047] In a possible implementation, the online document understanding module 11 also supports saving the generated md file as an unprocessed md file into a database so that the document active enhancement understanding module 12 can process it during the idle period of the online document understanding module 11 .

[0048] 2. Document active enhancement understanding module 12.

[0049] The document active enhancement understanding module 12 is the core innovation of the whole system. It uses multi-agent collaboration technology to deeply understand and enhance the unprocessed md file generated by the online document understanding module 11. The following is a more detailed introduction to the document active enhancement understanding module 12:

[0050] 2.1 Multi-agent collaboration framework.

[0051] The document active enhanced understanding module 12 is composed of a group of agents, including an agent leader 121, multiple agent members 122 (Agent Workers) and a knowledge integration agent.

[0052] 1) For the agent leader 121, the agent leader 121 receives the input of the unprocessed md file and performs the preliminary parsing task to extract the basic structure and content information of the md file to form a document object. Among them, the document object is a logical abstraction of the document content, including the chapter division, paragraph content, picture information, table data, etc. of the document. Then the document object is distributed to each agent member 122.

[0053] In a possible implementation, the agent group leader 121 can obtain a list of unprocessed md files from a database, and obtain corresponding unprocessed md files from the database based on document identifiers in the list of unprocessed md files, thereby ensuring that the system can efficiently process the unprocessed md files stored in the database.

[0054] 2) For multiple agent members 122, each agent member 122 is pre-configured with a specific specialized understanding task, such as summary extraction, paragraph relationship analysis, table understanding, image recognition, etc. When each agent member 122 receives a document object under the guidance of the agent team leader 121 based on the above embodiment, it can perform its own specific specialized understanding task on the document object to obtain a specialized document understanding result. Exemplarily, for any agent member 122, the agent member 122 can perform its own specific specialized understanding task on the document object based on a pre-trained deep learning large model (such as a large language model or a multimodal large model, etc.) for processing the specialized understanding task to obtain a specialized document understanding result.

[0055] In a possible implementation, the multiple agent members 122 may be configured with corresponding special understanding tasks according to the content elements and the relationships between the content elements contained in the document. Exemplarily, considering that the content elements contained in the document include abstracts, paragraphs, tables, pictures, formulas, and knowledge graphs, based on this, the multiple agent members 122 are responsible for performing the following special understanding tasks: abstract extraction, paragraph understanding, table content analysis, picture content understanding, formula content analysis, knowledge graph extraction, and relationship analysis between content elements.

[0056] In one example, after obtaining a special document understanding result, any agent member 122 can save the special document understanding result into a database for subsequent integration.

[0057] Figure 2 This is a schematic diagram of the internal processing flow of the document active enhancement understanding module 12 provided in the embodiment of the present application. Figure 2 As shown, the unprocessed md file is input into the agent leader 121. The agent leader 121 performs a preliminary parsing task on the md file, extracts the basic structure and content information of the md file, and forms a document object. The agent leader 121 then distributes the document object to each agent member 122. For any agent member 122, the agent member 122 can perform its own specific special understanding task on the document object based on a pre-trained deep learning large model (such as a large language model or a multimodal large model, etc.) for processing the special understanding task, so as to obtain the special document understanding result and save it in the database for subsequent integration and use.

[0058] 3) For the knowledge integration agent 123, the knowledge integration agent 123 can, under the arrangement of the agent group leader 121, integrate the special document understanding results generated by multiple agent members 122, and combine them with the document objects provided by the agent group leader 121 to form a deep understanding result of the document. Exemplarily, the knowledge integration agent 123 can receive input of various special document understanding results associated with a certain unprocessed md file, and perform knowledge integration tasks on each special document understanding result to obtain a deep understanding result of the document. For example, the table parsing results can be associated with the paragraph understanding results to show the contextual meaning of the table data; at the same time, the extracted knowledge graph can reveal the logical relationship between the various elements of the document.

[0059] In one example, the document deep understanding result is presented in a visual form, describing the attributes of each content element in the unprocessed md file and the relationship between them.

[0060] For example, in the database, the document depth understanding result of each document is a record, and the special document understanding result generated by each agent member 122 is a field in the record, such as Figure 3 As shown. For each document ID (ID), the summary extraction understanding result saves the output of the Agent team member for summary extraction on the understanding of this document; the paragraph understanding result is the output of the Agent team member for paragraph understanding on the understanding of this document; the table understanding result is the output of the Agent team member for table content analysis on the understanding of this document; the image understanding result is the output of the Agent team member for image content understanding on the understanding of this document; the formula understanding result is the output of the Agent team member for formula content analysis on the understanding of this document; the knowledge graph understanding result is the output of the Agent team member for knowledge graph extraction on the understanding of this document; the knowledge integration result is the output of the knowledge integration agent 123 on the understanding of this document obtained by integrating the understanding results of each special document.

[0061] In one example, the document active enhancement understanding module 12 in the active enhancement document understanding system can read the unprocessed md files in the database during the idle time of the online document understanding module 11, so as to make full use of the system's computing resources to process the unprocessed md files offline and avoid occupying the resources required by the online document understanding module 11. For example, during the day, the system realizes real-time online understanding of documents through the online document understanding module 11; at night, the system actively performs in-depth understanding of documents through the document active enhancement understanding module 12, and stores the results of the document in-depth understanding in the database, so that when users view the uploaded documents again during the day, they can see more comprehensive and in-depth results of the document in-depth understanding. Because the document active enhancement understanding module 12 can be processed offline and the timeliness requirement is not high, a more comprehensive and in-depth understanding of the document can be achieved when there is enough time and computing resources. Figure 4A schematic diagram of the workflow of a specific document active enhancement understanding module 12 provided in an embodiment of the present application. The engine of the document active enhancement understanding module 12 reads the list of unprocessed md files in the database at night, and then sends the document IDs to the agent leader 121 in sequence according to the document IDs in the list of unprocessed md files. The agent leader 121 obtains the corresponding unprocessed md file from the database according to the document ID, and performs a preliminary parsing task on the md file through the agent leader 121 to extract the basic structure and content information of the md file to form a document object. The agent leader 121 then distributes the document object to each intelligent agent member 122. Among them, each intelligent agent member 122 includes a summary extraction agent, a paragraph relationship agent, a table understanding agent, a picture understanding agent, a formula understanding agent, a graph construction agent, and a knowledge integration agent. For any agent member 122, the agent member 122 can perform its own specific special understanding task on the document object based on the pre-trained deep learning large model (such as a large language model or a multimodal large model, etc.) for processing the special understanding task, so as to obtain the special document understanding result and save it in the database for subsequent integration. Under the arrangement of the agent leader 121, the knowledge integration agent 123 can obtain the special document understanding results generated by multiple agent members 122 from the database, combine the special document understanding results with the document object provided by the agent leader 121, and integrate them into the document deep understanding result.

[0062] 2.2 Internal Mechanism of Agent

[0063] In this application, any agent includes a planning unit, an execution unit, and a memory unit;

[0064] The memory unit is used to store and complete short-term information and long-term information related to the understanding task corresponding to the agent; wherein the short-term information is used to store multiple rounds of interaction history, and the long-term information is used to store key knowledge required to complete the document understanding task;

[0065] The planning unit is used to decompose the understanding task into a plurality of subtasks through the memory unit;

[0066] The execution unit is used to construct prompt information corresponding to the subtask through the memory unit for the multiple subtasks decomposed by the planning unit; and execute the subtask based on the prompt information through a pre-trained deep learning large model for processing the subtask.

[0067] Figure 5A schematic diagram of the structure of an agent provided in an embodiment of the present application. In the present application, any agent includes a planning unit, an execution unit and a memory unit, and these units work together to complete the corresponding understanding task. Among them, the planning unit is responsible for analyzing the task requirements and generating an execution plan. Exemplarily, the planning unit decomposes the understanding task corresponding to the agent by reading the information in the memory unit, thereby decomposing a complex task into multiple subtasks. For example, in the formula content parsing task, the planning unit decomposes the formula into subtasks such as symbol definition, relationship between symbols and calculation steps according to the rules in the memory unit. The execution unit is responsible for executing each subtask decomposed by the planning unit, constructing corresponding prompt information for each subtask through the information stored in the memory unit, and then executing the corresponding subtask based on the prompt information corresponding to each subtask by calling the pre-trained deep learning large model. For example, in the table content parsing, the execution unit can extract the title, row and column data of the table and its corresponding semantic relationship based on the prompt information. The memory unit is responsible for storing and completing short-term information and long-term information related to the understanding task corresponding to the agent to support the learning and adaptation process of the agent. The short-term information includes the history records of multiple rounds of interactions, such as the input and output data of the current task. The long-term information is used to store the background knowledge required to complete the task, such as professional terms in the medical field, diagnosis rules, etc.

[0068] Among them, the deep learning big models used by the intelligent agent mainly include the big language model (LLM) and the large multimodal model (VLLM). The big language model is used to process text tasks, such as paragraph understanding and summary extraction; the multimodal big model processes cross-modal tasks, such as image content understanding and image and text matching analysis.

[0069] 3. Result merging and display module 13.

[0070] The result merging and displaying module 13 is used to display the document understanding content after processing. The document understanding content includes the following parts:

[0071] A. Lightweight markup language md file generated by the online document understanding module 11.

[0072] For the md file, the result merging and displaying module 13 can return the md file to the user as soon as it is obtained.

[0073] B. The document deep understanding result generated by the document active enhancement understanding module 12.

[0074] For the document deep understanding result, the result merging and displaying module 13 can output the document deep understanding result when receiving a user's query for the document again.

[0075] Among them, the document understanding content display forms include text, charts, flow charts, knowledge graphs and other visual methods, which intuitively present the important information in the document and the relationship between each element, and are not specifically limited here.

[0076] In one example, the result merging and displaying module 13 may receive a query operation input by a user. For example, a user may query the document understanding content of a certain document through a query button displayed on the display interface, or may input a requirement for querying the document understanding content of a certain document through a text box displayed on the display interface. Based on the query operation input by the user, the result merging and displaying module 13 may query the corresponding document understanding content in the database and display the document understanding content to the user.

[0077] The beneficial effects of this application are as follows:

[0078] 1. The online document understanding module performs preliminary analysis on the PDF files uploaded by users and quickly generates lightweight md files. This lightweight processing reduces the complexity of subsequent tasks and improves processing efficiency. At the same time, the document active enhancement understanding module further improves the understanding and processing quality of document content by deeply analyzing the md files output by the online document understanding module.

[0079] 2. Since the document active enhanced understanding module decomposes the document deep understanding task into three stages: preliminary analysis, special understanding task processing and result integration. The agent leader of the document active enhanced understanding module is responsible for the preliminary analysis and task allocation of the document, and multiple agent members of the document active enhanced understanding module complete special understanding tasks respectively. The knowledge integration agent of the document active enhanced understanding module is used to summarize and integrate special document understanding results. Among them, multiple agent members independently perform their respective special understanding tasks, which significantly improves the task processing speed of the system. And through the special understanding task processing of the agent members, the system can perform more accurate and in-depth analysis of different content elements in the document content. Through the task scheduling of the agent leader and the result integration of the knowledge integration agent, efficient document understanding and information processing are achieved, and deep fusion and semantic association of information are achieved, which is suitable for the processing and display of complex information.

[0080] Embodiment 2:

[0081] Based on the same inventive concept, the present application also provides a method for actively enhancing document understanding. Figure 6 A schematic diagram of a process of actively enhancing document understanding provided in an embodiment of the present application, the process comprising:

[0082] S1: receiving a PDF file uploaded by a user through an online document understanding module, and performing online document understanding on the PDF file to obtain a lightweight markup language md file of the PDF file.

[0083] S2: The md file is transferred to the document active enhancement understanding module, and the following steps are performed in sequence:

[0084] S21: The intelligent agent group leader in the active enhanced understanding module receives the input of the md file and performs a preliminary parsing task on the md file to parse and obtain the document object corresponding to the md file;

[0085] S22: Distribute the document object to multiple agent members in the active enhanced understanding module, and the agent members respectively perform their own special understanding tasks on the received document objects to generate special document understanding results;

[0086] S23: The knowledge integration agent in the active enhanced understanding module receives the input of the special document understanding results of the unprocessed md file, and performs a knowledge integration task on each of the special document understanding results to obtain a document in-depth understanding result.

[0087] S3: The md file generated by the online document understanding module and the document deep understanding result generated by the document active enhancement understanding module are displayed through the result merging and displaying module.

[0088] The active enhanced document understanding method provided in the present application is applied to a computer device, which may be an intelligent device, such as a mobile terminal, a computer, etc., or a server, such as a business server, an application server, etc.

[0089] Since the principle of solving the problem by the above-mentioned active enhanced document understanding method is similar to that of the active enhanced document understanding system, the details may be referred to the embodiment of the above-mentioned system, and the repeated parts will not be repeated here.

[0090] Embodiment 3:

[0091] See also Figure 7 , Figure 7 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present application, such as Figure 7As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.

[0092] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0093] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0094] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0095] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0096] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.

[0097] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0098] Embodiment 4:

[0099] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor implements the following steps when executing:

[0100] S1: receiving a PDF file uploaded by a user through an online document understanding module, and performing online document understanding on the PDF file to obtain a lightweight markup language md file of the PDF file;

[0101] S2: The md file is transferred to the document active enhancement understanding module, and the following steps are performed in sequence:

[0102] S21: The intelligent agent group leader in the active enhanced understanding module receives the input of the md file and performs a preliminary parsing task on the md file to parse and obtain the document object corresponding to the md file;

[0103] S22: Distribute the document object to multiple agent members in the active enhanced understanding module, and the agent members respectively perform their own special understanding tasks on the received document objects to generate special document understanding results;

[0104] S23: the knowledge integration agent in the active enhanced understanding module receives the input of the understanding results of each of the special documents of the unprocessed md file, and performs a knowledge integration task on each of the special document understanding results to obtain a document in-depth understanding result;

[0105] S3: The md file generated by the online document understanding module and the document deep understanding result generated by the document active enhancement understanding module are displayed through the result merging and displaying module.

[0106] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to that of the active enhanced document understanding method, the implementation of the above-mentioned computer-readable storage medium can refer to the embodiment of the method, and the repeated parts will not be repeated.

[0107] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. An active enhanced document understanding system, characterized in that: The system comprises: An online document understanding module, used for receiving a PDF file uploaded by a user, and performing online document understanding on the PDF file to obtain a lightweight markup language md file of the PDF file; The document active enhancement understanding module is connected to the online document understanding module, and is used to receive unprocessed md files and perform document deep understanding tasks on them; wherein, the document active enhancement understanding module includes: an agent group leader, multiple agent group members and a knowledge integration agent; The agent group leader is connected to the plurality of agent group members and the knowledge integration agent respectively, and is used to receive an unprocessed md file input, and perform a preliminary parsing task on the unprocessed md file to parse and obtain a document object corresponding to the unprocessed md file; The plurality of intelligent agent members are used to receive the input of the document object and perform their own corresponding special understanding tasks on the document object to obtain special document understanding results; The knowledge integration agent is used to receive the input of the understanding results of each of the special documents of the unprocessed md file, and perform the knowledge integration task on each of the special document understanding results to obtain the document in-depth understanding results; The result merging and displaying module is used to display the md file generated by the online document understanding module and the document deep understanding results generated by the document active enhanced understanding module.

2. The system according to claim 1, characterized in that The online document understanding module is also used to save the md file as an unprocessed md file in a database, so that the document active enhancement understanding module processes the unprocessed md file saved in the database during the idle time period of the online document understanding module.

3. The system according to claim 2, characterized in that The agent group leader is specifically used to: According to the document identifier recorded in the unprocessed md file list stored in the database, the unprocessed md file is obtained from the database.

4. The system according to claim 1, characterized in that Any intelligent agent includes a planning unit, an execution unit, and a memory unit; The memory unit is used to store and complete short-term information and long-term information related to the understanding task corresponding to the agent; wherein the short-term information is used to store multiple rounds of interaction history, and the long-term information is used to store key knowledge required to complete the document understanding task; The planning unit is used to decompose the understanding task into a plurality of subtasks through the memory unit; The execution unit is used to construct prompt information corresponding to the subtask through the memory unit for the multiple subtasks decomposed by the planning unit; and execute the subtask based on the prompt information through a pre-trained deep learning large model for processing the subtask.

5. The system according to claim 4, characterized in that The deep learning large model includes but is not limited to a large language model and a multimodal large model.

6. The system according to claim 1, characterized in that The multiple intelligent agent team members are respectively responsible for performing the following special understanding tasks: summary extraction, paragraph comprehension, table content analysis, image content understanding, formula content analysis, knowledge graph extraction, and relationship analysis between content elements.

7. The system according to claim 6, characterized in that The document in-depth understanding result is presented in a visual form, describing the attributes of each content element in the unprocessed md file and the relationship between them.

8. A method for actively enhancing document understanding, characterized in that: The method comprises: S1: receiving a PDF file uploaded by a user through an online document understanding module, and performing online document understanding on the PDF file to obtain a lightweight markup language md file of the PDF file; S2: The md file is transferred to the document active enhancement understanding module, and the following steps are performed in sequence: S21: The intelligent agent group leader in the active enhanced understanding module receives the input of the md file and performs a preliminary parsing task on the md file to parse and obtain the document object corresponding to the md file; S22: Distribute the document object to multiple agent members in the active enhanced understanding module, and the agent members respectively perform their own special understanding tasks on the received document objects to generate special document understanding results; S23: the knowledge integration agent in the active enhanced understanding module receives the input of the understanding results of each of the special documents of the unprocessed md file, and performs a knowledge integration task on each of the special document understanding results to obtain a document in-depth understanding result; S3: The md file generated by the online document understanding module and the document deep understanding result generated by the document active enhancement understanding module are displayed through the result merging and displaying module.

9. A computer device, characterized in that: The computer device includes a processor, and the processor is used to implement the steps of the actively enhanced document understanding method as described in any one of claims 1 to 7 when executing a computer program stored in a memory.

10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the steps of the actively enhanced document understanding method as described in any one of claims 1 to 7.