Method and system for document understanding

WO2026177359A1PCT designated stage Publication Date: 2026-08-27LG MANAGEMENT DEV INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/000287
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-12-16
Filing Date
2026-01-06
Publication Date
2026-08-27

Smart Images

  • Figure KR2026000287_27082026_PF_FP_ABST
    Figure KR2026000287_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for document understanding, and the method for document understanding according to the present invention may comprise the steps of: receiving a user request including at least one document; identifying, through analysis of the document, at least one layout region each including at least one content having different data characteristics; specifying one or more models configured to perform processing on each of the at least one content, on the basis of the data characteristics of the content included in each of the at least one layout region; extracting the at least one content included in each of the at least one layout region by using each of the one or more models; converting the document into a form corresponding to the user request by using the at least one content extracted using the one or more models; and as a response to the user request, providing a user terminal with the document converted into the form corresponding to the user request.
Need to check novelty before this filing date? Find Prior Art

Description

Document Understanding Methods and Systems

[0001] The present invention relates to a document understanding method and system, and provides a document understanding method and system using Deep Document Understanding (DDU) technology.

[0002] Documents are recorded through characters or symbols (or codes) to preserve and transmit thoughts, ideas, intentions, and information; in a broad sense, they can include everything that contains meaning, such as pictures, photographs, and videos. This means that information is recorded and preserved in various forms and can serve as a means of communication.

[0003] In this regard, a large volume of electronic documents created via computers is being utilized in various fields in modern society. These documents contain not only text but also various forms of information, such as tables, images, and graphs. For example, documents used in industrial settings, such as papers, patents, technical reports, and manuals, include not only text but also diverse forms of information like tables, graphs, diagrams, images, mathematical formulas, and molecular structural formulas.

[0004] While humans can read and understand the various forms of information contained in such documents, computers cannot fully comprehend complex information in the same way; therefore, analyzing and utilizing the document content requires the cumbersome task of converting it into a format that computers can recognize. However, as documents contain a mixture of diverse information, manually reviewing or analyzing them can consume a significant amount of time and cost.

[0005] To address this, various technologies for understanding and analyzing documents have been proposed in the past. However, conventional technologies are often optimized for specific domains or types of information (e.g., text-centric, table-centric, or graph-centric), and thus have limitations in that they are difficult to apply to documents containing information from different domains or complex forms.

[0006] Consequently, there is still a need for methods to deeply understand the various forms of information contained in documents and provide solutions optimized for various fields based on this understanding.

[0007] The present invention is intended to provide a document understanding method and system capable of effectively understanding various forms of documents.

[0008] More specifically, the present invention is intended to provide a document understanding method and system capable of effectively understanding complex forms of information contained in various types of documents.

[0009] In particular, the present invention is intended to provide a document understanding method and system that automatically converts multimodal information from various types of documents into data.

[0010] Furthermore, the present invention is intended to provide a document understanding method and system that can be universally utilized for various types of documents.

[0011] Furthermore, the present invention aims to provide a document understanding method and system capable of improving efficiency in various industrial or research fields and providing optimized solutions.

[0012] To solve the problem described above, the document understanding method according to the present invention, which is performed by a computer, is,

[0013] The method may include the steps of: receiving a user request including at least one document; identifying at least one layout area each including at least one content having different data characteristics through analysis of the document; specifying one or more models configured to perform processing on each of the at least one content based on the data characteristics of the content included in each of the at least one layout area; extracting the at least one content included in each of the at least one layout area using each of the at least one models; converting the document into a form corresponding to the user request using the at least one content extracted using the one or more models; and providing the document converted into a form corresponding to the user request to a user terminal as a response to the user request.

[0014] In an embodiment, the user request includes a request to convert the form of at least one document to be converted into at least one specific form among a plurality of forms, and in the conversion step, the form of at least one document can be converted into the specific form corresponding to the user request using the extracted at least one content.

[0015] In an embodiment, the step of identifying the plurality of layout areas involves processing the at least one document as input to a document layout analysis (DLA) model, and the document layout analysis model can analyze the at least one document to identify the plurality of layout areas, each containing at least one content.

[0016] In an embodiment, in the step of extracting each of the at least one content, each of the at least one content included in each of the plurality of layout areas identified by the layout analysis model can be extracted in each of the plurality of models.

[0017] In an embodiment, the document layout analysis model may be configured to output at least one of a label corresponding to each of the at least one content included in the at least one document, location information for each of the at least one content, and a score for each of the at least one content.

[0018] In the embodiment, in the conversion step, the relationship between the flow of the at least one document and the extracted at least one content can be understood to convert the form of the at least one document into the specific form.

[0019] In an embodiment, the specific form may include at least one form different from the form of at least one document, satisfying the user request.

[0020] In an embodiment, the method further includes a step of specifying sequence information for at least one content extracted from at least one document, and when the sequence information is specified, the conversion step may convert the form of the at least one document into the specified form based on the specified sequence information.

[0021] In an embodiment, the conversion step may convert the form of the at least one document into the specific form based on at least one of the sequence information and the output of the document layout analysis (DLA) model.

[0022] In an embodiment, the at least one content includes at least one of text, a table, a chart, a molecular structure, and chemical reaction information, and the user request may further include a request to reflect style information in the at least one document to be converted.

[0023] In an embodiment, if the style information is further included in the user request, the conversion step may convert the form of the at least one document into the specific form by reflecting the style information.

[0024] In an embodiment, the style information may be reflected in at least a part of the at least one content based on the user request.

[0025] In an embodiment, when a request to summarize the at least one document is received as a user request, the method may further include the step of summarizing the at least one document using the at least one content extracted from each of the plurality of models in response to the user request.

[0026] In the embodiment, in the extraction step, raw data can be extracted from at least some of the content of at least one of the content by using at least some of the plurality of models.

[0027] In an embodiment, the at least portion of the content may include at least one of the table and the chart.

[0028] In an embodiment, the method further includes the step of checking the extension of at least one document and converting the at least one document into a pre-set specific format. When the at least one document is converted into the specific format, the method identifies the plurality of layout areas each containing the at least one content through analysis of the at least one document converted into the specific format, extracts the at least one content included in each of the plurality of layout areas using each of the plurality of models, converts the at least one document converted into the specific format into a form corresponding to the user request using the at least one content extracted from each of the plurality of models, and provides the document converted into a form corresponding to the user request to the user terminal.

[0029] A document understanding system according to the present invention, comprising a memory configured to store executable instructions and one or more processors configured to perform operations by executing one or more instructions, wherein the system receives a user request including at least one document and, through analysis of the document, identifies at least one layout area each containing at least one content having different data characteristics, specifies one or more models configured to perform processing on each of the at least one content based on the data characteristics of the content included in each of the at least one layout area, extracts the at least one content included in each of the at least one layout area using each of the at least one models, converts the document into a form corresponding to the user request using the at least one content extracted using the one or more models, and provides the document converted into a form corresponding to the user request to a user terminal as a response to the user request.A program according to the present invention is a program that is executed by one or more processes in an electronic device and can be stored on a computer-readable recording medium, and may include instructions for performing the following steps: receiving a user request including at least one document; identifying at least one layout area each including at least one content having different data characteristics through analysis of the document; specifying one or more models configured to perform processing on each of the at least one content based on the data characteristics of the content included in each of the at least one layout area; extracting the at least one content included in each of the at least one layout area using each of the at least one models; converting the document into a form corresponding to the user request using the at least one content extracted using the one or more models; and providing the document converted into a form corresponding to the user request to a user terminal as a response to the user request.

[0030] As described above, the document understanding method and system according to the present invention can increase (or enhance) user convenience by recognizing the structure of various types of documents and the relationships of at least one content (e.g., text, table, graph, chart, molecular structural formula, etc.) included in the document, thereby automatically extracting accurate information. Through this, the user can minimize the time required to find necessary information in a document and reduce the hassle.

[0031] Furthermore, according to the document understanding method and system of the present invention, at least one piece of content included in a document is identified, and at least one piece of content is automatically extracted from the document by utilizing a plurality of models specialized for each of the at least one piece of content having different data characteristics. That is, the present invention utilizes a plurality of models to quickly and accurately extract necessary information even from a document containing various forms of content (or various forms of information). Through this, the user can receive the necessary information quickly and accurately.

[0032] Furthermore, according to the document understanding method and system of the present invention, by deeply analyzing the visual information of various forms of content included in a document, information omission can be prevented and the sophistication of the analysis can be enhanced. Through this, the present invention can accurately process a vast amount of documents, thereby improving the user's overall work or business processing speed.

[0033] Furthermore, according to the document understanding method and system of the present invention, by deeply analyzing various forms of content included in a document and providing a solution optimized for the user based on this, the efficiency of document processing can be maximized and the user's overall work efficiency can be improved.

[0034] Furthermore, according to the document understanding method and system of the present invention, it is possible to understand text, graphs, and tables contained in a document, while simultaneously interpreting complex molecular structural formulas and chemical reaction information. Through this, the present invention can enhance the research efficiency of chemists and support innovative discoveries in the field of chemistry. In other words, the present invention can satisfy the needs of the research field through a customized model that understands even information related to the chemical domain within the document.

[0035] Furthermore, according to the document understanding method and system of the present invention, molecular structural formulas can be accurately recognized in documents using a model specialized for molecular structural formulas, and a large-scale database of molecular structural formulas can be constructed based on this. Through this, researchers can utilize the database to search for chemical structural formulas and efficiently carry out large-scale analysis tasks required for research, thereby reducing the time required for research and / or development.

[0036] Furthermore, according to the document understanding method and system of the present invention, at least one piece of content extracted from a document using a plurality of models and output data generated based on said at least one piece of content can be visualized and provided through a user interface. Through this, the user can intuitively recognize the information needed and understand it more quickly, thereby improving the efficiency of work and / or research.

[0037] Furthermore, according to the document understanding method and system of the present invention, various content can be recognized from various types of documents, converted and translated into an editable form, and provided to the user. That is, the present invention can process documents of various types, thereby dramatically expanding the scope of document processing. Additionally, by automating the document conversion or translation process, users do not need to perform complex manual editing or translation, thus improving the efficiency and accuracy of document-based work.

[0038] Furthermore, according to the document understanding method and system of the present invention, raw data of tables, charts, and molecular structural formulas included in a document can be extracted and converted into an editable form to be provided to a user. Through this, the present invention solves the problem that detailed editing is impossible when tables, charts, and molecular structural formulas are processed as images, thereby improving the usability of document processing and minimizing the time or cost required for a user to re-edit the document.

[0039] FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented.

[0040] FIG. 2 illustrates an example of a block diagram of a computing device that may be included in a user computing device, a server computing system, and a training computing system, as an embodiment of a computing system in which the present invention can be implemented.

[0041] Figure 3 illustrates an example of a block diagram from another perspective of a computing device, which is one of the components of a computing system.

[0042] FIGS. 4, FIGS. 5, FIGS. 6a, FIGS. 6b, and FIGS. 7 are conceptual diagrams for explaining a document understanding system according to the present invention.

[0043] FIG. 8 is a flowchart illustrating a method for understanding documents according to the present invention.

[0044] FIGS. 9, FIGS. 10, FIGS. 11, FIGS. 12, FIGS. 13, FIGS. 14, FIGS. 15, FIGS. 16, FIGS. 17, FIGS. 18, FIGS. 19, FIGS. 20, FIGS. 21, FIGS. 22, FIGS. 23, FIGS. 24 and FIGS. 25 are conceptual diagrams for explaining a document understanding method according to the present invention.

[0045] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components are assigned the same reference number regardless of the drawing symbols, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention.

[0046] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.

[0047] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0048] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0049] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0050] Meanwhile, FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented. In this regard, the document understanding system according to the present invention can be implemented through a computing device described below and can perform data processing related to the document understanding method described above.

[0051] Referring to FIG. 1, a computing system (10000) that performs a method of automatically converting at least one or a plurality of contents extracted from various forms of documents according to an embodiment of the present invention into data and providing various services and / or functions based thereon may include at least one computing device. At this time, the at least one computing device may be a single processor or a multi-processor computing device.

[0052] The components of at least one computing device of the present invention may include various hardware components such as one or more processors, memory, other hardware, and a system bus (not shown) that connects various system components so that they can transmit and receive data to and from each other (e.g., telecommutatively connected, physically connected, electrically connected), and the components of at least one computing device are not limited thereto and may be very diverse.

[0053] Meanwhile, at least one computing device included in a computing system (10000) that performs a method of automatically converting at least one or multiple contents extracted from various types of documents into data and providing various services and / or functions based thereon may be connected to communicate via a network (1070). For example, at least one computing device included in the computing system (10000) may be clustered or may be part of a local area network (LAN). Additionally, at least one computing device may be part of a wide area network (WAN) or connected to at least one of a client-server network and a peer-to-peer network within the cloud.

[0054] Meanwhile, when at least one computing device is used in at least one of a network environment and a cloud computing environment, the at least one computing device may be connected to at least one of a public and private network through a network interface or adapter. In one embodiment, other communication connection devices, such as a modem, may be used to establish communication through the network. The modem may be at least one of an internal modem and an external modem, and may be connected to a system bus through a network interface or a specific mechanism, etc. A wireless network component consisting of an interface and an antenna may be coupled to the network through a device such as an access point, a peer computer, etc. In the present invention, the method of connecting at least one computing device to communicate through the network (1070) is not limited, and it may be connected to communicate in a manner different from the described example.

[0055] Furthermore, other computer-type devices and / or systems not shown in FIG. 1 may also interact technically with at least one computing device or other system through one or more connections to the network (1070) via a network interface. Here, the network interface may include network interface equipment such as a physical network interface controller (NIC) or a virtual network interface (VIF).

[0056] The network (1070) of the present invention may include various forms such as the Internet, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, Wireless USB (Wireless Universal Serial Bus), etc., and in the present invention, data transmission may be performed based on standard communication protocols such as TCP / IP, HTTP, SSL, etc.

[0057] A computing system (10000) that performs a method of automatically converting at least one or a plurality of contents extracted from various forms of documents according to the present invention into data and providing various services and / or functions based thereon may include at least one of a user computing device (1010), a training computing system (1050), and a server computing system (1030).

[0058] A user computing device (1010) according to the present invention may be understood as a computing device comprising at least one and / or at least one processor (1011) and a memory (1012) that perform a method of automatically converting at least one or a plurality of contents extracted from various types of documents into data and providing various services and / or functions based thereon. For example, the user computing device (1010) may include at least one computing device among a smartphone, a smart TV, a laptop computer, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, and a head-mounted display).

[0059] At least one and / or at least one processor (1011) constituting the user computing device (1010) may include one or more general-purpose processors and / or one or more special-purpose processors. For example, at least one and / or at least one processor (1011) constituting the user computing device (1010) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), an application integrated circuit, an application semiconductor (ASIC), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.

[0060] Furthermore, at least one and / or at least one processor (1011) may be configured to execute computer-readable instructions contained in memory (1012) and / or other instructions described herein.

[0061] The memory (1012) constituting the user computing device (1010) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media and / or other types of physically durable storage media.

[0062] For example, the memory (1012) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and may include web storage of a server that performs the storage function of memory on the internet. This memory (1012) may store data and instructions necessary for the operation of an application to automatically convert at least one or more contents extracted from various types of documents into data and to provide various services and / or functions based thereon.

[0063] A user computing device (1010) may include one or more user input components (1021) that detect user input. For example, the user input component (1021) may also be referred to as a user interface module. The user input component (1021) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of user input component (1021). In this case, the user input component (1021) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user. Meanwhile, the user of the present invention may refer to an automated agent, script, playback software, etc., that operates on behalf of one or more people.

[0064] A user can interact with a computing system (10000) including at least one computing device through input text, touch, voice, movement, computer vision, gestures and / or other forms of input / output using a user input component (1021). For example, the user input component (1021) may include one or more of a command line interface (CLI), a graphical user interface (GUI), a natural user interface (NUI), a voice command interface and / or other user interface (UI) representations.

[0065] Between the user input component (1021) and the user computing device (1010), one or more application programming interface (API) calls may be made based on user input received from the user interface and / or network.

[0066] Here, the expression "based on" may be interpreted to include cases where it is based on the use of a specific configuration, modified from, derived from, influenced by, dependent on, or otherwise derived from a specific configuration. In some embodiments, an API call may be configured for a specific API, which may be interpreted or converted into an API call configured for another API. Here, an API may refer to a defined interface or connection between computers or between computer programs.

[0067] In one embodiment, the user computing device (1010) may store at least one machine learning model (1020). For example, the user computing device (1010) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks), or other types of machine learning models including non-linear models and / or linear models, which perform a method of automatically converting at least one or a plurality of contents extracted from various types of documents into data and providing various services and / or functions based thereon, and may be composed of a combination thereof.

[0068] According to an embodiment of the present invention, a user computing device (1010) may use a local or / and external machine learning model (1020) to automatically convert at least one or more contents extracted from various types of documents into data, and provide various services and / or functions based thereon. Alternatively, the user computing device (1010) may use a machine learning model (1040) provided by a server to automatically convert at least one or more contents extracted from various types of documents into data, and provide various services and / or functions based thereon.

[0069] Additionally, according to another embodiment of the present invention, a server computing system (1030) communicating with a user computing device (1010) may provide various services and / or functions based on at least one content or a plurality of content extracted from various types of documents to the user computing device (1010) on an application or / and the web, in accordance with a request from a user received through the user computing device (1010).

[0070] In addition, according to another embodiment of the present invention, at least a part of a user computing device (1010) and a server computing system (1030) are interconnected to perform a process of automatically converting at least one or a plurality of contents extracted from various types of documents into data, and based on this, various services and / or functions can be provided to the user.

[0071] Additionally, according to various embodiments of the present invention, a user computing device (1010) and / or a server computing system (1030) can learn a machine learning model (1020, 1040) performed in a method of automatically converting at least one or a plurality of contents extracted from various forms of documents into data and providing various services and / or functions based thereon through interaction with a training computing system (1050) that is communicatedly connected via a network (1070). In this case, the training computing system (1050) may be a computing system separate from the server computing system (1030). Alternatively, in some embodiments, the training computing system (1050) may be a part of the server computing system (1030) or a part of the user computing device (1010).

[0072] Meanwhile, the server computing system (1030) may include at least one processor (1031) and memory (1032). Here, the processor (1031) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an application integrated circuit, an application semiconductor (ASIC), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions. For example, at least one processor (1031) may include a circuit and a transistor configured to execute instructions from memory (1032).

[0073] The memory (1032) constituting the server computing system (1030) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media, and / or other types of physically durable storage media. For example, the memory (1032) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs the storage function of memory over the internet. Additionally, the server computing system (1030) may further include a data storage (data store). For example, the data storage may be composed of at least one of a relational database, a NoSQL database, a data warehouse, and a local file system.

[0074] In the memory (1032) constituting the server computing system (1030) according to the present invention, data and instructions necessary for the operation of an application to automatically convert at least one or a plurality of contents extracted from various types of documents by the at least one processor (1031) into data and to provide various services and / or functions based thereon may be stored.

[0075] In one embodiment, the server computing system (1030) may be composed of a single device or a plurality of computing devices, and these may be configured to operate according to a sequential or parallel computing architecture. Additionally, a distributed processing system may be configured with a plurality of networked devices.

[0076] Meanwhile, the training computing system (1050) may include at least one processor (1051) and memory (1052). The model trainer (1060) is a logical component that executes the training of at least one machine learning model (1020, 1040) and may be implemented in the form of hardware, firmware, or software. For example, the model trainer (1060) may be executed by the processor (1051) after loading training data (1061) stored in a storage device into memory (1052). For example, the model trainer (1060) may be configured to execute one or more operations (e.g., model training, model reconstruction, model validation, model testing) on ​​at least one machine learning model.

[0077] The machine learning model of the present invention may include at least one of a statistical model, an algorithm, a neural network (NN), a convolutional neural network (CNN), a generative neural network (GNN), a Word2Vec model, a Bag of Words model, a TF-IDF (document frequency-inverse document frequency) model, a GPT (Generative Pre-trained Transformer) model (or other autoregressive models), a PPO (Proximal Policy Optimization) model, a nearest neighbor model (e.g., a k-nearest neighbor model), a linear regression model, a K-means clustering model, a Q-learning model, a TD (Temporal Difference) model, a Deep Adversarial Network model, and all other types of models further described herein.

[0078] Specifically, the model trainer (1060) may execute operations to train a machine learning model, and said operations may include at least one of adding, removing, and modifying model parameters. At this time, the training of the machine learning model may be at least one of supervised learning, semi-supervised learning, and unsupervised learning. In one embodiment, the training of the machine learning model may include the step of repeatedly inputting training data (1061) based on epochs and repeatedly performing the machine learning model training process configured in this way. Here, an epoch may refer to a unit in which the entire set of training data (1061) undergoes forward and backpropagation processing once. In some implementations, different levels of training methods (e.g., supervised learning, semi-supervised learning, unsupervised learning) may be used for different epochs.

[0079] The training data (1061) of the present invention may include input data and / or data previously output from at least one machine learning model (e.g., recursive learning feedback).

[0080] At least one parameter of a machine learning model may include at least one of a seed value, a model node, a model layer, an algorithm, a function, connections between different machine learning models, connections between parameters, machine learning model constraints, and other digital components that influence the output of the machine learning model. In this case, model connections between different machine learning models may include or represent relationships between model parameters and / or models, which may be dependent or interdependent, hierarchical, and / or static or dynamic. The combinations and configurations of model parameters described herein may be too complex to be maintained or utilized by human cognitive abilities.

[0081] In the present invention, the machine learning parameters described according to the embodiments are not limited, and a single machine learning model may further include a plurality of model parameters.

[0082] Meanwhile, FIG. 2 illustrates an example of a block diagram of a computing device (1100) that may be included in a user computing device (1010), a server computing system (1030), and a training computing system (1050), as an embodiment of a computing system (10000) in which the present invention can be implemented.

[0083] As illustrated in FIG. 2, the computing device (1100) may include at least one application (e.g., Application 1 to Application N), and each of the at least one application may include a machine learning library and a model execution environment for performing a method of automatically converting at least one or a plurality of contents extracted from various forms of machine learning-based documents into data and providing various services and / or functions based thereon. The at least one application included in the computing device (1100) may communicate with the sensor, context manager, device state manager, or additional component(s) within the computing device (1100) via an API (Application Programming Interface). In one embodiment, the at least one application may perform an interface with device components, such as receiving sensor data or state data or transmitting prediction results to an output device via a public or private API.

[0084] Meanwhile, FIG. 3 illustrates an example of a block diagram in another aspect of a computing device (1200) which is one of the components of a computing system (10000) that performs a method of automatically converting at least one or a plurality of contents extracted from various forms of documents according to an embodiment of the present invention into data and providing various services and / or functions based thereon.

[0085] A computing device (1200) according to the present invention may include at least one application (e.g., Application 1 to Application N), and at least one application may communicate with a central intelligence layer (1210). Each application may interact with a shared model within the central intelligence layer (1210) through an API (e.g., a common API).

[0086] The central intelligence layer (1210) includes one or more machine learning models and may share them among multiple applications or provide them independently to each. In one embodiment, the central intelligence layer (1210) may be integrated as part of an operating system or implemented as a separate logical layer.

[0087] Additionally, the central intelligence layer (1210) can communicate with the central device data layer (1220). The central device data layer (1220) can integrate and store various forms of documents containing at least one or multiple contents stored within the computing device (1200), automatically convert at least one or multiple contents from the documents into data, and provide this as input data necessary to provide various services and / or functions based thereon. Each device component (e.g., sensor, state manager, etc.) can communicate with the central device data layer (1220) through a private API, etc.

[0088] The technology described herein may be composed of a single or multiple computing devices, and a machine learning model that performs a method of automatically converting at least one or multiple contents extracted from various forms of documents into data and providing various services and / or functions based thereon may be executed sequentially or in parallel on a single component or multiple distributed components. The data storage, machine learning model, and application may be distributed and operated locally or over a network, and these configurations can be flexibly applied to various system architectures.

[0089] Meanwhile, the present invention relates to a document understanding method and system capable of effectively understanding various types of documents. More specifically, the document understanding system according to the present invention may be a system that provides a service and / or function for automatically converting multimodal information (e.g., layout, text, table, chart, image, graph, chemical molecular structural formula, mathematical formula, etc.) from various types of documents into data using Deep Document Understanding (DDU) technology.

[0090] The “document (or document to be analyzed)” described in the present invention may include information or content and may include any type of medium that can be stored, transmitted, processed, or referenced in electronic form. Such a document may include structured data or unstructured data and may consist of text, images, videos, charts, figures, tables (or tables), graphs, codes, mathematical expressions (or formulas), molecular structural formulas, chemical reaction formulas, or a combination thereof.

[0091] In one embodiment, the document may be data in the form of a file that is directly uploaded by the user. In this case, the document may include various electronic document formats such as a word processor document, a presentation file, a spreadsheet file, a text file, an image file, an audio file, a video file, and a PDF file.

[0092] In another embodiment, the document may include content accessible via a network. In this case, the document may include a web page, a Uniform Resource Locator (URL) of a web page, content connected by a hyperlink, a post on the web page, a blog post, an online article, a wiki document, a record in a database, and a screen on a web-based application or related metadata.

[0093] In another embodiment, the document is not limited to a single file or a single web page, but may be a set of documents consisting of multiple pages, sections or resources, and may include part or all of a database, content stored in cloud storage, or content provided via streaming.

[0094] In another embodiment, the document may consist of a single file or multiple files, and may be divided into one or more pages, sections, objects, elements, or data blocks for management or processing. Additionally, the document may include not only static content but also dynamic content whose content changes according to user input, time, location, or external events.

[0095] However, in the present invention, the term “document” is not limited to the examples described above and may be understood to comprehensively refer to all forms of content and data representations that can be used for the storage, transmission, display, or processing of information.

[0096] Furthermore, in the present invention, the document may be used in its original form, or it may be stored and utilized in a form with preprocessing, summarization, segmentation, transformation, indexing, or added metadata. Additionally, the document may be interpreted as a concept that includes not only static content but also content that is updated or dynamically generated over time.

[0097] Hereinafter, the document understanding system according to the present invention will be examined in more detail together with the attached drawings. FIGS. 4, FIGS. 5, FIGS. 6a, FIGS. 6b, and FIGS. 7 are conceptual diagrams for explaining the document understanding system according to the present invention.

[0098] Meanwhile, as illustrated in FIG. 4, the document understanding system (1000) according to the present invention may include at least one of an input unit (100), an output unit (200), a communication unit (300), a storage unit (400), a document understanding unit (500), and a control unit (600). However, the components of the document understanding system (1000) according to the present invention are not limited thereto and may further include various hardware components that perform the same or similar roles as described in the description of the present specification.

[0099] Although not illustrated, the document understanding system (1000) according to the present invention may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processor, tensor processing unit (TPU), graphics processing unit (GPU), neural network processing unit (NPU), application integrated circuit, application semiconductor (ASIC), field programmable gate array (FPGA), quantum processing unit (or quantum processor, QPU), etc.). One or more processors may be configured to execute instructions, computer-readable instructions, and / or other instructions described herein that are stored (or included) in the storage unit (400). The document understanding method and system according to the present invention may perform data processing described below in cooperation with memory and at least one processor. The processor may perform a series of operations and data processing using data and information stored in memory. In this case, memory may be a component of the storage unit (400).

[0100] In addition, the document understanding system (1000) according to the present invention can perform data processing and computation processes using quantum gates, quantum entanglement, and quantum superposition states, taking into consideration implementation in a quantum computer environment. For example, the present invention can perform parallel computations based on qubits, and such quantum computations can operate complementarily with existing classical computers.

[0101] Such quantum computers may include parallel computation using qubits and high-speed data processing devices utilizing quantum entanglement, and hardware-based computational optimization using FPGAs and ASICs is possible. In addition, quantum computers may utilize quantum processors capable of qubit-based parallel computation, and data processing efficiency can be improved through a hybrid structure with existing classical computers.

[0102] Meanwhile, the input unit (100) can be configured in various ways as a means of data input. For example, the input unit (100) can be configured to receive user input. The input unit (100) can be configured to receive user input from a user terminal (10). Here, “receiving input” may mean receiving an input signal (or selection signal) corresponding to the user’s input based on input made by the user through the input unit configuration provided in the user terminal (10).

[0103] Here, the user terminal (10) may include at least one of a mobile phone, a smartphone, a notebook computer, a laptop computer, a slate PC, a tablet PC, an ultrabook, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, and a wearable device (e.g., a smartwatch, a smart glass, a head-mounted display).

[0104] In addition, the input unit (100) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user.

[0105] The input unit (100) may also be referred to as a user interface module. The input unit (100) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of input unit (100).

[0106] Here, user input may include documents, text, images (or videos), voice, etc. In this case, the document understanding system (1000) may further include a module that converts voice into text.

[0107] Next, the output unit (200) can output information through an output unit configuration (e.g., a display unit, a touch screen, a speaker, etc.) provided in a user terminal (10) linked to the document understanding system (1000) according to the present invention. For example, the output unit (200) can output at least one page (2000, or service page) linked to the document understanding system (1000) according to the present invention to the display unit of the user terminal (10). Additionally, the output unit (200) does not necessarily mean a hardware means, but can be understood as a channel for outputting results to a user.

[0108] Next, the communication unit (300) may be connected via a wireless or wired network to a user terminal (10), a server (e.g., a central server, an external server, etc.), a device, and at least one network, etc., to receive or transmit overall data and information necessary for the operation of the document understanding system (1000) according to the present invention.

[0109] The communication unit (300) can support various communication methods depending on the communication standard of the communicating device.

[0110] For example, the communication unit (300) may be configured to communicate with a communication target using at least one of the following technologies: WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ RFID (Radio Frequency Identification), Infrared Communication (Infrared Data Association; IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus).

[0111] Next, the storage unit (400, or memory) serves to store various data related to the present invention and may include one or more non-transient computer-readable storage media that can be read and / or accessed by at least one of one or more processors.

[0112] One or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the storage unit (400) may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage device), whereas in other examples, the storage unit (400) may be implemented using two or more physical devices.

[0113] The storage unit (400) may include computer-readable instructions and additional data. The storage unit (400) may include a storage necessary to perform at least some of the methods, scenarios, and techniques described herein and / or at least some of the functions of the device and network.

[0114] Furthermore, at least a portion of the storage unit (400) may be a cloud storage or a cloud server. At least a portion of the data corresponding to user input received from the input unit (100) and the training data may be stored in the storage unit (400).

[0115] Additionally, the storage unit (400) may store at least one document collected (or received) from various sources (e.g., a web corpus, a document corpus, a database (DB) website, an API, a server linked to the document understanding system (1000), a central server, an external server, cloud storage, a user terminal (10), a large dataset, etc.). For example, the storage unit (400) may store at least one and / or at least one or more (multiple) documents collected from at least one of the various sources (e.g., a user terminal (10)). In this case, the collected documents may include documents related to at least one and / or at least one domain (or field). Alternatively, the collected documents may include documents containing at least one and / or at least one content.

[0116] That is, the storage unit (400) is sufficient as a space where information necessary for the operation of the document understanding system (1000) according to the present invention is stored, and it can be understood that there are no restrictions on the physical space.

[0117] Furthermore, the storage unit (400) may store a computer program including computer program instructions. Furthermore, the storage unit (400) may store a computer program including computer program instructions that control the operation of the system (100) or control the operation of the control unit (600) when loaded into the processor of the system (100).

[0118] Next, the document understanding unit (500) may be configured to effectively understand complex forms of information contained in various types of documents. The document understanding unit (500) may be configured to recognize the structure of the document (20) and to recognize the relationships of the content (e.g., Image-Text, Image-Image, Text-Text) contained in the document (20) to perform the role of extracting information. In the present invention, the document understanding unit (500) may also be named a “deep document understanding model,” a “document understanding model,” or a “DDU model.”

[0119] The document comprehension unit (500) can be configured to extract various forms of content (e.g., layout, text, table, chart, image, graph, chemical molecular structural formula, mathematical formula, etc.) from at least one and / or at least one document (e.g., paper, book, patent document, report, etc.).

[0120] More specifically, the document comprehension unit (500) may be a model trained to understand structured data, unstructured data, linguistic data (or linguistic elements) and non-linguistic data (or non-linguistic elements), etc., included in the document (20), and to extract various content and / or knowledge based on the understood content.

[0121] In one embodiment, the document understanding unit (500) can understand the chemical structure of a molecular structural formula included in the document (20) and, based on the result of understanding, convert the molecular structural formula into a SMILES string expression and extract it. Additionally, the document understanding unit (500) can understand the chemical structure of a molecular structural formula and, based on the result of understanding, perform a graph conversion corresponding to the molecular structural formula.

[0122] In another embodiment, the document understanding unit (500) can understand texts related to molecular structural formulas among the texts included in the document (20) and extract them as text data related to said molecular structural formulas.

[0123] In another embodiment, the document comprehension unit (500) can recognize rows and columns constituting a table associated with molecular structural formulas from the document (20) and convert them into structured data in a format such as HTML or Excel to extract them.

[0124] Additionally, the document comprehension unit (500) can extract relationship information (or relationships) between molecular structures included in the document (20).

[0125] In one embodiment, the document understanding unit (500) can understand the relationship between the first molecular structure and the second molecular structure included in the document (20) and extract relationship information in which a third molecular structure is generated through a chemical reaction between the first molecular structure and the second molecular structure. In this case, the relationship information between the molecular structures can be extracted by understanding the text included in the document (20) to be analyzed, or by understanding the non-verbal data included in the document (20).

[0126] In another embodiment, the document understanding unit (500) can understand the relationship between the first molecular structure and the second molecular structure through a symbol (e.g., plus sign, arrow, etc.) existing in one region among a plurality of regions included in the document (20) where the first molecular structure and the second molecular structure are located, and can extract relationship information in which a third molecular structure is generated through a chemical reaction between the first molecular structure and the second molecular structure.

[0127] Furthermore, the document comprehension unit (500) can extract various forms of content satisfying established content criteria from at least one and / or at least one document. Here, the established content criteria can be set in various ways and can be determined according to the purpose or use of the document comprehension system (1000). For example, if the purpose of use of the document comprehension system (1000) is chemistry, bio, new materials, new substances, and new drug development, the document comprehension unit (500) can be trained to understand and extract content related to chemistry, bio, new materials, new substances, and new drug development from the document (20). In this case, the established content criteria may include content related to molecular structures related to at least one of chemistry, bio, new materials, new substances, and new drug development. Here, the document comprehension unit (500) can extract content related to chemistry, bio, new materials, new substances, and new drug development from the document (20) according to the established content criteria. However, this is merely one embodiment, and the established content standards in the present invention are not necessarily limited thereto.

[0128] As described above, the document comprehension unit (500) can convert various forms of content (or information, data, etc.) contained in the document (20) into data that can be understood by a machine and / or artificial intelligence model. The data extracted using the document comprehension unit (500) can be sorted by page or document and stored in the storage unit (400, or memory). In this case, the document comprehension unit (500) can be configured (or constructed) to include at least one and / or at least one or more (multiple) artificial intelligence models (or models, modules, etc.).

[0129] Here, artificial intelligence models may include various models that can be utilized depending on various situations or purposes. For example, an artificial intelligence model may include at least one of a machine learning (ML) model, a deep learning model, a deep neural network (DNN), a language model (LM), a large language model (LLM), a massive foundation model, a generative AI model, a transformer-based model, a supervised learning (SL) model, a reinforcement learning (RL) model, a vision-language model (VLM), and a special-purpose model (e.g., a time series forecasting model (e.g., ARIMA model, SARIMA model, etc.), a time series foundation model, a graph neural network (GNN), a multimodal model, a natural language processing (NLP) model, a computer vision model, a speech recognition / synthesis model, a recommendation system model, etc.).

[0130] In this regard, as illustrated in FIG. 6a, the document comprehension unit (500) may include at least one of a layout analysis model (e.g., “DLA”, 510), a text extraction model (e.g., “OCR”, 520), a table extraction model (e.g., “Table”, 530), a chart extraction model (e.g., “Chart”, 540), a molecular structure formula extraction model (e.g., “Mol detect”, 550), a chemical reaction information extraction model (e.g., “Reaction”, 560), a molecular structure formula conversion model (e.g., “Chemical Structure Processing”, 570), a content order alignment model (e.g., “ROD”, 580), and a multimodal model (e.g., “VLM”, 590). However, at least some of the multiple models (510, 520, 530, 540, 550, 560, 570, 580, 590) may be implemented as a separate configuration from the document understanding unit (500). Alternatively, at least some of the multiple models (510, 520, 530, 540, 550, 560, 570, 580, 590) may be implemented as part of the document understanding system (1000). The present invention is not limited to any one of these. Furthermore, in the present invention, the term “model” may also be referred to as “module.”

[0131] Referring to FIGS. 5 and 6a, when a document understanding system (1000) receives user input for at least one document, it can identify (or identify) the extension of the received document (e.g., PDF, PPTX, HWPX, PNG, JPG, etc.) and convert the document (or each page of the document) into a specific format set in advance. For example, the specific format set in advance (or format, form, etc.) may include an image.

[0132] And, the document understanding system (1000) can process a document converted into a specific format (i.e., a document converted into an image, 610) as input to a document layout analysis model (510) included in the document understanding unit (500). The document layout analysis model (510) can be configured to perform the role of analyzing the structure of the document and separating the background.

[0133] Specifically, the document layout analysis model (510) may be configured to perform the role of identifying (or detecting) at least one and / or at least one layout area containing at least one and / or at least one content through analysis of the structure (or layout) of the document. Here, the identification of the layout area (or layout identification) may be a process of finding the location of an area (or object) corresponding (or for) each of a plurality of contents having different data characteristics (e.g., text, table (or chart), chart (or graph), molecular structural formula, chemical reaction information, etc.).

[0134] More specifically, the document layout analysis model (510) analyzes the document to include i) text (e.g., Body-text (e.g., a general body paragraph), Caption (e.g., special text located outside of a figure, table, chart, etc., and describing the element), List-item (e.g., an element of a list (or list) hanging, i.e., the paragraph is indented more from the second line onwards than the first line. In this case, each bbox has an entity, and each entity has a category), Title (e.g., the overall title of the document, generally displayed in a large font on the first page), Section-header (e.g., representing all types of text titles excluding the overall document title, in this case, each bbox has an entity, and each entity has a category), Footnote (e.g., small text generally located at the bottom of the page, containing numbers or symbols mentioned in the text above), Formula (e.g., a mathematical equation existing on its own line), Page-header (e.g., a repeating element such as a page number located outside the general text flow at the top Layout areas including at least one of the following can be identified: ii) displayed), page-footer (e.g., repeating elements such as page numbers outside the general text flow), iii) figures (e.g., graphics, photos, etc.), iv) charts (e.g., xy graphs, bar charts, line charts, etc.), iv) tables (e.g., data arranged in a row and column structure), v) other visual content (e.g., flowcharts, diagrams, infographics, maps, organizational charts, etc.).

[0135] At this time, when the document layout analysis model (510) identifies a layout area containing text, if the text is included in the continuous content flow (body flow) of the document, it may identify it as a layout area containing normal text. On the other hand, if the text is a page component that appears repeatedly outside the flow of the document, the document layout analysis model (510) may identify it as a layout area containing format text. For example, normal text is text that constitutes the actual content of the document and may include at least one of body text, caption, list item, title, section header, footnote, and formula (or mathematical expression). In contrast, format text is text for page formatting rather than the content of the document and may include at least one of page header and page footer.

[0136] In the following, we will focus on examples in which multiple contents are extracted from a document. However, as explained above, in the present invention, it is also possible to extract at least one content from a document. Furthermore, in the following, multiple layout areas are described as examples, but it is understood that this may also be at least one layout area.

[0137] As illustrated in FIG. 5, a document layout analysis model (510) can analyze a document converted into a specific format (or a document of a specific format, 610) to identify a plurality of layout areas (611, 612, 613, 614, 615, 617) each containing a plurality of contents having different data characteristics. In this case, the plurality of layout areas (611, 612, 613, 614, 615, 617) may include at least one of a layout area (611) containing a caption (e.g., “Caption”), a layout area (612) containing a page header (e.g., “Page-header”), a layout area (613) containing a table (e.g., “Tabel”), a layout area (614) containing body text (e.g., “Body-text”), a layout area (615) containing a list item (e.g., “List-item”), and a layout area (617) containing a chart (e.g., “Chart”). For example, “idx” included in the document layout analysis results (e.g., “Document Layout Analysis”, 620) may represent a unique number used when referencing the corresponding layout area, and “bbox” may represent location information that each layout area (or object) occupies within the page of the document.

[0138] That is, the document layout analysis model (510) receives an image of each page of a document as input and can output at least one of a label corresponding to each of the multiple contents (e.g., text label, table label, chart label, etc.), location information of each of the multiple contents (e.g., bounding box coordinates in the format [x1, y1, x2, y2]), and a score for each of the multiple contents (e.g., a probability value indicating whether the identified content belongs to the content). The output of this document layout analysis model (510) can also be understood as a “document layout identification result” or a “document layout analysis result.”

[0139] Subsequently, the document comprehension unit (500) can extract multiple contents included in each of the multiple layout areas by utilizing multiple models specialized for each of the multiple contents. To this end, the document comprehension unit (500) can process multiple layout areas containing multiple contents as inputs for each of the multiple models. In the present invention, processing multiple layout areas as inputs for each of the multiple models can also be interpreted to mean calling models to extract multiple contents based on the output of the document layout analysis model (510). Alternatively, it can be interpreted to mean processing at least one of the output of the document layout analysis model (510) and a document of a specific format as inputs for each of the multiple models.

[0140] More specifically, as illustrated in FIGS. 5 and 6a, the document understanding unit (500) can process at least one of a document (610) of a specific format including the output of a document layout analysis model (510) (e.g., a label corresponding to each of a plurality of contents, location information of each of a plurality of contents within a document (or image), etc.) and a plurality of layout areas (611, 612, 613, 614, 615, 617) as the input of each of a plurality of models.

[0141] First, when a document (610) of a specific format is input, the text extraction model (520) can recognize layout areas (611, 614, 615) containing general text and layout areas (612) containing format text from a document (610) of a specific format, based on the output of the document layout analysis model (510) (e.g., a label corresponding to general text and location information of general text within the document (610) of the specific format, a label corresponding to format text and location information of format text within the document (610) of the specific format). Then, the text extraction model (520) can extract text from each of the multiple layout areas (611, 612, 614, 615) containing text. This text extraction model (520) can also be understood as a model that performs optical character recognition (OCR) to convert characters (621, or text images) contained in an image or document into text data (621a) that can be recognized and read by a machine (or computer). In the present invention, the text extraction model (520) may also be named an “OCR model”, a “text model”, a “text processing model”, a “text understanding model”, or a “text specialized model”.

[0142] Additionally, when a document (610) of a specific format is input, the table extraction model (530) can recognize a layout area (613) containing a table (613a) from a document (610) of a specific format based on the output of the document layout analysis model (510) (e.g., a label corresponding to the table and location information of the table within the document (610) of the specific format). Then, the table extraction model (530) can extract the table (613a) from the layout area (613) containing the table (613a). In the present invention, the table extraction model (530) may also be named a “table model,” a “table processing model,” a “table understanding model,” or a “table specialized model.”

[0143] Additionally, the table extraction model (530) may be configured to perform table structure recognition (TSR). Table structure recognition may mean extracting the logical and physical structure of an unstructured table image into a machine-readable format. For example, the table extraction model (530) may recognize the logical and physical structure of a table (613a) and convert the table (613a) into a specific pre-set format (e.g., “HTML”, 613b).

[0144] Furthermore, when a document (610) of a specific format is input, the chart extraction model (540) can recognize a layout area (617) containing a chart (617a) from a document (610) of a specific format based on the output of the document layout analysis model (510) (e.g., a label corresponding to the chart and location information of the chart within the document (610) of the specific format). Then, the chart extraction model (540) can extract the chart (617a) from the layout area (617) containing the chart (617a). In the present invention, the chart extraction model (540) may also be named a “chart model,” a “chart processing model,” a “chart understanding model,” or a “chart specialized model.”

[0145] Additionally, the chart extraction model (540) can perform the role of analyzing the chart (617a) to extract raw data of the chart. For example, the chart extraction model (540) can process the chart (617a) to structure information about the entire chart (617a) into a specific format (e.g., “JSON”) and generate (or output) structured data (617b) based on this. Additionally, the chart extraction model (540) can reconstruct the visual result of the chart based on the structured data (617b) to generate (or output) reconstructed data. Additionally, the chart extraction model (540) can recognize numerical information such as axis labels and data points from the layout area (617) containing the chart (617a) and convert it into numerical data in the form of a table.

[0146] Meanwhile, if a document received from a user terminal (10) contains content (or data, information, etc.) related to a chemical domain (or chemical field), the document understanding system (1000) may input the document into at least one model specialized in the chemical domain. For example, in the present invention, the model specialized in the chemical domain may include at least one of a molecular structure formula extraction model (550), a chemical reaction information extraction model (560), and a molecular structure formula conversion model (570).

[0147] First, the document understanding system (1000) can process a document (610) of a specific format as input to a molecular structure formula extraction model (550) in order to extract at least one molecular structure formula among a plurality of contents included in the document.

[0148] Here, the molecular structure extraction model (550) may be configured to accurately detect regions corresponding to (or corresponding to) molecular structure formulas within a document. More specifically, the molecular structure extraction model (550) may perform molecular detection (or molecular detection) by identifying and locating at least one and / or at least one molecular structure formula contained in a document. This molecular structure extraction model (550) may be a model trained using at least one training data set to accurately detect molecular structure formulas.

[0149] In one embodiment, when a document (610) of a specific format is input, the molecular structure formula extraction model (550) recognizes a layout area (616) containing a molecular structure formula and can extract a molecular structure formula from the recognized layout area (616). In the present invention, the molecular structure formula extraction model (550) may also be named a “molecular structure formula specialized model,” a “molecular structure formula model,” a “molecular structure formula processing model,” or a “molecular structure formula understanding model,” etc.

[0150] Additionally, the document understanding system (1000) can process a document (610) of a specific format as input to a chemical reaction information extraction model (560) in order to extract at least one chemical reaction information among a plurality of contents included in the document.

[0151] Here, the chemical reaction information extraction model (560) may be configured to perform the role of extracting chemical reaction information consisting of reactants, reaction conditions, and products within a document. The chemical reaction information extraction model (560) may perform the role of identifying the location of these components from each chemical reaction information and assigning an appropriate class to the corresponding area. For example, the chemical reaction information extraction model (560) may perform reaction parsing (or reaction parsing) to extract and classify at least one and / or at least one chemical reaction information contained in a document. This chemical reaction information extraction model (560) may be a model trained using at least one training data set to extract and classify reaction roles within a chemical reaction diagram.

[0152] In one embodiment, when a document (610) of a specific format is input, the chemical reaction information extraction model (560) recognizes a layout area (not shown) containing chemical reaction information and can extract chemical reaction information from the recognized layout area. In the present invention, the chemical reaction information extraction model (560) may also be named a “chemical reaction information specialized model,” a “chemical reaction information model,” a “chemical reaction information processing model,” or a “chemical reaction information understanding model,” etc.

[0153] However, in the present invention, recognizing a layout area containing molecular structural formulas and chemical reaction information included in a document, and extracting molecular structural formulas and chemical reaction information from the recognized layout area, respectively, may be performed by a document layout analysis model (510).

[0154] Furthermore, the molecular structure formula conversion model (570) may be configured to perform the role of converting the molecular structure formula into at least one pre-set chemical structure representation format (e.g., SMILES, InChI, Mol, etc.). For example, the molecular structure formula conversion model (570) may recognize the molecular structure formula, analyze the bonding relationships between the atoms constituting the molecular structure formula and the atoms, and convert the molecular structure formula into at least one chemical structure representation format. Here, the chemical structure representation format may mean that the molecular structure formula has been converted into a representation that a machine (or computer) can understand and use.

[0155] In this case, the document understanding unit (500) inputs the molecular structural formula extracted from the molecular structural formula extraction model (550) into the molecular structural formula conversion model (570) and can obtain a chemical structure representation format for the molecular structural formula converted through the molecular structural formula conversion model (570). Alternatively, the document understanding system (1000) inputs a document (610) of a specific format into the molecular structural formula conversion model (570) and can recognize the molecular structural formula contained in the document (610) of the specific format in the molecular structural formula conversion model (570) and convert it into a chemical structure format.

[0156] Next, the content order sorting model (580) may be configured to perform the role of specifying (or determining) order information of at least one and / or at least one content extracted from the document.

[0157] Additionally, the content order sorting model (580) may be configured to perform the role of sorting at least one and / or at least one piece of content extracted from a document in a pre-set order (e.g., human reading order). This content order sorting model (580) may also be understood as a model that performs Reading Order Detection (ROD). Here, Reading Order Detection may mean sorting (or arranging) text in a document (or document image) having various layouts in a logical order in which a human reads. In one embodiment, Reading Order Detection may mean a rule for sorting multiple text areas (or blocks) within a document in the order in which a human actually reads. In the present invention, the expression "sorting content in order" may also be referred to as "arranging content in order" or "arranging content in order."

[0158] As seen above, the document layout analysis model (510) identifies multiple layout areas each containing multiple contents through analysis of each page of the document. At this time, in order to understand a document that is visually diverse, it is necessary to organize the contents included in the multiple layout areas in the same way as a human reads.

[0159] To this end, the content ordering model (580) can perform tasks of arranging the content extracted from the document according to a preset order. For example, as illustrated in FIG. 7 (a) and (b), the content ordering model (580) can arrange the order of the content in the order in which the text is read, starting from a first layout area (e.g., “id 1”) located at the top left of the page of the document. At this time, since the content ordering model (580) arranges (or reconstructs) the order according to the body flow (or body text flow) of the document, i) figures, tables, charts, ii) format text such as page headers, page footers, journal names, etc., may be excluded during the content ordering process. Here, the text included in the fifth layout area (e.g., “id 5”) and the text included in the tenth layout area (e.g., “id 10”) may have a structure that leads to the same paragraph or consecutive sentences. In this case, the content ordering model (580) can arrange the text in the 10th layout area (e.g., “id 10”) so that it follows the text in the 5th layout area (e.g., “id 5”).

[0160] That is, the content ordering model (580) can arrange multiple layout areas (e.g., layout areas containing text) within a page of a document in a pre-set order (e.g., left → right, top → bottom, etc.), and when the content is consecutive, such as with the 5th layout area (e.g., “id 5”) and the 10th layout area (e.g., “id 10”), the two layout areas can be arranged to be connected continuously. However, in the present invention, the arrangement of the content order can also be performed by a multimodal model (590).

[0161] Next, the multimodal model (or multimodal artificial intelligence model, 590) may include a Vision-Language Model (VLM) that simultaneously understands and processes text and images. This multimodal model can integrally understand visual information (or visual content, visual elements) and text information (or text content, text elements, linguistic information, linguistic elements, etc.) contained in a document, and perform data processing to provide an optimal output (or output data, output result, outcome) according to a user request (or user purpose). In the present invention, the multimodal model (590) may also be named a “vision language model,” a “generative artificial intelligence model,” or a “generative model.”

[0162] In one embodiment, the multimodal model (590) can simultaneously understand text content and visual content (e.g., pictures, tables, charts, molecular structural formulas, etc.) included in each page (or image data) of the input document and / or converted document, identify semantic relationships between the contents, and perform data processing to provide optimal output according to user requests. For example, the data processing may include at least one of i) converting the document into an editable form, ii) converting the document into a form that meets the user's request (or needs), iii) summarizing the document, iv) translating the document, v) structuring the document into semantic units to generate structured data, or vi) reconstructing the document to generate a reconstructed document.

[0163] In this case, the input data of the multimodal model (590) may be data of a specific format (or form, style). For example, the specific format may include JSON (JavaScript Object Notation). This specific format of data may refer to a data format for structuring and expressing layout areas containing each of multiple contents extracted from a document, and order information, type information, characteristics (or attribute information), location information, etc., for each of the multiple contents. This specific format of data may be generated based on the output of the document layout analysis model (510) (e.g., identification results of multiple layout areas) and the output of the content order sorting model (580) (e.g., order of contents). Through this, the relationships between contents and contextual flow within the document are organized into a form that can be mechanically interpreted, and can be utilized as input data to enable the multimodal model (590) to perform processes such as converting the document into an editable form, reconstructing the document, converting the document into a form that meets the user's request, or translating the document in subsequent processes. Additionally, data of a specific format can be utilized as information necessary to rearrange or visualize content within a document in a user interface. This data of a specific format can be understood as data generated by a document understanding system (1000) or a document understanding unit (500).

[0164] Next, the control unit (600) can perform the role of controlling the overall operation of the document understanding system (1000) related to the present invention. The control unit (600) can process signals, data, information, etc. that are input or output through the components of the document understanding system (1000) described above, or perform a series of data processing to provide or process appropriate information and functions to the user. The control unit (600) can be physically implemented by the processor described above.

[0165] Meanwhile, in the present invention, depending on the purpose or use of the document understanding system (1000), it may be implemented with at least one of the plurality of models (510, 520, 530, 540, 550, 560, 570, 580, 590) described above excluded. Alternatively, in the present invention, depending on the purpose or use of the document understanding system (1000), the user may disable (i.e., turn OFF) at least one of the plurality of models (510, 520, 530, 540, 550, 560, 570, 580, 590) so that it is not calculated and / or called during the document processing process. That is, in the present invention, by allowing a model to be selected according to the purpose or use of the document understanding system (1000), the processing speed and / or resource usage of the system can be optimized.

[0166] In addition, the present invention may be implemented to include at least one additional model depending on the purpose or use of the document understanding system (1000). Alternatively, the present invention may provide an environment in which a user can additionally configure a model depending on the purpose or use of the document understanding system (1000).

[0167] In one embodiment, as illustrated in FIG. 6b, the document comprehension unit (500) may further include a style module (e.g., “style”, 575) capable of processing a document by reflecting user preference style information (e.g., text style (font), character size, line spacing, boldness, etc.), writing style and / or tone style (e.g., tone of voice, sentence length, summary level, etc.), format style (e.g., table style, PPT style, report / thesis, patent document style, etc.), and color and / or design (e.g., color tone, accent color type, background brightness, etc.). More specific details regarding the style module (575) will be described later.

[0168] As described above, the present invention provides a document understanding system (1000) capable of effectively understanding various types of documents. More specifically, the document understanding system (1000) according to the present invention can provide a service and / or function that automatically converts multimodal information from various types of documents into data using Deep Document Understanding (DDU) technology. Below, we will examine in more detail the document understanding method performed by the document understanding system (1000) together with the attached drawings. FIG. 8 is a flowchart for explaining the document understanding method according to the present invention, and FIGS. 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25 are conceptual diagrams for explaining the document understanding method according to the present invention.

[0169] Meanwhile, as illustrated in FIG. 8, the document understanding method according to the present invention may include the steps of: receiving at least one document (S810); identifying a plurality of layout areas each containing a plurality of contents having different data characteristics through analysis of at least one document (S820); extracting a plurality of contents each containing a plurality of contents each by using a plurality of models specialized for each of the plurality of contents (S830); performing multimodal data processing on the plurality of contents having different data characteristics extracted from each of the plurality of models to generate a structured document understanding result related to the plurality of contents (S840); and providing an editing function for at least one document to a user terminal based on the structured document understanding result (S850).

[0170] Multimodal data processing for multiple contents involves integrating the processing of multiple contents corresponding to different modalities, which can also be understood as “performing integrated multimodal data processing for multiple contents having different data characteristics.”

[0171] A document understanding system (1000) can receive at least one document from a user terminal (10). For example, as illustrated in FIG. 9, the document understanding system (1000) can activate a graphic object (2001) based on the selection of a graphic object (e.g., “Edit”, 2001) associated with a request to provide a document editing function included in at least one page (2000) from the user terminal (10). Then, while the graphic object (2001) is activated, the document understanding system (1000) can receive (or identify) a document (900a) corresponding to the user input based on the input of a document (900a) into a document input area included in the page (2000) from the user terminal (10). Here, receiving user input for the document (900a) while the graphic object (2001) is active can also be understood as receiving a user request to convert the document (900a) into an editable form.

[0172] And, the document understanding system (1000) can check the extension (e.g., PDF, PPTX, HWPX, PNG, JPG, etc.) of the received document (900a) and convert the document (900a) with the confirmed extension into a specific format (or form, style, etc.) that is pre-set. For example, the specific format that is pre-set may include images, and this may be a process of converting at least one and / or at least one page included in the document (900a) into images.

[0173] Subsequently, the document understanding system (1000) can process a document converted into a specific format (or a document of a specific format, a document converted into an image, an image, a document image, etc.) as input to the document understanding unit (500). In this case, the document understanding unit (500) can process the document converted into a specific format as input to the document layout analysis model (510). Here, processing the document converted into a specific format as input to the document layout analysis model (510) may be performed by at least one processor.

[0174] A document layout analysis model (510) can analyze the structure of a document to identify multiple layout areas each containing multiple contents having different data characteristics. Here, analyzing the structure of a document may mean a series of processing steps to recognize the arrangement and interrelationships of various contents constituting the document and to identify the logical and spatial structure of the document. For example, it may include a process of extracting document constituent units such as multiple contents (e.g., text (plain text and formatted text, etc.), tables, charts, pictures, etc.) and determining (or identifying, judging, etc.) at least one of the location, boundary, hierarchical structure, order, and relationship of each constituent unit to model the layout structure of the entire document. This may provide structural information necessary for reconstructing the visual arrangement structure and logical flow of the document, or for content extraction and / or document editing and / or document conversion processes.

[0175] As illustrated in FIG. 10, the document layout analysis model (510) can identify a plurality of layout areas (901, 902, 903, 904, 905, 907, 908, 910, 911, 912, 913, 914, 915, 916, 917, 918) each containing a plurality of contents through analysis of a document (900) converted into a specific format. Here, identification is a task of specifically distinguishing what the detected area and / or object is in the document (or image corresponding to each page of the document), and accurately confirming (or identifying) the identity (or type, kind, etc.) of the content included in the area. For example, recognition is generally a process of identifying the specific identity of a detected object, and in the present invention, the expression “identify” may also be referred to as “recognize.”

[0176] In one embodiment, the document layout analysis model (510) can identify a plurality of layout areas (901, 902, 904, 905, 907, 908, 910, 911, 913, 916, 918) containing text from a document (900) converted into a specific format. In the present invention, the layout areas containing text may also be named “text layout areas,” “text areas,” or “text blocks,” etc.

[0177] In another embodiment, the document layout analysis model (510) can identify a layout area (912) containing a table from a document (900) converted into a specific format. In the present invention, the layout area (912) containing a table may also be named a “table layout area,” a “table area,” or a “table block,” etc.

[0178] In another embodiment, the document layout analysis model (510) can identify a layout area (915) containing a chart from a document (900) converted into a specific format. In the present invention, the layout area containing a chart may also be named a “chart layout area,” a “chart area,” or a “chart block,” etc.

[0179] In another embodiment, the document layout analysis model (510) can identify a plurality of layout areas (903, 917) containing pictures from a document (900) converted into a specific format. In the present invention, the layout areas containing pictures may also be named “picture layout areas,” “picture areas,” or “picture blocks,” etc. In this case, the pictures may be content that is recognized and processed as information different from tables and charts.

[0180] Furthermore, the document understanding unit (500, or at least one processor) can extract multiple contents included in each of the multiple layout areas by using multiple models specialized for each of the multiple contents. More specifically, the document comprehension unit (500) specifies a plurality of models for extracting a plurality of contents included in each of a plurality of layout areas (901, 902, 903, 904, 905, 907, 908, 910, 911, 912, 913, 914, 915, 916, 917, 918), and by using (or calling) each of the specified plurality of models, the plurality of contents included in each of the plurality of layout areas (901, 902, 903, 904, 905, 907, 908, 910, 911, 912, 913, 914, 915, 916, 917, 918) can be extracted. Alternatively, each of the multiple models can recognize multiple contents included in each of the multiple layout areas (901, 902, 903, 904, 905, 907, 908, 910, 911, 912, 913, 914, 915, 916, 917, 918). Alternatively, each of the multiple models can recognize multiple contents included in each of the multiple layout areas (901, 902, 903, 904, 905, 907, 908, 910, 911, 912, 913, 914, 915, 916, 917, 918) and extract the recognized content.

[0181] In this regard, the term “recognizing content” in the present invention may be understood to encompass both the process of recognizing content and the process of extracting recognized content. Alternatively, the term “extracting content” in the present invention may be understood to encompass both the process of recognizing content and the process of extracting recognized content. Alternatively, the process of “recognizing content” and the process of “extracting content” in the present invention may be implemented as separate processes. However, for the convenience of explanation, the following description will not be limited to any one of these methods.

[0182] The document comprehension unit (500) can process at least one of the documents (900) converted into a specific format including the output of the document layout analysis model (510) and a plurality of layout areas (901, 902, 903, 904, 905, 907, 908, 910, 911, 912, 913, 914, 915, 916, 917, 918) as input to each of the plurality of models.

[0183] Alternatively, the document comprehension unit (500) may call a plurality of models for extracting a plurality of contents included in a document (900) converted into a specific format based on the output of the document layout analysis model (510) (e.g., labels corresponding to each of the plurality of contents, location information of each of the plurality of contents included in the document (900) converted into a specific format, etc.). In this case, each of the called plurality of models may extract a plurality of contents included in each of the plurality of layout areas based on the output of the document layout analysis model (510).

[0184] In one embodiment, the document understanding unit (500) may call a text extraction model (520) for extracting text from each of a plurality of layout areas (901, 902, 904, 905, 907, 908, 910, 911, 913, 916, 918) containing text. The text extraction model (520) may extract text from each of the plurality of layout areas (901, 902, 904, 905, 907, 908, 910, 911, 913, 916, 918) identified as containing text through the document layout analysis model (510).

[0185] In another embodiment, the document understanding unit (500) may call a table extraction model (530) for extracting a table from a layout area (912) containing a table. The table extraction model (530) may extract a table from a layout area (912) identified as containing a table through the document layout analysis model (510) and convert it into a specific format (e.g., “HTML”). In this case, the table extraction of the table extraction model (530) may be understood as extracting raw data of the table.

[0186] In another embodiment, the document comprehension unit (500) may call a chart extraction model (540) for extracting a chart from a layout area (915) containing a chart. The chart extraction model (540) may extract a chart from a layout area (915) identified as containing a chart through the document layout analysis model (510), and structure information about the entire extracted chart into a specific format (e.g., JSON) that is pre-set. In this case, the chart extraction of the chart extraction model (540) may be understood as extracting raw data of the chart.

[0187] Additionally, the document understanding system (1000) can process the document (900) converted into a specific format as input to the molecular structure formula extraction model (550) and the chemical reaction information extraction model (560), respectively, in order to identify (or recognize) and extract at least one of the molecular structure formula and chemical reaction information contained in the document (900) converted into a specific format.

[0188] In one embodiment, the molecular structural formula extraction model (550) recognizes a layout area (909) containing a molecular structural formula within a document (900) converted into a specific format, and can extract a molecular structural formula from the recognized layout area (909). In the present invention, the layout area containing a molecular structural formula may also be named a “molecular structural formula layout area,” a “molecular structural formula area,” or a “molecular structural formula block,” etc.

[0189] In another embodiment, the chemical reaction information extraction model (560) recognizes a layout area (906) containing chemical reaction information consisting of reactants, reaction conditions, and products within a document (900) converted into a specific format, and can extract chemical reaction information from the recognized layout area (906). In the present invention, the layout area containing chemical reaction information may also be named a “chemical reaction information layout area,” a “chemical reaction information area,” or a “chemical reaction information block,” etc.

[0190] In another embodiment, the document comprehension unit (500) processes the molecular structural formula extracted from the molecular structural formula extraction model (550) as input to the molecular structural formula conversion model (570), and the molecular structural formula conversion model (570) can convert the molecular structural formula into at least one pre-set chemical structure representation format (e.g., SMILES).

[0191] Furthermore, the document understanding system (1000, or document understanding unit (500)) can perform multimodal data processing on the plurality of contents having different data characteristics to generate a structured document understanding result related to the plurality of contents. And, based on the structured document understanding result, the document understanding system (1000) can provide an editing function for at least one document to a user terminal.

[0192] Here, multimodal can refer to a technology in which artificial intelligence simultaneously understands and processes various forms of data (modalities), such as text, images, voice, and video.

[0193] In the present invention, performing multimodal data processing on a plurality of contents to generate a structured document understanding result related to a plurality of contents may mean structuring a plurality of contents of different modalities or formats (e.g., text, images, charts, tables, formulas, molecular structural formulas, chemical reaction information, etc.) extracted from a document through different models.

[0194] Furthermore, as mentioned above, “multimodal data processing for multiple contents” involves integrating and performing processing for multiple contents corresponding to different modalities, which can also be understood as “performing integrated multimodal data processing for multiple contents having different data characteristics.”

[0195] That is, the term “multimodal data processing” in this specification may mean a process of integrating multiple contents having different data characteristics or modalities while maintaining semantic and structural relationships between contents, taking into account the expression format, semantic characteristics, and positional relationships within the document of each content, rather than simply combining or listing multiple contents in parallel.

[0196] The above multimodal data processing is performed after the content extraction step and can be performed to generate a structured document understanding result by aligning multiple extracted contents into a common presentation space, matching or inferring relationships based on spatial arrangement, semantic similarity, or document context between contents, and fusing features of different modalities.

[0197] In other words, the above multimodal data processing can be understood as an integrated processing step for reorganizing multiple contents based on document semantics and structure.

[0198] More specifically, it may mean integratively analyzing each piece of content extracted from a document by considering its semantic associations, spatial arrangement, and logical relationships, and based on this, including (or reflecting) the document's meaning, structure (e.g., title, body text, tables, captions, relationships between each piece of content, etc.) and relationships between components, and organizing (or reconstructing) it into a structured data form (or representation) that can be interpreted by a machine (or computer).

[0199] As an example, a multimodal data integration processing process for multiple contents may include integrally understanding information between different modalities (i.e., multiple different contents) and interpreting the meaning, structure, and layout of the document to generate a document understanding result that is editable, convertible, searchable, or translated. In this case, the multimodal data integration processing process may include a process of combining the recognition results of different modalities into a single integrated expression by mutually merging (or aligning) them based on positional relationships, semantic associations, and document context, rather than simply merging them. In this case, the process may include mapping text and images corresponding to the same area, establishing correspondence relationships between table and chart images and corresponding raw data, and analyzing semantic connection relationships between body text, titles, comments, captions, shapes, and formulas.

[0200] Alternatively, the multimodal data integration processing process for multiple contents may include a processing process that processes (or interprets) contents having different data characteristics through a pre-configured processing technique (or method) and generates a structured document understanding result by reflecting the semantic and structural relationships between each content.

[0201] Here, the pre-configured processing technique may include a multimodal data integration processing technique and / or a multimodal analysis technique. In one embodiment, the pre-configured processing technique may include at least one of i) a technique for normalizing features by vectorizing or embedding extracted information and mapping them to a common representation space, ii) a technique for matching and / or inferring relationships between contents based on the spatial arrangement or semantic similarity of the contents, iii) a technique for fusing features of different modalities in an Early fusion and / or Late fusion and / or Hybrid fusion manner, and v) a technique for aligning semantic correspondence relationships between contents having different data characteristics. In the present invention, the multimodal data integration processing process for content may include a process of processing contents having different data characteristics according to the pre-configured processing technique and generating a structured document understanding result by reflecting the semantic and structural relationships between each content.

[0202] Here, the pre-configured processing technique may include at least one of a multimodal data integration processing technique and a multimodal analysis technique. For example, the processing technique may include vectorization or embedding processing for mapping features of extracted content to a common representation space, relationship matching or inference based on the spatial arrangement or semantic similarity of the content, and at least one of early fusion, late fusion, or mixed fusion methods for fusing features of different modalities.

[0203] Based on this processing, the document understanding result can be structured while maintaining semantic correspondence between content with different data characteristics.

[0204] The structured document understanding result generated based on this can include an outcome in which each component within the document is expressed in a structurally defined data form, along with information on its role, attributes, and interrelationships, by reflecting both the visual composition and semantic content of the document. In other words, it can include an outcome generated by understanding and reconstructing the various contents constituting the document based on meaning and structure.

[0205] Furthermore, providing an editing function for at least one document to a user terminal may include the process of converting a document (900) of a specific format into an editable form based on (or based on, using, etc.) the result of understanding the structured document, and providing it to the user terminal (10).

[0206] This process can also be understood as a process of converting a document (900) of a specific format into an editable form using multiple contents extracted from each of multiple models, and providing the document converted into an editable form to a user terminal (10).

[0207] Alternatively, based on the results of understanding the structured document, the process may include converting a document (900) of a specific format into an editable form and providing an editing function for the document (or a plurality of contents) together with the document converted into an editable form to the user terminal (10).

[0208] In this specification, the term “structured document understanding result” may mean a result in which a plurality of contents extracted from a document are not merely a collection of visual objects, image data, or unstructured information, but are expressed in a data structure in which the type, location, role, attributes, and semantic and structural relationship information of each content within the document are explicitly defined.

[0209] The above-mentioned structured document understanding result can be generated by reflecting the visual layout structure and logical compositional flow of the document, and may include data configured to enable mechanical processing, editing, conversion, or reconstruction at the content unit level and at the relationship unit level between content.

[0210] Accordingly, the results of the structured document understanding described above can be utilized as technical input data for providing document editing functions, converting document formats, or subsequent multimodal processing. Here, a document converted into an editable form may refer to a document that has been reconstructed to allow modification, addition, deletion, etc., while maintaining the meaning and structure of multiple contents included in the document. That is, it may refer to a document that has been reconstructed to allow for free modification and editing in various document formats (e.g., PPT, Excel, Word, etc.) by recognizing and extracting each content constituting the document from a document in which modification and / or structural editing of the contents is restricted.

[0211] In this case, the document understanding system (1000) can determine (or determine) the order information of each of the multiple contents extracted from a document (900) of a specific format by using a content order sorting model (580).

[0212] As seen above, the content order sorting model (580) can determine a logical reading order (e.g., ROD) corresponding to the flow of a person actually reading a document for multiple contents extracted within a document, and perform the role of sorting the content according to the determined order.

[0213] In one embodiment, the content order sorting model (580) receives the output of the document layout analysis model (510) as input, analyzes the structural features of the layout included in the document (900) of a specific format, and assigns a sequential index to each layout area (or each content). Through this, the content order sorting model (580) enables multiple contents included in the document (900) of a specific format to be sorted (or arranged, placed, etc.) in a semantically consistent order, and can provide a standard that is referenced when the document comprehension unit (500) subsequently generates data of a specific format (e.g., “JSON”). That is, the content order sorting model (580) restores the structural context from the document and provides information that enables the subsequent multimodal model (590) to generate various content following the flow of the original document.

[0214] Subsequently, the multimodal model (590) can convert a document (900) of a specific format into an editable form based on at least one of the order information specified through the content order sorting model (580), the output of the document layout analysis model (510), and multiple contents extracted from each of the multiple models.

[0215] In this case, the document understanding system (1000) may configure a prompt including at least one of the order information specified through the content order sorting model (580), the output of the document layout analysis model (510), multiple contents extracted from each of the multiple models, and a user request, and input the prompt into the multimodal model (590). The multimodal model (590) may convert a document (900) of a specific format into an editable form using (or based on) the input prompt.

[0216] Alternatively, the document understanding system (1000) may integrate the identification results of each of a plurality of layout areas identified through the document layout analysis model (510) with the order information of each of a plurality of contents specified through the content order sorting model (580), and input the integrated data into the multimodal model (590). The multimodal model (590) may use the input integrated data to convert a document (900) of a specific format into an editable form.

[0217] Alternatively, the document understanding system (1000) may input data of a specific format (e.g., JSON) into the multimodal model (590). The multimodal model (590) may use the input data of the specific format to convert a document (900) of the specific format into an editable form.

[0218] In another embodiment, the multimodal model (590) can i) analyze the meaning of each of the multiple contents to summarize the content of the document, ii) reconstruct sentences, or iii) extract key keywords. Additionally, the multimodal model (590) can recognize and connect reference relationships between text, images, charts, and tables among the extracted multiple contents, and perform structural and descriptive transformations necessary for an output form (e.g., file format) corresponding to a user request. Furthermore, the multimodal model (590) can generate natural narratives that match the context or style of the document, and, if necessary, perform additional semantic judgments for visual editing, such as changing the background or highlighting elements, through image-based context analysis.

[0219] Furthermore, as illustrated in FIG. 11, the document understanding system (1000) may provide a document (1300) converted into an editable form through at least one page (2000) output (or provided) to the user terminal (10). The document (1300) converted into an editable form may include a plurality of contents (1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318) extracted from each of a plurality of models. For example, the plurality of contents may include at least one of text (1301, 1302, 1304, 1305, 1307, 1308, 1310, 1311, 1313, 1314, 1316, 1318), a table (1312), a chart (1315), a molecular structural formula (1309), chemical reaction information (1306), and a figure (1303, 1317).

[0220] In this case, the document understanding system (1000) can provide an editing function for each of the multiple contents (1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318) extracted together with the document (1300) converted into an editable form, based on the fact that the user request received from the user terminal (10) is a request to provide a document editing function.

[0221] Here, providing an editing function may mean providing at least one and / or at least one editing tool that provides an editing function for each of the multiple contents so that the user can edit each of the multiple contents.

[0222] In one embodiment, the document understanding system (1000) may provide a text editing tool (1330) that provides an editing function for text (1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318) among a plurality of extracted contents (1301, 1302, 1304, 1305, 1307, 1308, 1310, 1311, 1313, 1314, 1316, 1318). For example, the text editing tool (1330) may include a tool capable of editing (or modifying) at least one of the font, size, color, and paragraph of the text. However, the tools included in the text editing tool (1330) are not limited to the examples mentioned above, and it is obvious that various additional tools may be included.

[0223] In another embodiment, the document understanding system (1000) may provide a table editing tool (e.g., “table editing tool”, 1340) that provides an editing function for a table (1312) among a plurality of extracted contents (1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318). For example, the table editing tool (1340) may include a tool capable of editing at least one of the cells of the table, the rows of the table, and the columns of the table. However, the tools included in the table editing tool (1340) are not limited to the examples mentioned above, and it is obvious that various additional tools may be included.

[0224] In another embodiment, the document understanding system (1000) may provide a chart editing tool (1350) that provides an editing function for a chart (1315) among a plurality of extracted contents (1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318). For example, the chart editing tool (1350) may include a tool capable of editing at least one of the chart style and the chart data. However, the tool included in the chart editing tool (1350) is not limited to the examples mentioned above, and it is obvious that various additional tools may be included.

[0225] Although not illustrated, the document understanding system (1000) may provide a document background editing tool that provides editing functions for the background of a document (1300) converted into an editable form. For example, the document background editing tool may include a tool capable of editing at least one of the overall layout of the document, the theme of the document, the background style of the document, the color of the document (or the page color of the document), and the margins of the document. However, the tools included in the document background editing tool are not limited to the examples mentioned above, and it is obvious that various additional tools may be included.

[0226] These extracted multiple contents (1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318) may be configured to be edited based on user input to at least one editing tool. In one embodiment, when editing is performed on at least one of the extracted multiple contents (1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318), the document understanding system (1000) can update the document (1300) converted into an editable form in real time so that at least one edited content is displayed on the user terminal (10).

[0227] Meanwhile, at least one editing tool provided together with the document converted into an editable form may be provided to the user terminal (10) based on at least one data characteristic that each of the extracted multiple contents has.

[0228] Here, data characteristics include unique data attributes (or properties) possessed by the content (or by the information constituting the content), and may include technical and logical attributes of the information constituting the content, such as the format, structure, mode of representation, internal composition, arrangement of data elements, encoding method, and file format of the content. For example, data characteristics may include i) morphological attributes of the data, such as text, tables, charts, graphs, formulas, molecular structural formulas, and chemical reaction information; ii) structural attributes, such as internal layer configuration, hierarchical structure, metadata, and whether the data is vector or raster; iii) functional attributes, such as the extractability, manipulability, transformability, and rules affecting the processing of the data; and iv) logical and technical attributes, such as relationships between data elements and methods of semantic expression. Furthermore, data characteristics include technical attributes formed during the process of creating, storing, transmitting, and displaying the content, and may be utilized as a criterion for selecting or providing appropriate editing tools or processing methods for the content. In the present invention, data characteristics may also be referred to as “data attributes.”

[0229] Specifically, the document understanding system (1000) may provide at least one and / or at least one editing tool that provides an editing function for each content based on the data characteristics of each content extracted through a plurality of models.

[0230] In one embodiment, among a plurality of contents, a first content (e.g., text) may have a first data characteristic including a unique data characteristic of the text. In this case, the document understanding system (1000) may provide an editing tool that provides an editing function for the first content based on the fact that the first content has a single data characteristic. This single editing tool may be configured to enable editing of the first content having a single data characteristic.

[0231] Alternatively, among the multiple contents, a second content (e.g., a table) different from the first content may have a first data characteristic including the unique data characteristic of the table and a second data characteristic including the unique data characteristic of the image. In this case, the document understanding system (1000) may provide a plurality of editing tools to a user terminal (10) that provide an editing function for the second content based on the fact that the second content has a plurality of data characteristics (or a plurality of different data characteristics). These plurality of editing tools may be configured to enable editing of the second content having both the first data characteristic and the second data characteristic simultaneously.

[0232] The document understanding system (1000) can receive an editing request for selected content based on the selection of at least one of a plurality of contents included in a document converted into an editable form from a user terminal (10). For example, as illustrated in FIG. 12, a document understanding system (1000) receives a user input selecting one of a plurality of contents (1401, 1402, 1403, 1404, 1405, 1406, 1407, 1408, 1409, 1410, 1411, 1412, 1413, 1414, 1415, 1416, 1417, 1418, 1419, 1420, 1421, 1422, 1423) included in the document (1400) converted into an editable form from a user terminal (10) on which at least one page (2000) including the document (1400) converted into an editable form is output, and requests to edit one of the contents (1406). It can be received. For convenience of explanation, “one of the contents (1406)” will be referred to as “the first content (1406)”.

[0233] Here, the statement that a user input for selecting content is received can also be understood as “receiving a user input for selecting an area corresponding to the content” or “receiving a user input for selecting a graphic object associated with an editing request receiving function output (or provided) to an area surrounding the content.”

[0234] And, the document understanding system (1000) can determine (or analyze, specify, etc.) at least one data characteristic of the first content (1406) for which an editing request has been received. Based on the result of the determination (or analysis result, specific result, etc.), the document understanding system (1000) can determine that the first content (1406) has multiple data characteristics simultaneously, based on the fact that the first content (1406) has a first data characteristic including a unique data characteristic of a chart and a second data characteristic including a unique data characteristic of an image.

[0235] Subsequently, the document understanding system (1000) may provide a plurality of editing tools (1430, 1440) that provide an editing function for the first content (1406) to the user terminal (10, or service page (2000)) based on the above judgment result. Here, the first content (1406) may include a chart corresponding to at least some of the content from which raw data was extracted during the content extraction process. As described above, the chart extraction model (540) can recognize the chart included in the chart area of ​​the document and extract the raw data of the chart. That is, the document understanding system (1000) may provide a plurality of editing tools (1430, 1440) that provide an editing function for the chart from which raw data was extracted to the user terminal (10).

[0236] These multiple editing tools (1430, 1440) can be configured to enable editing (or modification) of the first content (1406) having the first data characteristics and the second data characteristics simultaneously.

[0237] For example, among a plurality of editing tools (1430, 1440), a first editing tool (e.g., “chart editing tool”, 1430) may provide an editing function for a first content (1406) having a first data characteristic. In this case, the first editing tool (1430) may include a tool capable of editing at least one of the style and data (or data value) of the first content (1406). However, the tool included in the first editing tool (1430) is not limited to the examples mentioned above, and it is obvious that it may include various additional tools capable of editing charts.

[0238] As another example, among the multiple editing tools (1430, 1440), the second editing tool (1430) may provide an editing function for the first content (1406) having the second characteristic. In this case, the second editing tool (1440) may include a tool capable of editing at least one of the style, visual appearance, and size of the first content (1406). However, the tool included in the second editing tool (1440) is not limited to the example mentioned above, and it is obvious that it may include various additional tools capable of editing images.

[0239] In another embodiment, as illustrated in FIG. 13, a document understanding system (1000) receives a user input selecting one of a plurality of contents (1401, 1402, 1403, 1404, 1405, 1406, 1407, 1408, 1409, 1410, 1411, 1412, 1413, 1414, 1415, 1416, 1417, 1418, 1419, 1420, 1421, 1422, 1423) included in the document (1400) converted into an editable form from a user terminal (10) on which at least one page (2000) including the document (1400) converted into an editable form is output, and edits the one of the contents (1408). A request can be received. For convenience of explanation, “one of the contents (1408)” will be referred to as “the second content (1408)”.

[0240] And, the document understanding system (1000) can determine (or analyze, specify, etc.) at least one data characteristic of the second content (1408) for which an editing request has been received. Based on the result of the determination (or analysis result, specific result, etc.), the document understanding system (1000) can determine that the second content (1408) has multiple data characteristics simultaneously, based on the fact that the second content (1408) has a first data characteristic including a unique data characteristic of a table and a second data characteristic including a unique data characteristic of an image.

[0241] Subsequently, the document understanding system (1000) may provide a plurality of editing tools (1450, 1460) that provide an editing function for the second content (1408) to the user terminal (10, or service page (2000)) based on the above judgment result. Here, the second content (1408) may include a table corresponding to at least some of the content from which raw data was extracted during the content extraction process. As seen above, the table extraction model (530) can recognize a table included in the table area of ​​the document and extract the raw data of the table. That is, the document understanding system (1000) may provide a plurality of editing tools (1450, 1460) that provide an editing function for the table from which raw data was extracted to the user terminal (10).

[0242] These multiple editing tools (1450, 1460) can be configured to enable editing of a second content (1408) having both a first data characteristic and a second data characteristic simultaneously.

[0243] For example, among a plurality of editing tools (1450, 1460), a first editing tool (e.g., “table editing tool”, 1450) may provide an editing function for a second content (1408) having a first data characteristic. In this case, the first editing tool (1450) may include a tool capable of editing at least one of the rows and columns, cells, size (e.g., cell size, table size, etc.), and data (or data value) of the second content (1408). However, the tool included in the first editing tool (1450) is not limited to the examples mentioned above, and it is obvious that it may include various additional tools capable of editing a table.

[0244] As another example, among the multiple editing tools (1450, 1460), the second editing tool (1460) may provide an editing function for the second content (1408) having the second characteristic. In this case, the second editing tool (1460) may include a tool capable of editing at least one of the attributes, visual appearance, border, and effect of the second content (1408). However, the tool included in the second editing tool (1460) is not limited to the example mentioned above, and it is obvious that it may include various additional tools capable of editing images.

[0245] Meanwhile, the editing tool described above may include at least one editing tool that provides an editing function for at least one molecular structure among multiple contents.

[0246] The document understanding system (1000) can receive an editing request for a selected molecular structural formula from a user terminal (10) based on the selection of at least one molecular structural formula among a plurality of contents included in a document converted into an editable form. For example, as illustrated in FIG. 14, a document understanding system (1000) receives a user input selecting a molecular structural formula (1421) among a plurality of contents (1401, 1402, 1403, 1404, 1405, 1406, 1407, 1408, 1409, 1410, 1411, 1412, 1413, 1414, 1415, 1416, 1417, 1418, 1419, 1420, 1421, 1422, 1423) included in the document (1400) converted into an editable form from a user terminal (10) on which at least one page (2000) including the document (1400) converted into an editable form is output, and requests to edit the molecular structural formula (1421). Can receive.

[0247] Here, the statement that a user input for selecting a molecular structure formula is received can also be understood as “receiving a user input for selecting an area corresponding to the molecular structure formula” or “receiving a user input for selecting a graphic object associated with an editing request receiving function output (or provided) to an area surrounding the molecular structure formula.”

[0248] Subsequently, the document understanding system (1000) may provide an editing tool to a user terminal (10, or service page (2000)) that provides an editing function for a selected molecular structural formula (1421). Here, providing an editing tool that provides an editing function for a selected molecular structural formula (1421) can also be understood as providing an editing interface that provides an editing function for a selected molecular structural formula (1421). For example, as illustrated in FIG. 15, the document understanding system (1000) may provide an editing interface (1510) to a user terminal (10) that provides an editing function for a molecular structural formula (1421). This editing interface (1510) may include nodes corresponding to each of the atoms constituting the molecular structural formula (1421) and edges representing the bonding relationship of the atoms.

[0249] Furthermore, the document understanding system (1000) can update the editing interface (1510) in real time so that when editing of the molecular structural formula (1421) is performed based on user input for at least one editing tool included in the editing interface (1510), the edited molecular structural formula is displayed on the user terminal (10). In one embodiment, the editing of the molecular structural formula (1421) may be the deletion or relocation of at least one of the nodes corresponding to each of the atoms constituting the molecular structure and the edges representing the bonding relationships of the atoms, or the addition of a new node corresponding to a new atom or the addition of a new edge creating a new bonding relationship for the atoms.

[0250] Meanwhile, the document understanding system (1000) can provide recommendation information for an editing tool that provides an editing function for each of the extracted multiple contents.

[0251] For example, as illustrated in FIG. 16, a document understanding system (1000) receives a user input selecting one of a plurality of contents (1401, 1402, 1403, 1404, 1405, 1406, 1407, 1408, 1409, 1410, 1411, 1412, 1413, 1414, 1415, 1416, 1417, 1418, 1419, 1420, 1421, 1422, 1423) included in the document (1400) converted into an editable form, from a user terminal (10) to which at least one page (2000) including the document (1400) converted into an editable form is output, and based on this input, the system selects one of the contents (1406) that is the subject of an editing request. You can receive a request to edit the content (1406).

[0252] At this time, the document understanding system (1000) may provide recommendation information for recommending an editing tool that provides an editing function for any one of the content (1406) selected from the user terminal (10). For example, the document understanding system (1000) may provide recommendation information for an editing tool that provides an editing function for any one of the content (1406) based on the data characteristics of the one of the content (1406) (e.g., “If you want to modify the style or data of the chart, I recommend using Chart Editing Tool 1!”, “If you want to modify the overall size of the chart, I recommend using Chart Editing Tool 2!”, 1610).

[0253] Meanwhile, the document understanding system (1000) can provide recommendation information about the background of the document to the user terminal (10).

[0254] For example, as illustrated in FIG. 17(a), the document understanding system (1000) may provide the user terminal (10) with recommendation information regarding the background of a document converted into an editable form (e.g., “It would be better to change the background A currently applied to the background of the document to the recommended background below!”, 1710). In this case, the recommendation information (1710) may include an example comparing the background applied to the current document with the document to which the recommended background is applied, so that the user can intuitively recognize the recommended background. Additionally, the recommendation information (1710) may include at least one and / or at least one graphic object (1710a, 1710b, 1710c). In this case, the first graphic object (1710a) included in the recommendation information (1710) is associated with a function that maintains the current background, the second graphic object (1710b) is associated with a function that applies a recommended background, and the third graphic object (1710c) can be associated with a function that recommends a different recommended background.

[0255] Meanwhile, the document understanding system (1000) can provide recommendation style information for at least one of the extracted multiple contents to the user terminal (10).

[0256] For example, as illustrated in FIG. 17(b), the document understanding system (1000) may provide the user terminal (10) with recommended style information (e.g., “It would be better to change the font currently applied to the document’s font A to the recommended font below!”, 1720) for at least one of the multiple contents (e.g., text) included in a document converted into an editable form. In this case, the recommended style information (1720) may include an example comparing the style applied to the content with the content to which the recommended style is applied, so that the user can intuitively recognize the recommended style information. Additionally, the recommended style information (1720) may include at least one and / or at least one graphic object (1720a, 1720b, 1720c). In this case, the first graphic object (1720a) included in the recommended style information (1720) is associated with a function that maintains the current style of the content, the second graphic object (1720b) ​​is associated with a function that applies a recommended style to the content, and the third graphic object (1720c) can be associated with a function that recommends a different style.

[0257] The recommendation information regarding recommendation backgrounds and recommendation style information examined above may be information provided in accordance with the context of the document entered by the user, or information provided based on (or derived from) user history information. For example, user history information may include background and style information that the user frequently uses or prefers.

[0258] Meanwhile, the document understanding system (1000) can provide a translation for the entire document. For example, the control unit (600) can activate the graphic object (2002) based on the selection of a graphic object (e.g., “Translate”, 2002) associated with a request to provide a translation document included in at least one page (2000) from the user terminal (10) (see FIG. 9). And, as shown in FIG. 18, the document understanding system (1000) can receive user input for at least one document (e.g., document image, document in image form, image, etc., 1800) from the user terminal (10) while the graphic object (2002) is activated. Here, receiving user input for the document (1800) while the graphic object (2002) is activated can also be understood as receiving a user request to translate the entire content of the document (1800).

[0259] When a document (1800) to be translated is received, the document understanding system (1000) can use a document layout analysis model (510) to analyze the structure and separate the background of the document (1800). More specifically, as illustrated in FIG. 19, a document layout analysis model (510) can analyze a document (1800) and identify a plurality of layout areas (1801, 1802, 1803, 1804, 1805, 1806, 1807, 1808, 1809, 1810, 1811, 1812, 1813, 1814, 1815, 1816, 1817, 1818, 1819, 1820, 1821, 1822, 1823, 1824) each containing a plurality of contents in the document (1800).

[0260] And, the document understanding system (1000) can use a plurality of models to extract a plurality of contents included in each of a plurality of layout areas (1801, 1802, 1803, 1804, 1805, 1806, 1807, 1808, 1809, 1810, 1811, 1812, 1813, 1814, 1815, 1816, 1817, 1818, 1819, 1820, 1821, 1822, 1823, 1824), and perform a translation of at least one of the extracted plurality of contents. In this case, performing a translation can be understood as performing a translation of content corresponding to text, or performing a translation of text included in visual content (e.g., pictures, charts, tables, etc.). Alternatively, it can be understood as integrating the plurality of contents extracted from each of the plurality of models in a multimodal manner to generate a structured document understanding result, and performing translation of the document based on the structured document understanding result.

[0261] Furthermore, the document understanding system (1000) can provide a translated document to a user terminal (10). For example, as illustrated in FIG. 20, the document understanding system (1000) can provide a document (1850) to a user terminal (10) in which at least a portion of a plurality of contents has been translated.

[0262] Meanwhile, the document understanding system (1000) according to the present invention described above can understand the relationship between the flow of the document and the contents (e.g., text-images) and convert the document into a form (or format, etc.) that meets the user's needs, or summarize the contents of the document and provide it to the user.

[0263] In this regard, as illustrated in FIG. 26, the document understanding method according to the present invention may include: receiving a user request including at least one document from a user terminal (10) (S2610); identifying a plurality of layout areas each containing a plurality of contents having different data characteristics through analysis of at least one document (S2620); extracting a plurality of contents included in each of the plurality of layout areas each by using a plurality of models specialized for each of the plurality of contents each (S2630); converting at least one document into a form corresponding to the user request using a plurality of contents extracted from each of the plurality of models (S2640); and providing the document converted into a form corresponding to the user request to the user terminal as a response to the user request (S2650).

[0264] A document understanding system (1000) may receive a user request (or user input) containing at least one document from a user terminal (10). For example, as illustrated in FIG. 21, the document understanding system (1000) may activate the graphic object (2003) based on the selection of a graphic object (e.g., “convert / generate”, 2003) associated with a document conversion (or generation) request included in at least one page (2000) from the user terminal (10). Then, while the graphic object (2003) is activated, the document understanding system (1000) may receive a user request (2150) containing at least one document (e.g., “Document A_20251130”, 2100) based on the selection of a graphic object (e.g., “confirm”, 2150a) associated with a document and user request reception function from the user terminal (10). Here, the user request (2150) may further include user input (e.g., “file format”, 2101) for converting the form of the document (2100) to be converted into at least one specific form desired by the user among a plurality of forms. In this case, receiving a user request including at least one document may also be understood as receiving together the document and the user request (or user input, user selection, etc.) for converting the document (2100) into a form that meets the user’s needs (or is desired by the user). Additionally, the specific form corresponds to the user request and may mean at least one form different from the form of at least one document (for example, if the form of the original document is the first form, the form of the converted document is the second form).

[0265] Subsequently, the document understanding system (1000) can check the extension of the document (2100) and convert the document (2100) into a specific format that has been pre-set. Then, the document understanding system (1000) can identify multiple layout areas each containing multiple contents having different data characteristics through analysis of the document converted into the specific format, and can extract multiple contents included in each of the multiple layout areas by using multiple models specialized for each of the multiple contents. Since more specific details regarding this have been described above, they will be omitted below to avoid duplication of explanation.

[0266] Furthermore, the document understanding system (1000) can use multiple contents extracted from a document converted into a specific format to convert the document converted into a specific format into a form corresponding to a user request (2150), and as a response to the user request (2150), provide the document converted into a form corresponding to the user request to a user terminal (10). Here, the term "document converted to correspond to the user request (2150)" may include a document converted into a specific form corresponding to the user request.

[0267] More specifically, the document understanding system (1000) can understand the relationship between the flow of the document and the extracted multiple contents, and convert the form of the document (or the document converted into a specific format) into a specific form corresponding to the user request. Here, understanding the relationship between the flow of the document and the extracted multiple contents may include the process of identifying the logical connection structure between paragraphs in the document and analyzing the technical and explanatory purpose conveyed by each paragraph. In addition, it may include the process of identifying the interrelationships regarding how images, charts, tables, etc. included in the document complement or explain certain texts. Through this, the correspondence between the concepts, steps, and components presented by the text and the visual information provided by the images can be structurally understood. Furthermore, the logical consistency of the entire document can be ensured by automatically determining whether an image supports the explanation of a specific paragraph or assists the flow of the technical composition. Based on this analysis, text-image mapping, the reorganization of the information order, and the rearrangement of visual elements become possible, and distortion of the original text's meaning can be prevented during the document conversion and / or creation process.

[0268] Furthermore, converting into a form corresponding to (or corresponding to) a user request may mean converting the structure and / or format (or form, style) of a document based on (or criterion for) the content of the user's request. This may include a process of generating a document in a form that satisfies the user's intended purpose and requirements by reconstructing the document to correspond to the content and format of the user's request using multiple contents extracted from the document.

[0269] Alternatively, converting into a form that corresponds to a user request involves analyzing the user's request based on multiple contents extracted from the document and rearranging and / or converting the content in a manner corresponding to the conditions of the request. Through this, by satisfying the level of information provision and presentation format requested by the user, a document that satisfies the user's request can be provided.

[0270] In other words, “response” can be understood as organizing content (e.g., selecting, combining, rearranging, etc.) to correspond to the intent, purpose, requirements, and output format included in the user request, and “satisfaction” as generating a document that satisfies the scope of information and presentation method requested by the user.

[0271] For example, let us assume that a user request (2150) to convert a document converted to a specific format into a PPT is received. As illustrated in FIG. 22 (a), the document understanding system (1000) can convert the document converted to a specific format into a first form (e.g., PPT) corresponding to the user request (2150) and provide the converted document (2100a) to the user terminal (10).

[0272] Let us assume, as another example, that a user request (2150) is received to convert a document converted into a specific format into the form of a blog post. As illustrated in FIG. 22 (b), the document understanding system (1000) can convert the document converted into a specific format into a second form (e.g., a blog post) corresponding to the user request (2150) and provide the converted document (2100b) to the user terminal (10).

[0273] Let us assume, as another example, that a user request (2150) is received to convert a document converted into a specific format into a poster form. As illustrated in (c) of FIG. 22, the document understanding system (1000) can convert the document (2100) into a third form (e.g., a poster) corresponding to the user request (2150) and provide the converted document (2100c) to the user terminal (10).

[0274] Meanwhile, the document understanding system (1000) according to the present invention can convert the form of a document into a form desired by the user by reflecting style information. More specifically, the document understanding system (1000) can convert the form of a document into a form desired by the user by reflecting the user's preferred style information and provide it to the user.

[0275] For example, as illustrated in FIG. 23, the document understanding system (1000) can activate the graphic object (2003) based on the selection of a graphic object (e.g., “conversion / creation”, 2003) associated with a document conversion (or creation) request included in at least one page (2000) from the user terminal (10). Then, with the graphic object (2003) activated, the document understanding system (1000) can receive a user request (2350) including at least one document (e.g., “Document B_20251201”, 2300) based on the selection of a graphic object (2350a) associated with a document and user request receiving function from the user terminal (10). At this time, the user request (2350) may further include at least one of style information (2301) to be reflected in the document (2300) to be converted, paragraph information (2302), and user input (e.g., “file format”, 2303) for converting the document (2300) into a form desired by the user.

[0276] However, the style information (2301) and / or paragraph information (2302) may be information that is automatically reflected according to the context of the document entered by the user, or information that is automatically reflected based on (or derived from) user history information. In one embodiment, the user history information may include style information and / or paragraph information that the user frequently uses or prefers.

[0277] Subsequently, the document understanding system (1000) can check the extension of the document (2300) and convert the document (2300) into a specific format that has been pre-set. Then, the document understanding system (1000) can identify multiple layout areas each containing multiple contents having different data characteristics through analysis of the document converted into the specific format, and can extract multiple contents included in each of the multiple layout areas by using multiple models specialized for each of the multiple contents. Since more specific details regarding this have been described above, they will be omitted below to avoid duplication of explanation.

[0278] Furthermore, the document understanding system (1000) can use multiple contents extracted from a document converted into a specific format to convert the document converted into a specific format into a form corresponding to a user request (2350), and as a response to the user request (2350), provide the document converted into a form corresponding to the user request to a user terminal (10). At this time, the document understanding system (1000) can convert the form of the document converted into a specific format into a form corresponding to the user request by reflecting the style information based on the fact that the user request (2350) includes style information.

[0279] In this case, the document understanding system (1000) can determine the order information of each of the multiple contents extracted from a document converted into a specific format by using a content order sorting model (580). For example, as shown in FIG. 24, the content order sorting model (580) can determine the order information of each of the multiple contents extracted from a document (2300a) converted into a specific format and sort the contents according to the determined order information.

[0280] At this time, at least some of the multiple contents being aligned may correspond to content that has a style corresponding to user input reflected through the style module (575). More specifically, the document understanding system (1000) may reflect (or apply) a style corresponding to user input to at least some of the multiple contents extracted from a document converted to a specific format based on the inclusion of style information in the user request (2350). For example, the document understanding system (1000) may reflect at least one of the style information (2301) and paragraph information (2302) included in the user request (2350) to at least some of the multiple contents (e.g., text) extracted from a document converted to a specific format.

[0281] As seen above, the content ordering model (580) can perform the role of arranging multiple contents extracted within a document by specifying (or determining) a logical reading order (e.g., ROD) corresponding to the flow of a person actually reading the document. This can be understood as a process of reconstructing the information flow structure of the entire document, arranging the order by combining the location information and semantic location of each content, and enabling the multimodal model (590) to remap multiple contents to their respective correct locations.

[0282] In one embodiment, the content order sorting model (580) receives the output of the document layout analysis model (510) as input, analyzes the structural features of the layout included in the document converted to a specific format, and assigns a sequential index to each layout area (or each content). Through this, the content order sorting model (580) enables multiple contents included in the document converted to a specific format to be sorted (or arranged, placed, etc.) in a semantically consistent order, and can provide a standard that is referenced when the document comprehension unit (500) subsequently generates data of a specific format (e.g., “JSON”). That is, the content order sorting model (580) restores the structural context from the document and provides information that enables the subsequent multimodal model (590) to generate various content following the flow of the original document.

[0283] Subsequently, the multimodal model (590) can convert the form of a document converted into a specific format into a form corresponding to a user request based on at least one of the order information specified through the content order sorting model (580), the output of the document layout analysis model (510), and multiple contents extracted from each of the multiple models.

[0284] In this case, the document understanding system (1000) may configure a prompt including at least one of the order information specified through the content order sorting model (580), the output of the document layout analysis model (510), multiple contents extracted from each of the multiple models, and a user request, and input the prompt into the multimodal model (590). The multimodal model (590) may use (or use as a basis for) the input prompt to convert the form of the document converted into a specific format into a form corresponding to the user request.

[0285] Alternatively, the document understanding system (1000) may integrate the identification results of each of a plurality of layout areas identified through the document layout analysis model (510) with the order information of each of a plurality of contents specified through the content order sorting model (580), and input the integrated data into the multimodal model (590). The multimodal model (590) may use the input integrated data to convert the form of a document converted into a specific format into a form corresponding to a user request.

[0286] Alternatively, the document understanding system (1000) may input data in a specific format (e.g., JSON) into the multimodal model (590). The multimodal model (590) may use the input data in the specific format to convert the form of the document converted into the specific format into a form corresponding to the user request.

[0287] In another embodiment, the multimodal model (590) can i) analyze the meaning of each of the multiple contents to summarize the content of the document, ii) reconstruct sentences, or iii) extract key keywords. Additionally, the multimodal model (590) can recognize and connect reference relationships between text, images, charts, and tables among the extracted multiple contents, and perform structural and descriptive transformations necessary for an output form (e.g., file format) corresponding to a user request. Furthermore, the multimodal model (590) can generate natural narratives that match the context or style of the document, and, if necessary, perform additional semantic judgments for visual editing, such as changing the background or highlighting elements, through image-based context analysis.

[0288] Finally, the document understanding system (1000) can provide a document converted to correspond to a user request (2350) to a user terminal (10). Here, the term "document converted to correspond to a user request (2350)" may include a document that reflects at least one of the style information and paragraph information entered by the user and is converted into a specific form according to the user request (2350).

[0289] For example, let us assume that a user request (2350) to convert a document into PPT is received. As illustrated in FIG. 25 (a) to (d), the document understanding system (1000) may provide a converted document to a user terminal (10) that includes a plurality of slides (2401, 2402, 2403, 2404), reflecting at least one of the style information and paragraph information entered by the user. At this time, the plurality of slides (2401, 2402, 2403, 2404) may include slides generated by arranging a plurality of contents according to the order information specified from the content order sorting model (580) in the multimodal model (590).

[0290] As described above, the document understanding method and system according to the present invention can increase (or enhance) user convenience by recognizing the structure of various types of documents and the relationships between multiple contents (e.g., text, tables, graphs, charts, molecular structural formulas, etc.) included in the documents, thereby automatically extracting accurate information. Through this, users can minimize the time required to find necessary information in documents and reduce inconvenience.

[0291] Furthermore, according to the document understanding method and system of the present invention, multiple contents included in a document are each identified, and multiple contents are automatically extracted from the document by utilizing multiple models specialized for each of the multiple contents having different data characteristics. That is, the present invention utilizes multiple models to quickly and accurately extract necessary information even from a document containing various forms of content (or various forms of information). Through this, the user can receive the necessary information quickly and accurately.

[0292] Furthermore, according to the document understanding method and system of the present invention, by deeply analyzing the visual information of various forms of content included in a document, information omission can be prevented and the sophistication of the analysis can be enhanced. Through this, the present invention can accurately process a vast amount of documents, thereby improving the user's overall work or business processing speed.

[0293] Furthermore, according to the document understanding method and system of the present invention, by deeply analyzing various forms of content included in a document and providing a solution optimized for the user based on this, the efficiency of document processing can be maximized and the user's overall work efficiency can be improved.

[0294] Furthermore, according to the document understanding method and system of the present invention, it is possible to understand text, graphs, and tables contained in a document, while simultaneously interpreting complex molecular structural formulas and chemical reaction information. Through this, the present invention can enhance the research efficiency of chemists and support innovative discoveries in the field of chemistry. In other words, the present invention can satisfy the needs of the research field through a customized model that understands even information related to the chemical domain within the document.

[0295] Furthermore, according to the document understanding method and system of the present invention, molecular structural formulas can be accurately recognized in documents using a model specialized for molecular structural formulas, and a large-scale database of molecular structural formulas can be constructed based on this. Through this, researchers can utilize the database to search for chemical structural formulas and efficiently carry out large-scale analysis tasks required for research, thereby reducing the time required for research and / or development.

[0296] Furthermore, according to the document understanding method and system of the present invention, a plurality of contents extracted from a document using a plurality of models and output data generated based on said plurality of contents can be visualized and provided through a user interface. Through this, the user can intuitively recognize the information needed and understand it more quickly, thereby improving the efficiency of work and / or research.

[0297] Furthermore, according to the document understanding method and system of the present invention, various content can be recognized from various types of documents, converted and translated into an editable form, and provided to the user. That is, the present invention can process documents of various types, thereby dramatically expanding the scope of document processing. Additionally, by automating the document conversion or translation process, users do not need to perform complex manual editing or translation, thus improving the efficiency and accuracy of document-based work.

[0298] Furthermore, according to the document understanding method and system of the present invention, raw data of tables, charts, and molecular structural formulas included in a document can be extracted and converted into an editable form to be provided to a user. Through this, the present invention solves the problem that detailed editing is impossible when tables, charts, and molecular structural formulas are processed as images, thereby improving the usability of document processing and minimizing the time or cost required for a user to re-edit the document.

[0299] Meanwhile, the present invention described above can be implemented based on a quantum computer. The present invention implemented based on a quantum computer may include a qubit-based quantum processor and quantum memory, and may include software and hardware interfaces optimized for quantum computation.

[0300] Quantum processors in quantum computers utilize qubits to efficiently process complex operations through parallel computation, quantum entanglement, and quantum superposition, which cannot be performed by the binary bits of classical computers. Quantum processors process data using quantum gates and can provide exponential speed improvements for specific problems.

[0301] Meanwhile, the present invention described above can be implemented as a program that is executed by one or more processes on a computer and can be stored on a computer-readable medium (or recording medium).

[0302] Furthermore, the present invention described above can be implemented as computer-readable code or instructions on a medium on which a program is recorded. That is, the present invention can be provided in the form of a program.

[0303] Meanwhile, computer-readable media include all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0304] Furthermore, the computer-readable medium may be a server or cloud storage that includes a storage and is accessible to an electronic device via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage via wired or wireless communication.

[0305] A computer program may reach the system (100) through various suitable transmission mechanisms. The transmission mechanism may be, for example, a computer-readable storage medium, a computer program product, a memory device, a recording medium such as a CD-ROM or DVD, or a product that tangibly embodies the computer program. The transmission mechanism may be a signal configured to reliably transmit the computer program through air or an electrical connection. The system (100) may propagate or transmit the computer program as a computer data signal.

[0306] Furthermore, references to 'computer-readable storage media,' 'computer program products,' 'computer programs embodied in a tangible form,' etc., or to 'controller,' 'computer,' 'processor,' etc., should be understood to include not only computers with various architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures, but also specialized circuits such as Field-Programmable Gate Arrays (FPGAs), Application Specific Circuits (ASICs), signal processing units, and other devices. References to computer programs, instructions, code, etc., should be understood to include software for programmable processors or firmware, such as programmable content for hardware devices, whether it is instructions for a processor or configuration settings for a fixed-function device, gate array, or programmable logic device.

[0307] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, namely a CPU (Central Processing Unit), and no special limitations are placed on its type.

[0308] Meanwhile, the above detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.

Claims

1. Regarding methods performed by a computer, A step of receiving a user request including at least one document; A step of identifying at least one layout area each containing at least one content having different data characteristics through analysis of the above document; A step of specifying one or more models configured to perform processing for each of the at least one content based on the data characteristics of the content included in each of the at least one layout areas; A step of extracting at least one content included in each of the at least one layout areas using each of the above one or more models; A step of converting the document into a form corresponding to the user request using the at least one content extracted using the above one or more models; and A document understanding method characterized by including the step of providing a document converted into a form corresponding to the user request to a user terminal as a response to the above user request.

2. In Paragraph 1, The above user request is, A request to convert the form of at least one document to be converted into at least one specific form among a plurality of forms, and In the above conversion step, A document understanding method characterized by using at least one extracted content to convert the form of at least one document into a specific form corresponding to the user request.

3. In Paragraph 1, In the step of identifying the plurality of layout areas mentioned above, The above at least one document is processed as input to a Document Layout Analysis (DLA) model, and A document understanding method characterized by analyzing at least one document in the above document layout analysis model to identify a plurality of layout areas each containing at least one content.

4. In Paragraph 3, In the step of extracting each of the above-mentioned at least one content, A document understanding method characterized by extracting, in each of the plurality of models, at least one content included in each of the plurality of layout areas identified by the layout analysis model.

5. In Paragraph 4, The above document layout analysis model is, A document understanding method characterized by outputting at least one of a label corresponding to each of each of the at least one content included in the at least one document, location information for each of the at least one content, and a score for each of the at least one content.

6. In Paragraph 2, In the above conversion step, A document understanding method characterized by understanding the relationship between the flow of at least one document and at least one extracted content, and converting the form of at least one document into the specific form.

7. In Paragraph 6, A method for understanding a document, characterized in that the above specific form satisfies the above user request and includes at least one form different from the above at least one document form.

8. In Paragraph 2, The method further includes the step of specifying sequence information for at least one content extracted from at least one document, and When the above sequence information is specified, in the above conversion step, A document understanding method characterized by converting the form of at least one document into the specific form based on the specified order information.

9. In Paragraph 8, In the above conversion step, A document understanding method characterized by converting the form of at least one document into the specific form based on at least one of the above sequence information and the output of the document layout analysis (DLA) model.

10. In Paragraph 2, The above at least one content is, It includes at least one of text, table, chart, molecular structure, and chemical reaction information, and The above user request is, A document understanding method characterized by further including a request to reflect style information in at least one document that is the subject of conversion.

11. In Paragraph 10, If the above user request further includes the above style information, in the above conversion step, A document understanding method characterized by converting the form of at least one document into the specific form by reflecting the style information above.

12. In Paragraph 10, A document understanding method characterized by the style information being reflected in at least a part of at least one of the above-mentioned content based on the above-mentioned user request.

13. In Paragraph 1, When receiving a summary request for at least one document in response to the above user request, A document understanding method characterized by further including the step of summarizing at least one document using at least one content extracted from each of the plurality of models in response to the above user request.

14. In Paragraph 1, In the above extraction step, A document understanding method characterized by extracting raw data from at least some of the content among at least one of the content by using at least some of the plurality of models.

15. In Paragraph 7, A method for understanding a document characterized in that at least some of the above-mentioned content includes at least one of the above-mentioned table and the above-mentioned chart.

16. In Paragraph 1, The method further includes the step of checking the extension of at least one document and converting at least one document into a pre-set specific format. When at least one of the above documents is converted to a specific format, Through analysis of at least one document converted to the above specific format, the plurality of layout areas each containing at least one content are identified, and Each of the above plurality of models is used to extract each of the at least one content included in each of the above plurality of layout areas, and Using the at least one content extracted from each of the plurality of models, at least one document converted into the specific format is converted into a form corresponding to the user request, and A document understanding method characterized by providing a document converted into a form corresponding to the user request to the above-mentioned user terminal.

17. A system comprising memory configured to store executable instructions and one or more processors configured to perform operations by executing one or more instructions, The above system is, Receive a user request including at least one document, and Through analysis of the above document, at least one layout area containing at least one content having different data characteristics is identified, and Based on the data characteristics of the content included in each of the at least one layout area, one or more models configured to perform processing for each of the at least one content are specified, and Using each of the above one or more models, extract the at least one content included in each of the at least one layout areas, and Using the at least one content extracted using the above one or more models, the document is converted into a form corresponding to the user request, and A document understanding system characterized by providing a document converted into a form corresponding to the user request to a user terminal as a response to the above user request.

18. A program that is executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The above program is, A step of receiving a user request including at least one document; A step of identifying at least one layout area each containing at least one content having different data characteristics through analysis of the above document; A step of specifying one or more models configured to perform processing for each of the at least one content based on the data characteristics of the content included in each of the at least one layout areas; A step of extracting at least one content included in each of the at least one layout areas using each of the above one or more models; A step of converting the document into a form corresponding to the user request using the at least one content extracted using the above one or more models; and A program stored on a computer-readable recording medium characterized by including instructions that perform the step of providing a document converted into a form corresponding to the user request to a user terminal as a response to the user request.