Method and system for document understanding

WO2026177406A1PCT designated stage Publication Date: 2026-08-27LG MANAGEMENT DEV INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/001477
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2026-01-05
Filing Date
2026-01-26
Publication Date
2026-08-27

Smart Images

  • Figure KR2026001477_27082026_PF_FP_ABST
    Figure KR2026001477_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for document understanding, and provides a method and system for document understanding using a deep document understanding (DDU) technology.
Need to check novelty before this filing date? Find Prior Art

Description

Document Understanding Methods and Systems

[0001] The present invention relates to a document understanding method and system, and provides a document understanding method and system using Deep Document Understanding (DDU) technology.

[0002] Documents are recorded through characters or symbols (or codes) to preserve and transmit thoughts, ideas, intentions, and information; in a broad sense, they can include everything that contains meaning, such as pictures, photographs, and videos. This means that information is recorded and preserved in various forms and can serve as a means of communication.

[0003] In this regard, a large volume of electronic documents created via computers is being utilized in various fields in modern society. These documents contain not only text but also various forms of information, such as tables, images, and graphs. For example, documents used in industrial settings, such as papers, patents, technical reports, and manuals, include not only text but also diverse forms of information like tables, graphs, diagrams, images, mathematical formulas, and molecular structural formulas.

[0004] While humans can read and understand the various forms of information contained in such documents, computers cannot fully comprehend complex forms of information in the same way. Consequently, analyzing and utilizing document content requires the cumbersome task of converting it into a format recognizable by computers. However, as documents contain a mixture of diverse information, manually reviewing or analyzing them can consume a significant amount of time and cost.

[0005] To address this, various technologies have been proposed to understand and analyze diverse forms of information contained in documents and provide solutions optimized for various fields. However, conventional technologies are often optimized for specific domains or types of information (e.g., text-centric, table-centric, or graph-centric), which presents limitations in that they are difficult to apply to documents containing information from different domains or complex forms.

[0006] More specifically, technology that automatically detects, analyzes, and visualizes chemical molecular structural formulas and chemical reaction information contained within documents is essential for efficiently processing large-scale document data, increasing research efficiency, and reducing errors in various fields (e.g., chemistry, pharmaceuticals, life sciences, etc.).

[0007] However, conventional technologies related to this have limitations in terms of the visual intuitiveness of analysis results, data processing scope, user convenience, and the usability of analysis results. For example, because the analysis results do not visually match the original document, it is difficult for users to easily verify detection accuracy. Furthermore, there is a lack of editing functions to incorporate user feedback, and mechanisms to convey the reliability of model inference to the user are insufficient.

[0008] Consequently, there is still a need for methods to precisely analyze document data and intuitively visualize the analysis results, enabling users to practically utilize the findings.

[0009] The present invention is intended to provide a document understanding method and system capable of effectively understanding various forms of documents.

[0010] More specifically, the present invention aims to provide a document understanding method and system capable of efficiently processing documents in various fields and increasing research efficiency.

[0011] In particular, the present invention is intended to provide a document understanding method and system capable of automatically detecting and analyzing data related to the field of chemistry contained in a document and intuitively visualizing the analysis results.

[0012] Furthermore, the present invention aims to provide a document understanding method and system that can be universally utilized in various forms of documents and fields.

[0013] Furthermore, the present invention aims to provide a document understanding method and system capable of improving efficiency in various industrial or research fields and providing optimized solutions.

[0014] To solve the problem described above, a document understanding method according to the present invention, performed by a computer, may include the steps of: specifying at least one document to be analyzed; inputting the document to be analyzed into at least one module to detect at least one molecular structural formula from the document to be analyzed; inputting the document to be analyzed into at least one other module to detect at least one chemical reaction formula from the document to be analyzed; analyzing atoms constituting the at least one detected molecular structural formula and the bonding relationships between the atoms to derive an analysis result for the at least one detected molecular structural formula; analyzing the relationships between components constituting the at least one detected chemical reaction formula to derive an analysis result for the at least one detected chemical reaction formula; and providing the analysis result for the at least one detected molecular structural formula and the analysis result for the at least one detected chemical reaction formula to the user terminal.

[0015] In an embodiment, the analysis result of the at least one detected molecular structure formula and the analysis result of the at least one detected chemical reaction formula are provided through a service page output to the user terminal, and the service page may include at least one of a first area in which at least one graphic object corresponding to at least one page included in the document to be analyzed is provided, a second area in which the document to be analyzed is provided, and a third area in which at least one of the analysis result of the at least one detected molecular structure formula and the analysis result of the at least one detected chemical reaction formula is provided.

[0016] In an embodiment, so that it is possible to identify that the at least one molecular structural formula and the at least one chemical reaction formula have been detected from the document to be analyzed, the detection result of the at least one molecular structural formula and the detection result of the at least one chemical reaction formula may be displayed in the at least one area containing the at least one molecular structural formula and the at least one area containing the at least one chemical reaction formula of the document to be analyzed provided in the second area.

[0017] In an embodiment, the second region may be provided with at least one page included in the document to be analyzed, and the third region may be provided with an analysis result for at least one data related to a chemical domain detected from the at least one page provided in the second region.

[0018] In an embodiment, when at least one molecular structural formula is detected from at least one page, in order to identify that the at least one molecular structural formula has been detected from the at least one page, the detection result of the at least one molecular structural formula is displayed in at least one area containing the at least one molecular structural formula of the at least one page provided in the second area, and the analysis result for the at least one molecular structural formula detected from the at least one page may be provided in the third area.

[0019] In an embodiment, the analysis result for the at least one molecular structural formula includes at least one molecular structural formula corresponding to the layout of the at least one detected molecular structural formula, at least one molecular structural formula standardized based on the at least one detected molecular structural formula, and at least one chemical structural representation format corresponding to the at least one molecular structural formula, wherein the at least one chemical structural representation format may be obtained by analyzing the atoms constituting the at least one molecular structural formula and the bonding relationships between the atoms in at least one module, and converting the at least one molecular structural formula into the at least one chemical structural representation format.

[0020] In an embodiment, when a user input selecting one of the analysis results for at least one molecular structural formula is received, at least one interface may be provided that displays in detail the analysis result for the selected molecular structural formula according to the user input.

[0021] In an embodiment, when at least one chemical reaction formula is detected from at least one page, in order to identify that at least one chemical reaction formula has been detected from at least one page, at least one area containing at least one chemical reaction formula of at least one page provided in the second area may display a detection result of said at least one chemical reaction formula, and a third area may provide an analysis result for said at least one chemical reaction formula detected from at least one page.

[0022] In an embodiment, the detection result of the at least one chemical reaction equation includes the detection result of each of the components constituting the at least one chemical reaction equation, and the detection result of each of the components is displayed in the second area with different visual appearances, and the components may include at least one of reactants, reaction conditions, and products.

[0023] In an example, the analysis result for the at least one chemical reaction equation may include at least one chemical reaction information derived by analyzing the relationship between the components constituting the at least one chemical reaction equation.

[0024] In an embodiment, when a user input selecting one of the analysis results for at least one chemical reaction formula is received, at least one interface may be provided that displays in detail the analysis result for the selected chemical reaction formula according to the user input.

[0025] In the embodiment, the third region may display a reliability grade for the analysis result of the at least one molecular structural formula so as to enable identification of the quality of the analysis result of the at least one molecular structural formula.

[0026] In the embodiment, the third region may display a reliability grade for the analysis result of the at least one chemical reaction formula so as to identify the quality of the analysis result for the at least one chemical reaction formula.

[0027] In an embodiment, when a user request to edit an analysis result for the at least one molecular structural formula is received from the user terminal, the method further includes the step of providing an editing interface to the user terminal that provides an editing function for the analysis result for the at least one molecular structural formula in response to the user request, and when editing of the analysis result for the at least one molecular structural formula is performed through the editing interface, the edited analysis result may be stored in a specified storage.

[0028] In an embodiment, when a user request to edit an analysis result for the at least one chemical reaction formula is received from the user terminal, the method further includes the step of providing an editing interface to the user terminal that provides an editing function for the analysis result for the at least one chemical reaction formula in response to the user request, and when editing of the analysis result for the at least one chemical reaction formula is performed through the editing interface, the edited analysis result may be stored in a specified storage.

[0029] In an embodiment, the at least one module detects the at least one molecular structural formula from the document to be analyzed and outputs at least one of a label, location information, and score of the at least one detected molecular structural formula, and the at least one other module detects the at least one chemical reaction formula from the document to be analyzed and outputs at least one of a class of each component constituting the at least one detected chemical reaction formula and location information of each of said components, and said components may include at least one of reactants, reaction conditions, and products.

[0030] In an embodiment, the method further includes the step of converting the document to be analyzed into a pre-set specific format, and when the document to be analyzed is converted into the specific format, the document to be analyzed converted into the specific format is input into the at least one module to detect the at least one molecular structural formula, and the document to be analyzed converted into the specific format is input into the at least one other module to detect the at least one chemical reaction formula.

[0031] In an embodiment, the second area is provided with at least one page image corresponding to at least one page included in the document to be analyzed, and the at least one page image can be reconstructed with the same layout as the at least one page included in the document to be analyzed.

[0032] A document understanding system according to the present invention, comprising a memory configured to store executable instructions and one or more processors configured to perform operations by executing one or more instructions, can specify at least one document to be analyzed, input the document to be analyzed into at least one module to detect at least one molecular structural formula from the document to be analyzed, input the document to be analyzed into at least one other module to detect at least one chemical reaction formula from the document to be analyzed, analyze atoms constituting the at least one molecular structural formula and the bonding relationships between the atoms to derive an analysis result for the at least one molecular structural formula to be detected, analyze the relationships between components constituting the at least one chemical reaction formula to derive an analysis result for the at least one chemical reaction formula to be detected, and provide the analysis result for the at least one molecular structural formula to the user terminal.

[0033] A program according to the present invention is a program that is executed by one or more processes in an electronic device and can be stored on a computer-readable recording medium, and may include instructions for performing the steps of: specifying at least one document to be analyzed; inputting the document to be analyzed into at least one module to detect at least one molecular structural formula from the document to be analyzed; inputting the document to be analyzed into at least one other module to detect at least one chemical reaction formula from the document to be analyzed; analyzing atoms constituting the at least one molecular structural formula and the bonding relationships between the atoms to derive an analysis result for the at least one molecular structural formula to be detected; analyzing the relationships between components constituting the at least one chemical reaction formula to be detected to derive an analysis result for the at least one chemical reaction formula to be detected; and providing the analysis result for the at least one molecular structural formula to the user terminal and the analysis result for the at least one chemical reaction formula to be detected.

[0034] As described above, the document understanding method and system according to the present invention can automatically detect and analyze molecular structural formulas and chemical reaction information contained in a document, and visualize and provide this information to the user. Through this, the user can intuitively recognize the necessary information and understand it more quickly, thereby increasing the accuracy and efficiency of research. In other words, the user can receive the necessary information from the document quickly and accurately, thus reducing the time and cost required for research or development.

[0035] Furthermore, according to the document understanding method and system of the present invention, by providing the user with an image reconstructed with the same layout as the original document, the user is supported in easily comparing the original document with the analysis results and enhancing reliability. That is, when reconstructing the detection results of molecular structural formulas and chemical reaction information contained in a document, the present invention maintains (or preserves) the layout of the original document, thereby providing an environment in which the user can intuitively verify the analysis results.

[0036] Furthermore, according to the document understanding method and system of the present invention, an editing interface can be provided that allows for real-time modification of the analysis results of molecular structural formulas detected from a document and the analysis results of chemical reaction formulas detected from a document. That is, the present invention provides an intuitive editing environment based on the visual comparison of the analysis results and the original document, and can store data edited (or modified) by the user. Through this, the present invention enables the effective collection of user feedback and allows for continuous learning and performance optimization of the model.

[0037] Furthermore, according to the document understanding method and system of the present invention, by displaying and providing the reliability grade of the analysis results for molecular structural formulas detected from documents and the analysis results for chemical reaction formulas detected from documents, the user is supported in intuitively understanding the quality of each analysis result and making quick decisions.

[0038] Furthermore, according to the document understanding method and system of the present invention, by visually providing the analysis results of detected data in conjunction with the location of the original document, the user can easily verify the analysis results and efficiently modify them when necessary. In other words, the present invention can simultaneously improve the accuracy of analysis and productivity through a user-friendly interface.

[0039] Furthermore, the present invention can store edited analysis results, detected data, and location information within documents. This data can be utilized for document search and database construction, thereby providing high value in research and commercial applications.

[0040] FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented.

[0041] FIG. 2 illustrates an example of a block diagram of a computing device that may be included in a user computing device, a server computing system, and a training computing system, as an embodiment of a computing system in which the present invention can be implemented.

[0042] Figure 3 illustrates an example of a block diagram from another perspective of a computing device, which is one of the components of a computing system.

[0043] FIGS. 4, FIGS. 5, and FIGS. 6 are conceptual diagrams for explaining a document understanding system according to the present invention.

[0044] FIG. 7 is a flowchart illustrating a method for understanding documents according to the present invention.

[0045] FIGS. 8, FIGS. 9, FIGS. 10, FIGS. 11, FIGS. 12, FIGS. 13, FIGS. 14 and FIGS. 15 are conceptual diagrams for explaining a document understanding method according to the present invention.

[0046] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components are assigned the same reference number regardless of the drawing symbols, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not have distinct meanings or roles in themselves. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention.

[0047] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.

[0048] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0049] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0050] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0051] Meanwhile, FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented. In this regard, the document understanding system according to the present invention can be implemented through a computing device described below and can perform data processing related to the document understanding method described in this specification.

[0052] Referring to FIG. 1, a computing system (10000) that performs a method of automatically detecting and analyzing various data included in a document according to an embodiment of the present invention and intuitively visualizing the analysis results of the detected data may include at least one computing device. At this time, the at least one computing device may be a single processor or a multi-processor computing device.

[0053] The components of at least one computing device of the present invention may include various hardware components such as one or more processors, memory, other hardware, and a system bus (not shown) that connects various system components so that they can transmit and receive data to and from each other (e.g., telegraphically connected, physically connected, electrically connected), and the components of at least one computing device are not limited thereto and may be very diverse.

[0054] Meanwhile, at least one computing device included in a computing system (10000) that performs a method of automatically detecting and analyzing various data included in a document and intuitively visualizing the analysis results of the detected data may be connected to communicate via a network (1070). For example, at least one computing device included in the computing system (10000) may be clustered or may be part of a local area network (LAN). Additionally, at least one computing device may be part of a wide area network (WAN) or connected to at least one of a client-server network and a peer-to-peer network within the cloud.

[0055] Meanwhile, when at least one computing device is used in at least one of a network environment and a cloud computing environment, the at least one computing device may be connected to at least one of a public and a private network through a network interface or an adapter. In one embodiment, other communication connection devices, such as a modem, may be used to establish communication through the network. The modem may be at least one of an internal modem and an external modem, and may be connected to a system bus through a network interface or a specific mechanism, etc. A wireless network component consisting of an interface and an antenna may be coupled to the network through a device such as an access point, a peer computer, etc. In the present invention, the method of connecting at least one computing device to communicate through the network (1070) is not limited, and it may be connected to communicate in a manner different from the described example.

[0056] Furthermore, other computer-type devices and / or systems not shown in FIG. 1 may also interact technically with at least one computing device or other system through one or more connections to the network (1070) via a network interface. Here, the network interface may include network interface equipment such as a physical network interface controller (NIC) or a virtual network interface (VIF).

[0057] The network (1070) of the present invention may include various forms such as the Internet, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, Wireless USB (Wireless Universal Serial Bus), etc., and in the present invention, data transmission may be performed based on standard communication protocols such as TCP / IP, HTTP, SSL, etc.

[0058] A computing system (10000) that performs a method of automatically detecting and analyzing various data included in a document according to the present invention and intuitively visualizing the analysis results of the detected data may include at least one of a user computing device (1010), a training computing system (1050), and a server computing system (1030).

[0059] A user computing device (1010) according to the present invention may be understood as a computing device comprising at least one processor (1011) and at least one memory (1012) that perform a method of automatically detecting and analyzing various data included in a document and intuitively visualizing the analysis results of the detected data. For example, the user computing device (1010) may include at least one computing device among a smartphone, a smart TV, a laptop computer, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, and a head-mounted display).

[0060] At least one and / or at least one processor (1011) constituting the user computing device (1010) may include one or more general-purpose processors and / or one or more special-purpose processors. For example, at least one and / or at least one processor (1011) constituting the user computing device (1010) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), an application integrated circuit, an application semiconductor (ASIC), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.

[0061] Furthermore, at least one and / or at least one processor (1011) may be configured to execute computer-readable instructions contained in memory (1012) and / or other instructions described herein.

[0062] The memory (1012) constituting the user computing device (1010) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media and / or other types of physically durable storage media.

[0063] For example, the memory (1012) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and may include web storage of a server that performs the storage function of memory on the internet. This memory (1012) may store data and instructions necessary for the at least one and / or at least one processor (1011) to automatically detect and analyze various data contained in a document and to perform the operation of an application for intuitively visualizing the analysis results of the detected data.

[0064] A user computing device (1010) may include one or more user input components (1021) that detect user input. For example, the user input component (1021) may also be referred to as a user interface module. The user input component (1021) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of user input component (1021). In this case, the user input component (1021) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user. Meanwhile, the user of the present invention may refer to an automated agent, script, playback software, etc., that operates on behalf of one or more people.

[0065] A user can interact with a computing system (10000) including at least one computing device through input text, touch, voice, movement, computer vision, gestures and / or other forms of input / output using a user input component (1021). For example, the user input component (1021) may include one or more of a command line interface (CLI), a graphical user interface (GUI), a natural user interface (NUI), a voice command interface and / or other user interface (UI) representations.

[0066] Between the user input component (1021) and the user computing device (1010), one or more application programming interface (API) calls may be made based on user input received from the user interface and / or network.

[0067] Here, the expression "based on" may be interpreted to include cases where it is based on the use of a specific configuration, modified from, derived from, influenced by, dependent on, or otherwise derived from a specific configuration. In some embodiments, an API call may be configured for a specific API, which may be interpreted or converted into an API call configured for another API. Here, an API may refer to a defined interface or connection between computers or between computer programs.

[0068] In one embodiment, the user computing device (1010) may store at least one machine learning model (1020). For example, the user computing device (1010) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks), or other types of machine learning models including non-linear models and / or linear models, which perform a method of automatically detecting and analyzing various data contained in a document and intuitively visualizing the analysis results of the detected data, and may be composed of a combination thereof.

[0069] According to an embodiment of the present invention, a user computing device (1010) may use a local or / and external machine learning model (1020) to automatically detect and analyze various data contained in a document and to intuitively visualize the analysis results of the detected data. Alternatively, the user computing device (1010) may use a machine learning model (1040) provided by a server to automatically detect and analyze various data contained in a document and to intuitively visualize the analysis results of the detected data.

[0070] In addition, according to another embodiment of the present invention, a server computing system (1030) communicating with a user computing device (1010) may provide the analysis results of data detected from a document to the user computing device (1010) on an application or / and the web in accordance with a request from a user received through the user computing device (1010).

[0071] In addition, according to another embodiment of the present invention, by linking at least a part of a user computing device (1010) and a server computing system (1030) with each other to automatically detect and analyze various data contained in a document and intuitively visualize the analysis results of the detected data, the analysis results of the data detected from the document can be provided to the user.

[0072] Additionally, according to various embodiments of the present invention, a user computing device (1010) and / or a server computing system (1030) can learn a machine learning model (1020, 1040) that is performed in a method of automatically detecting and analyzing various data contained in a document and intuitively visualizing the analysis results of the detected data through interaction with a training computing system (1050) that is communicatedly connected via a network (1070). In this case, the training computing system (1050) may be a computing system separate from the server computing system (1030). Alternatively, in some embodiments, the training computing system (1050) may be part of the server computing system (1030) or part of the user computing device (1010).

[0073] Meanwhile, the server computing system (1030) may include at least one processor (1031) and memory (1032). Here, the processor (1031) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an application integrated circuit, an application semiconductor (ASIC), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions. For example, at least one processor (1031) may include a circuit and a transistor configured to execute instructions from memory (1032).

[0074] The memory (1032) constituting the server computing system (1030) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media, and / or other types of physically durable storage media. For example, the memory (1032) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs the storage function of memory over the internet. Additionally, the server computing system (1030) may further include a data storage (data store). For example, the data storage may be composed of at least one of a relational database, a NoSQL database, a data warehouse, and a local file system.

[0075] In the memory (1032) constituting the server computing system (1030) according to the present invention, data and instructions necessary for the operation of an application to automatically detect and analyze various data included in a document and intuitively visualize the analysis results of the detected data may be stored.

[0076] In one embodiment, the server computing system (1030) may be composed of a single device or a plurality of computing devices, and may be configured to operate according to a sequential or parallel computing architecture. Additionally, a distributed processing system may be configured with a plurality of networked devices.

[0077] Meanwhile, the training computing system (1050) may include at least one processor (1051) and memory (1052). The model trainer (1060) is a logical component that executes the training of at least one machine learning model (1020, 1040) and may be implemented in the form of hardware, firmware, or software. For example, the model trainer (1060) may be executed by the processor (1051) after loading training data (1061) stored in a storage device into memory (1052). For example, the model trainer (1060) may be configured to execute one or more operations (e.g., model training, model reconstruction, model validation, model testing) on ​​at least one machine learning model.

[0078] The machine learning model of the present invention may include at least one of a statistical model, an algorithm, a neural network (NN), a convolutional neural network (CNN), a generative neural network (GNN), a Word2Vec model, a Bag of Words model, a TF-IDF (document frequency-inverse document frequency) model, a GPT (Generative Pre-trained Transformer) model (or other autoregressive models), a PPO (Proximal Policy Optimization) model, a nearest neighbor model (e.g., a k-nearest neighbor model), a linear regression model, a K-means clustering model, a Q-learning model, a TD (Temporal Difference) model, a Deep Adversarial Network model, and all other types of models further described herein.

[0079] Specifically, the model trainer (1060) may execute operations to train a machine learning model, and said operations may include at least one of adding, removing, and modifying model parameters. At this time, the training of the machine learning model may be at least one of supervised learning, semi-supervised learning, and unsupervised learning. In one embodiment, the training of the machine learning model may include the step of repeatedly inputting training data (1061) based on epochs and repeatedly performing the machine learning model training process configured in this way. Here, an epoch may refer to a unit in which the entire set of training data (1061) undergoes forward and backpropagation processing once. In some implementations, different levels of training methods (e.g., supervised learning, semi-supervised learning, unsupervised learning) may be used for different epochs.

[0080] The training data (1061) of the present invention may include input data and / or data previously output from at least one machine learning model (e.g., recursive learning feedback).

[0081] At least one parameter of a machine learning model may include at least one of a seed value, a model node, a model layer, an algorithm, a function, connections between different machine learning models, connections between parameters, machine learning model constraints, and other digital components that influence the output of the machine learning model. In this case, model connections between different machine learning models may include or represent relationships between model parameters and / or models, which may be dependent or interdependent, hierarchical, and / or static or dynamic. The combinations and configurations of model parameters described herein may be too complex to be maintained or utilized by human cognitive abilities.

[0082] In the present invention, the machine learning parameters described according to the embodiments are not limited, and a single machine learning model may further include a plurality of model parameters.

[0083] Meanwhile, FIG. 2 illustrates an example of a block diagram of a computing device (1100) that may be included in a user computing device (1010), a server computing system (1030), and a training computing system (1050), as an embodiment of a computing system (10000) in which the present invention can be implemented.

[0084] As illustrated in FIG. 2, the computing device (1100) may include at least one application (e.g., Application 1 to Application N), and each of the at least one application may include a machine learning library and a model execution environment for performing methods to automatically detect and analyze various data contained in a machine learning-based document and to intuitively visualize the analysis results of the detected data. The at least one application included in the computing device (1100) may communicate with the sensor, context manager, device state manager, or additional component(s) within the computing device (1100) via an Application Programming Interface (API). In one embodiment, the at least one application may interface with device components, such as receiving sensor data or state data or transmitting prediction results to an output device via a public or private API.

[0085] Meanwhile, FIG. 3 illustrates an example of a block diagram in another aspect of a computing device (1200), which is one of the components of a computing system (10000) that automatically detects and analyzes various data included in a document according to an embodiment of the present invention and intuitively visualizes the analysis results of the detected data.

[0086] A computing device (1200) according to the present invention may include at least one application (e.g., Application 1 to Application N), and at least one application may communicate with a central intelligence layer (1210). Each application may interact with a shared model within the central intelligence layer (1210) through an API (e.g., a common API).

[0087] The central intelligence layer (1210) includes one or more machine learning models and may share them among multiple applications or provide them independently to each. In one embodiment, the central intelligence layer (1210) may be integrated as part of an operating system or implemented as a separate logical layer.

[0088] Additionally, the central intelligence layer (1210) can communicate with the central device data layer (1220). The central device data layer (1220) can store at least one and / or at least one document stored within the computing device (1200), automatically detect and analyze various data contained in the document, and provide it as input data necessary to intuitively visualize the analysis results of the detected data. Each device component (e.g., sensor, state manager, etc.) can communicate with the central device data layer (1220) through a private API, etc.

[0089] The technology described in this specification may be composed of a single or multiple computing devices, and a machine learning model that performs a method of automatically detecting and analyzing various data contained in a document and intuitively visualizing the analysis results of the detected data may be executed sequentially or in parallel on a single component or multiple distributed components. The data storage, machine learning model, and application may be distributed and operated locally or over a network, and these configurations can be flexibly applied to various system architectures.

[0090] Meanwhile, the present invention relates to a document understanding method and system capable of effectively understanding various types of documents. The document understanding system according to the present invention may be a system that provides a service and / or function for automatically converting multimodal information (e.g., layout, text, table, chart, image, graph, chemical molecular structural formula, chemical reaction formula, etc.) from various types of documents into data using Deep Document Understanding (DDU) technology.

[0091] The present invention aims (or has the purpose) to efficiently process documents in various fields and increase research efficiency. Below, we will examine the document understanding system according to the present invention in more detail together with the attached drawings. Figures 4, 5, and 6 are conceptual diagrams for explaining the document understanding system according to the present invention.

[0092] Meanwhile, as illustrated in FIG. 4, the document understanding system (1000) according to the present invention may include at least one of an input unit (100), an output unit (200), a communication unit (300), a storage unit (400), a document understanding unit (500), and a control unit (600). However, the components of the document understanding system (1000) according to the present invention are not limited thereto and may further include various hardware components that perform the same or similar roles as described in the description of the present specification.

[0093] Although not illustrated, the document understanding system (1000) according to the present invention may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processor, tensor processing unit (TPU), graphics processing unit (GPU), neural network processing unit (NPU), application integrated circuit, application semiconductor (ASIC), field programmable gate array (FPGA), quantum processing unit (or quantum processor, QPU), etc.). One or more processors may be configured to execute instructions, computer-readable instructions, and / or other instructions described herein that are stored (or included) in the storage unit (400). The document understanding method and system according to the present invention may perform data processing described below in cooperation with memory and at least one processor. The processor may perform a series of operations and data processing using data and information stored in memory. In this case, memory may be a component of the storage unit (400).

[0094] In addition, the document understanding system (1000) according to the present invention can perform data processing and computation processes using quantum gates, quantum entanglement, and quantum superposition states, taking into consideration implementation in a quantum computer environment. For example, the present invention can perform parallel computations based on qubits, and such quantum computations can operate complementarily with existing classical computers.

[0095] Such quantum computers may include parallel computation using qubits and high-speed data processing devices utilizing quantum entanglement, and hardware-based computational optimization using FPGAs and ASICs is possible. In addition, quantum computers may utilize quantum processors capable of qubit-based parallel computation, and data processing efficiency can be improved through a hybrid structure with existing classical computers.

[0096] Meanwhile, the input unit (100) can be configured in various ways as a means of data input. For example, the input unit (100) can be configured to receive user input. The input unit (100) can be configured to receive user input from a user terminal (10). Here, “receiving input” may mean receiving an input signal (or selection signal) corresponding to the user’s input based on input made by the user through the input unit configuration provided in the user terminal (10).

[0097] Here, the user terminal (10) may include at least one of a mobile phone, a smartphone, a notebook computer, a laptop computer, a slate PC, a tablet PC, an ultrabook, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, and a wearable device (e.g., a smartwatch, a smart glass, a head-mounted display).

[0098] In addition, the input unit (100) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user.

[0099] The input unit (100) may also be referred to as a user interface module. The input unit (100) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of input unit (100).

[0100] Here, user input may include documents, text, images (or videos), voice, etc. In this case, the document understanding system (1000) may further include a module that converts voice into text.

[0101] Next, the output unit (200) can output information through an output unit configuration (e.g., a display unit, a touch screen, a speaker, etc.) provided in a user terminal (10) linked to the document understanding system (1000) according to the present invention. For example, the output unit (200) can output at least one page (2000, or service page) linked to the document understanding system (1000) according to the present invention to the display unit of the user terminal (10). Additionally, the output unit (200) does not necessarily mean a hardware means, but can be understood as a channel for outputting results to a user.

[0102] Next, the communication unit (300) may be connected via a wireless or wired network to a user terminal (10), a server (e.g., a central server, an external server, etc.), a device, and at least one network, etc., to receive or transmit overall data and information necessary for the operation of the document understanding system (1000) according to the present invention.

[0103] The communication unit (300) can support various communication methods depending on the communication standard of the communicating device.

[0104] For example, the communication unit (300) may be configured to communicate with a communication target using at least one of the following technologies: WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus).

[0105] Next, the storage unit (400, or memory) serves to store various data related to the present invention and may include one or more non-transient computer-readable storage media that can be read and / or accessed by at least one of one or more processors.

[0106] One or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the storage unit (400) may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage device), whereas in other examples, the storage unit (400) may be implemented using two or more physical devices.

[0107] The storage unit (400) may include computer-readable instructions and additional data. The storage unit (400) may include a storage necessary to perform at least some of the methods, scenarios, and techniques described herein and / or at least some of the functions of the device and network.

[0108] Furthermore, at least a portion of the storage unit (400) may be a cloud storage or a cloud server. At least a portion of the data corresponding to user input received from the input unit (100) and the training data may be stored in the storage unit (400).

[0109] Additionally, the storage unit (400) may store at least one document collected (or received) from various sources (e.g., a web corpus, a document corpus, a database (DB) website, an API, a server linked to the document understanding system (1000), a central server, an external server, cloud storage, a user terminal (10), a large dataset, etc.). For example, the storage unit (400) may store at least one and / or at least one document collected from at least one of the various sources (e.g., a user terminal (10)). In this case, the collected document may include documents related to at least one and / or at least one domain (or field). Alternatively, the collected document may include documents containing at least one and / or at least one content.

[0110] That is, the storage unit (400) is sufficient as a space where information necessary for the operation of the document understanding system (1000) according to the present invention is stored, and it can be understood that there are no restrictions on the physical space.

[0111] Furthermore, the storage unit (400) can store a computer program including computer program instructions. Furthermore, the storage unit (400) can store a computer program including computer program instructions that control the operation of the system (1000) or control the operation of the control unit (600) when loaded into the processor of the system (1000).

[0112] Unlike general documents, chemical documents often contain a mixture of various forms of chemical domain data, such as molecular structural formulas, chemical equations, reaction conditions, and experimental result tables, in addition to text. Accurate detection and interpretation of this chemical domain data are difficult using standard document analysis techniques, and its layout within the document is frequently unstructured. Consequently, there is a growing demand for document understanding technologies specialized for chemical documents.

[0113] Accordingly, the document understanding unit (500) can be configured to effectively understand complex forms of information contained in various types of documents. The document understanding unit (500) can be configured to recognize the structure of the document to be analyzed (20) and to recognize the relationships of the content (e.g., Image-Text, Image-Image, Text-Text) contained in the document to be analyzed (20) to perform the role of extracting information. In the present invention, the document understanding unit (500) may also be named a “deep document understanding model,” a “document understanding model,” or a “DDU model.”

[0114] The document comprehension unit (500) may be configured to extract various forms of content (e.g., layout, text, table, chart, image, graph, chemical molecular structural formula, chemical reaction formula, etc.) from at least one and / or at least one document (e.g., paper, book, patent document, report, etc.). Here, the various forms of content may also be understood as multimodal information.

[0115] More specifically, the document comprehension unit (500) may be a model trained to understand structured data, unstructured data, linguistic data (or linguistic elements) and non-linguistic data (or non-linguistic elements), etc., included in the document (20) to be analyzed, and to extract various content and / or knowledge based on the understood content.

[0116] In one embodiment, the document understanding unit (500) can understand the chemical structure of a molecular structure formula included in the document to be analyzed (20), and based on the result of understanding, convert the molecular structure formula into a SMILES string expression and extract it. Additionally, the document understanding unit (500) can understand the chemical structure of the molecular structure formula and perform a graph conversion corresponding to the molecular structure formula based on the result of understanding.

[0117] In another embodiment, the document understanding unit (500) can understand texts related to molecular structural formulas among the texts included in the document to be analyzed (20) and extract them as text data related to said molecular structural formulas.

[0118] In another embodiment, the document comprehension unit (500) can recognize rows and columns constituting a table associated with molecular structural formulas from the document to be analyzed (20), and convert them into structured data in a format such as HTML or Excel to extract them.

[0119] Additionally, the document understanding unit (500) can extract relationship information (or relationships) between molecular structures included in the document to be analyzed (20).

[0120] In one embodiment, the document understanding unit (500) can understand the relationship between the first molecular structure and the second molecular structure included in the document to be analyzed (20), and extract relationship information in which a third molecular structure is generated through a chemical reaction between the first molecular structure and the second molecular structure. In this case, the relationship information between the molecular structures can be extracted by understanding the text included in the document to be analyzed (20) or by understanding the non-verbal data included in the document to be analyzed (20).

[0121] In another embodiment, the document understanding unit (500) can understand the relationship between the first molecular structure and the second molecular structure through a symbol (e.g., plus sign, arrow, etc.) located in one region among a plurality of regions included in the document (20) to be analyzed, and can extract relationship information in which a third molecular structure is generated through a chemical reaction between the first molecular structure and the second molecular structure.

[0122] Furthermore, the document comprehension unit (500) can extract various forms of content satisfying established content criteria from at least one and / or at least one document. Here, the established content criteria can be set in various ways and can be determined according to the purpose or use of the document comprehension system (1000). For example, if the purpose of use of the document comprehension system (1000) is chemistry, bio, new materials, new substances, and new drug development, the document comprehension unit (500) can be trained to understand and extract content related to chemistry, bio, new materials, new substances, and new drug development from the document to be analyzed (20). In this case, the established content criteria may include content related to molecular structures related to at least one of chemistry, bio, new materials, new substances, and new drug development. Here, the document comprehension unit (500) can extract content related to chemistry, bio, new materials, new substances, and new drug development from the document to be analyzed (20) according to the established content criteria. However, this is merely one embodiment, and the established content standards in the present invention are not necessarily limited thereto.

[0123] As described above, the document comprehension unit (500) can convert various forms of content (or information, data, etc.) contained in the document to be analyzed (20) into data that can be understood by a machine and / or artificial intelligence model. The data extracted using the document comprehension unit (500) can be sorted by page or document and stored in the storage unit (400, or memory). In this case, the document comprehension unit (500) can be configured (or constructed) to include at least one and / or at least one artificial intelligence model (or module).

[0124] Here, artificial intelligence models may include various models that can be utilized depending on various situations or purposes. For example, an artificial intelligence model may include at least one of a machine learning (ML) model, a deep learning model, a deep neural network (DNN), a language model (LM), a large language model (LLM), a massive foundation model, a generative AI model, a transformer-based model, a supervised learning (SL) model, a reinforcement learning (RL) model, a vision-language model (VLM), and a special-purpose model (e.g., a time series forecasting model (e.g., ARIMA model, SARIMA model, etc.), a time series foundation model, a graph neural network (GNN), a multimodal model, a natural language processing (NLP) model, a computer vision model, a speech recognition / synthesis model, a recommendation system model, etc.).

[0125] In this regard, as illustrated in FIGS. 5 and 6, when a document understanding system (1000) receives at least one document to be analyzed (20), it can process the received document to be analyzed (20) as input to a document understanding unit (e.g., “Manager API”, 500). The document understanding unit (500) can check the extension (e.g., PDF, PPTX, HWPX, PNG, JPG, etc.) of the input document to be analyzed (20, or at least one page included in the document to be analyzed) and convert the document to be analyzed (20, or at least one page included in the document to be analyzed) into a specific format set in advance. For example, the specific format set in advance (or format, form, etc.) may include an image.

[0126] Subsequently, the document understanding system (1000) can process the document to be analyzed converted into a specific format (i.e., the document to be analyzed converted into an image, 21) as input to the document understanding unit (500). In this case, the document to be analyzed converted into a specific format (21) can be distributed (or input) to each module (or each application programming interface (API)).

[0127] More specifically, the document comprehension unit (500) can process the document to be analyzed (21) converted into a specific format as an input to each of the plurality of modules (510, 520). Here, “processing the document to be analyzed (21) converted into a specific format as an input to each of the plurality of modules (510, 520)” can also be understood as inputting the document to be analyzed (21) converted into a specific format by calling each of the plurality of application programming interfaces (APIs). In the present invention, the application programming interface may also be implemented as an “agent” or an “artificial intelligence agent.”

[0128] The document comprehension unit (500) can call a first application programming interface (e.g., “DLA API”, 511) to detect (or extract) at least one and / or at least one molecular structural formula (or molecular structure) from an analysis target document (21) converted into a specific format. The first application programming interface (511) can input the analysis target document (21) converted into a specific format into a first module (510) to detect at least one molecular structural formula included in the analysis target document (21) converted into a specific format. In the present invention, the first application programming interface may also be expressed as a “first agent” or a “first artificial intelligence agent,” etc.

[0129] In the present invention, the term “detecting a molecular structural formula” may be understood to encompass both the process of detecting (or identifying) a molecular structural formula from a document or image and the process of converting the detected molecular structural formula into structural data (e.g., chemical structure representation format). Alternatively, the “process of detecting a molecular structural formula” and the “process of converting the detected molecular structural formula into structural data” may each be implemented as separate processes. The present invention is not limited to either of these.

[0130] In the document understanding method according to the present invention, a molecular structural formula can be automatically detected from a document to be analyzed, and an analysis result of the molecular structural formula can be derived by analyzing the atoms constituting the detected molecular structural formula and the bonding relationships between atoms. The analysis result may include at least one of a form maintaining the layout of the molecular structural formula, a standardized molecular structural formula form, or a form converted into a specific chemical structure representation format. Such molecular structural formula analysis provides foundational information for quantitatively and structurally understanding compound information contained in chemical documents.

[0131] The above chemical structure representation format may include, for example, SMILES, InChI, Molfile, or similar structural chemical representation formats, but is not limited thereto.

[0132] A first module (510, or a first application programming interface (511)) can detect at least one of a molecular structural formula and table data from an analysis target document (21) converted into a specific format and output (or return) a detection result of at least one of the molecular structural formula and table data. For example, the detection result output from the first module (510) may include at least one of a label of the detected data (e.g., “Label - mol”), location information (e.g., “Boxes: [x1, y1, x2, y2]”), and a score (e.g., “Score: 0.9”). Here, the label is a value indicating the type of what the detected data (or object) is, and the location information may mean (or include) the location coordinates of the detected data in the document. Additionally, the score may be a probability (or confidence) value regarding how reliable the detection result of the data is.

[0133] Subsequently, the document comprehension unit (500) may call a second application programming interface (e.g., “Reaction API”, 521) to detect (or extract) at least one and / or at least one chemical reaction formula from the document to be analyzed (21) converted into a specific format. The second application programming interface (521) may input the document to be analyzed (21) converted into a specific format into the second module (520) to detect at least one chemical reaction formula included in the document to be analyzed (21) converted into a specific format. In the present invention, the second application programming interface (521) may also be expressed as a “second agent” or a “second artificial intelligence agent,” etc.

[0134] In the present invention, the phrase “detecting a chemical reaction formula” can also be understood as detecting chemical reaction information from a chemical reaction formula included in a document or image.

[0135] Alternatively, in the present invention, the “process of detecting a chemical reaction equation” may also be understood as a process of recognizing a chemical reaction diagram that schematically represents a chemical reaction. For example, the process of recognizing a chemical reaction diagram may refer to a process of extracting reactants, reaction conditions, and products from the chemical reaction diagram.

[0136] Alternatively, the process of “detecting chemical reaction information” in the present invention may also be understood as a process of recognizing chemical reaction information from a document or image. For example, chemical reaction information recognition may refer to a process of recognizing chemical reaction information from a document or image and structuring it into reactants, reaction conditions, and products.

[0137] The detection of the above molecular structural formula or chemical reaction formula may be performed on the entire structure included in the document, or may be performed on only a part of the document or some components.

[0138] The second module (520, or the second application programming interface (521)) can detect a chemical reaction equation from an analysis target document (21) converted into a specific format and output a detection result of the chemical reaction equation. For example, the detection result output from the second module (520) may include components constituting the detected chemical reaction equation (i.e., reactants, reaction conditions, products, etc.) and location information for each of the components (e.g., “[x1, y1, x2, y2]”). Here, the location information of the reactants refers to an area within the document where the starting material (reactant) participating in the reaction is located, and the location information of the reaction conditions may refer to the location of text and / or areas where the conditions under which the chemical reaction occurs (reagent, catalyst, solvent, temperature, time, etc.) are indicated. Additionally, the location information of the products may refer to an area within the document where a molecular structural formula corresponding to the product (product) generated as a result of the reaction is located. That is, the second module (520) can extract the separated reaction structure information of the components and identify and output the location (or location information) of the components.

[0139] In addition, the document understanding method according to the present invention can automatically detect a chemical reaction equation from a document to be analyzed and analyze the relationships between the components constituting the chemical reaction equation. For example, a chemical reaction equation may include at least one component among reactants, reaction conditions, and products, and each component may be distinguished by a different class or visual appearance. Through the analysis of the relationships between these components, chemical reaction information including reaction pathways, reaction conditions, or product information can be derived.

[0140] Furthermore, the document comprehension unit (500) can generate a chemical structure representation format corresponding to (or for) the detected molecular structure formula. Additionally, the document comprehension unit (500) can generate a chemical structure representation format corresponding to at least one molecular structure formula included in the detected chemical reaction formula (e.g., a molecular structure formula corresponding to a product).

[0141] In one embodiment, the document comprehension unit (500) may call a third application programming interface (e.g., “Analyze API”, 531) to convert the detected molecular structural formula into a chemical structural representation format. The third application programming interface (531) may input the detected molecular structural formula into a third module (530) to convert the detected molecular structural formula into a chemical structural representation format (e.g., “SMILES”).

[0142] In another embodiment, the document comprehension unit (500) may call the third application programming interface (531) to convert at least one molecular structural formula included in the detected chemical reaction equation into a chemical structural representation format. The third application programming interface (531) may input at least one molecular structural formula included in the detected chemical reaction equation into the third module (530) to convert the at least one molecular structural formula into a chemical structural representation format (e.g., “SMILES”).

[0143] Finally, the document comprehension unit (500) can perform data processing on at least one of the detection results of the data detected from the document (detection results of molecular structural formulas and detection results of chemical reaction formulas) and the analysis results of the detected data (analysis results of molecular structural formulas and analysis results of chemical reaction formulas) to generate (or output) structured data (e.g., “Output”, 22) related to the document to be analyzed (20, or data included in the document to be analyzed (20)). For example, the structured data (22) may be data structured in at least one specific form (e.g., JSON). In the present invention, the structured data (22) may also be named as “structured data format,” “structured document comprehension result,” “structured data of a specific form (or format, etc.),” “structured result data,” or “structured output data.”

[0144] Here, “data processing” may include multimodal data processing for data with different characteristics (or content with different data characteristics). Multimodal may refer to a technology in which artificial intelligence (AI) simultaneously understands and processes various forms of data (modalities), such as text, images, voice, and video.

[0145] “Performing multimodal data processing on data having different characteristics to generate structured data (22)” may mean structuring data of different modalities or formats (e.g., molecular structural formulas, tables, chemical reaction formulas, etc.) detected from the document (20) to be analyzed.

[0146] This multimodal data processing involves integrating the processing of data corresponding to different modalities, which can also be understood as “performing multimodal data integration processing for data with different data characteristics.”

[0147] More specifically, it may mean to comprehensively analyze each content detected in the document (20) to be analyzed by considering semantic associations, spatial arrangements, logical relationships, etc., and based on this, include (or reflect) the meaning, structure (e.g., title, body text, table, caption, relationships between each content, etc.) and relationships between components (e.g., atoms constituting a molecular structural formula and bonding relationships between said atoms, relationships between components constituting a chemical reaction formula, etc.) of the document (20) to be analyzed, and to organize (or reconstruct) it into a structured data form (or representation) that can be interpreted by a machine (or computer).

[0148] As an example, the “multimodal data integration processing process for at least one of the detection result of data detected from the document to be analyzed (20) and the analysis result of the detected data” may include integratively understanding information between different modalities and interpreting the meaning, structure, and layout of the document to be analyzed (20) to generate results such as a document image (or page image), molecular structural formula, and chemical reaction formula composed of the same layout as the document to be analyzed (20). At this time, the multimodal data integration processing process may include a process of combining the recognition results of different modalities into a single integrated expression by mutually merging (or aligning) them based on positional relationships, semantic associations, and document context, rather than simply merging them. In this case, it may be performed by including mapping text and images (e.g., molecular structural formula, chemical reaction formula, etc.) corresponding to the same area, establishing correspondence relationships between table and chart images and the corresponding raw data, and analyzing semantic connection relationships between the body text, title, comment, caption, figure, and formula.

[0149] Alternatively, the “multimodal data integration processing process for data having different characteristics” may include a processing process that processes (or interprets) data having different characteristics through a pre-set processing technique (or method) and generates structured data (22) by reflecting the semantic and structural relationships between each data.

[0150] Here, the pre-configured processing technique may include a multimodal data integration processing technique and / or a multimodal analysis technique. In one embodiment, the pre-configured processing technique may include at least one of i) a technique for normalizing features by vectorizing or embedding extracted information and mapping them to a common representation space, ii) a technique for matching and / or inferring relationships between contents based on the spatial arrangement or semantic similarity of the contents, iii) a technique for fusing features of different modalities in an Early fusion and / or Late fusion and / or Hybrid fusion manner, and v) a technique for aligning semantic correspondence relationships between contents having different data characteristics.

[0151] The structured data (22) generated based on this may include a result in which each component within the document is expressed in a structurally defined data form along with information on roles, attributes, and interrelationships, by reflecting both the visual composition and semantic content of the document (20) to be analyzed. That is, it may include a result generated by understanding and reconstructing the various contents constituting the document (20) to be analyzed based on meaning and structure.

[0152] For example, the structured data (22) may be a data format that structures and expresses the detection and / or analysis results of data related to the chemical domain within the document (20) to be analyzed. In this case, the structured data (22) may include unique identifiers, labels, location information indicating the location within the document (20), and score information indicating the detection reliability for the molecular structural formulas and chemical reaction formulas detected within the document (20). At this time, for the detected molecular structural formulas, a chemical structural representation format that expresses the molecular structural formula in a form understandable by a machine (or computer) may be mapped and included. Additionally, for chemical reaction formulas, the reactants, reaction conditions, and products constituting the chemical reaction formula may be defined as separate objects, and information on the reaction relationships between them may be included. This structured data (22) may be configured to simultaneously preserve (or maintain) the layout and semantic information of the original document and to be utilized in various subsequent processes such as data visualization, editing, searching, databases, and training data.

[0153] Next, the control unit (600) can perform the role of controlling the overall operation of the document understanding system (1000) related to the present invention. The control unit (600) can process signals, data, information, etc. that are input or output through the components of the document understanding system (1000) described above, or perform a series of data processing to provide or process appropriate information and functions to the user. The control unit (600) can be physically implemented by the processor described above.

[0154] Meanwhile, as described above, the present invention provides a document understanding system (1000) that automatically converts multimodal information from various types of documents into data using Deep Document Understanding (DDU) technology. In particular, the document understanding system (1000) according to the present invention can automatically detect and analyze data related to the chemical field contained in the document and intuitively visualize the analysis results. In this regard, the document understanding method performed by the document understanding system (1000) will be examined in more detail below in conjunction with the attached drawings. Figure 7 is a flowchart for explaining the document understanding method according to the present invention, and Figures 8, 9, 10, 11, 12, 13, 14, and 15 are conceptual diagrams for explaining the document understanding method according to the present invention.

[0155] Meanwhile, as illustrated in FIG. 7, the document understanding method according to the present invention may include the steps of: specifying at least one document to be analyzed (S710); inputting the document to be analyzed into at least one module to detect at least one molecular structural formula from the document to be analyzed (S720); inputting the document to be analyzed into at least one other module to detect at least one chemical reaction formula from the document to be analyzed (S730); analyzing atoms constituting the detected at least one molecular structural formula and the bonding relationship between the atoms to derive an analysis result for the detected at least one molecular structural formula (S740); analyzing the relationship between components constituting the detected at least one chemical reaction formula to derive an analysis result for the detected at least one chemical reaction formula (S750); and providing the analysis result for the detected at least one molecular structural formula and the analysis result for the detected at least one chemical reaction formula to a user terminal (S760).

[0156] The document understanding system (1000) can identify at least one document to be analyzed.

[0157] In the present invention, there may be various methods (or methods or criteria) for identifying the document to be analyzed.

[0158] For example, as illustrated in FIG. 8, the document understanding system (1000) can receive a document (800) corresponding to user input based on at least one document (800) being input into a document input area included in a service page (2000) from a user terminal (10). In this case, the document understanding system (1000) can identify the received document (800) as a document to be analyzed.

[0159] In another example, the document understanding system (1000) can identify the input document as the document to be analyzed (800) based on the input of at least one document corresponding to the user's selection among the documents stored (or embedded) in the storage (or memory or storage space or database) of the user terminal (10) to the document upload page (or interface) provided on the service page (2000).

[0160] As another example, the document understanding system (1000) may receive link information of a document (e.g., a URL) or link information of an external storage service storing the document from a user terminal (10). Then, the document understanding system (1000) may directly access the document or download the document through the link information of the document to identify the document to be analyzed (800).

[0161] However, the method of identifying the document to be analyzed in the present invention is not necessarily limited to the embodiments described above. For the convenience of explanation, the following description will be made on the premise that the document received through the user terminal (10) on which the service page (2000) is displayed is identified as the document to be analyzed (800).

[0162] Additionally, the document understanding system (1000) can process a specific document to be analyzed (800) as input to the document understanding unit (500). The document understanding unit (500) can check the extension of the document to be analyzed (e.g., PDF, PPTX, HWPX, PNG, JPG, etc.) and convert the document to be analyzed (800) into a specific format that has been set. For example, the specific format that has been set (or format, form, etc.) may include an image, and this may involve converting the document to be analyzed (800) into a specific format to generate (or obtain) a document to be analyzed of a specific format (or a document to be analyzed corresponding to a specific format, a document to be analyzed having a specific format, an image of a document to be analyzed, etc.). Alternatively, it may involve converting at least one and / or at least one or more pages (multiple pages) included in (or constituting) the document to be analyzed (800) into an image. In this case, it can be understood that at least one page image corresponding to at least one page included in the document (800) to be analyzed is generated.

[0163] Subsequently, the document comprehension unit (500) can detect (or extract) at least one and / or at least one data (e.g., molecular structural formula, chemical reaction formula, etc.) related to the chemical domain included in the document to be analyzed (800). In the present invention, the phrase “detected from the document to be analyzed (800)” can be interpreted to mean that it was detected from the document to be analyzed converted into a specific format.

[0164] The document comprehension unit (500) can input the document to be analyzed, converted into a specific format, into a module specialized in detecting molecular structure formulas and a module specialized in detecting chemical reaction formulas, respectively, in order to detect at least one molecular structure formula and at least one chemical reaction formula included in the document to be analyzed (800).

[0165] Specifically, the document comprehension unit (500) may input the document to be analyzed, converted into a specific format, into at least one module in order to detect a molecular structural formula from the document to be analyzed (800). Here, the at least one module may include a first module (510) specialized for detecting molecular structural formulas. The first module (510) may be at least one module, and may also be expressed as “at least one module” in this specification. For convenience, the at least one module is described as the “first module (510)” in this specification. At this time, the first module (510) may be implemented with a singular or multiple components.

[0166] The first module (510) analyzes a document to be analyzed converted into a specific format and can detect at least one and / or at least one molecular structural formula from the document to be analyzed converted into a specific format. For example, the first module (510) can perform the role of accurately detecting at least one and / or at least one region corresponding to (or corresponding to) the molecular structural formula as a bounding box within the document to be analyzed converted into a specific format. This can be understood as detecting a region containing (or corresponding to) the molecular structural formula within the document converted into a specific format.

[0167] This first module (510) can detect at least one molecular structural formula from an analysis target document converted into a specific format and output a detection result of said molecular structural formula. For example, as described above, the detection result of the molecular structural formula output from the first module (510) may include at least one of a label assigned to the at least one detected molecular structural formula, location information of the at least one detected molecular structural formula, and a score. In this case, the location information refers to the coordinates of the detected molecular structural formula, and may include a bounding box for the area corresponding to the molecular structural formula.

[0168] Next, the document comprehension unit (500) may input the document to be analyzed, converted into a specific format, into at least one other module in order to detect a chemical reaction formula from the document to be analyzed (800). Here, the at least one other module may include a second module (520) specialized in detecting a chemical reaction formula (or chemical reaction information). The second module (520) may be at least one other module and may also be expressed as “at least one other module.” For convenience, the at least one other module is described in this specification as “the second module (520),” and the second module (520) may be implemented with a singular or multiple components. In the present invention, “the first module (510)” is a term referring to “at least one module” for convenience, and “the second module (520)” is a term referring to “at least one other module” for convenience.

[0169] Here, “another module” may mean a module that is not identical to the first module (510) (a physically separate component or a logically distinct module). Accordingly, in this specification, the expression “first module (510)” refers to “at least one module,” and the expression “second module (520)” refers to “at least one other module,” and the two terms are not confused with each other. Additionally, each module may be implemented as a singular or plural component. Specifically, the modules mentioned in the present invention may be implemented as a singular (1) or as a plural (e.g., 2, 3, 4 or more). That is, for convenience, they are described in this specification as “first module (510)” and “second module (520),” respectively, but these modules may be physically separated or may be multiple functional units within a single piece of hardware.

[0170] The second module (520) can analyze a document to be analyzed converted into a specific format and detect at least one and / or at least one chemical reaction formula from the document to be analyzed converted into a specific format. The chemical reaction formula may consist of reactants, reaction conditions, and products. For example, the second module (520) may perform the role of identifying the location of each component constituting at least one chemical reaction formula and assigning a label (or class) to the identified location. This can be understood as identifying the area within the document converted into a specific format that contains (or corresponds to) each component constituting the chemical reaction formula and assigning a label (or class) to the identified area. Such a label can also be understood as information for distinguishing (or distinguishing, identifying, etc.) the reactants, reaction conditions, and products included in the detected chemical reaction formula.

[0171] This second module (520) can detect at least one chemical reaction formula from a document to be analyzed converted into a specific format and output the detection result of the chemical reaction formula. For example, as described above, the detection result of the chemical reaction formula output from the second module (520) may include at least one of a label (or class) assigned to each of the components constituting the at least one detected chemical reaction formula and location information for each of the components. Here, the label assigned to each of the components may refer to a label assigned to an area containing the reactant, reaction conditions, and product, respectively. Additionally, location information refers to the coordinates of each of the components constituting the detected chemical reaction formula, and may include a bounding box for the area corresponding to each of the components. That is, the second module (520) can detect (or extract) a structured chemical reaction formula (or chemical reaction information) from a document to be analyzed converted into a specific format.

[0172] Furthermore, the document comprehension unit (500) can derive (or generate) analysis results for each of at least one molecular structural formula and at least one chemical reaction formula detected from the document to be analyzed (800).

[0173] The document comprehension unit (500) can analyze the bonding relationships between the atoms constituting at least one and / or at least one molecular structural formula detected and derive an analysis result for at least one and / or at least one molecular structural formula detected.

[0174] In the present invention, “analyzing the atoms constituting the molecular structural formula and the bonding relationships between the atoms to derive the results of the analysis of the molecular structural formula” may mean a process of identifying the type, location, and geometric arrangement of each atom included in the molecular structural formula recognized from the document to be analyzed (800), and structurally interpreting the bonding type (e.g., single bond, double bond, triple bond, directional bond, etc.) and connection relationships formed between each atom.

[0175] In this case, the “analysis process” may include defining the topology of the molecular structure, including the type of atoms, bond order, bond direction, presence of a ring structure, and substituent linkage relationships. Additionally, it may include a process of normalizing and expressing the molecular structural information into a machine (or computer) readable format based on the analysis results of the atoms and the bonding relationships between the atoms. For example, the detected molecular structural formula may be converted into a chemical structural representation format (e.g., SMILES, InChI, Mol, etc.) containing atomic and bonding information, which can be utilized for the reconstruction, storage, and subsequent processing of the molecular structure.

[0176] In another embodiment, the “analysis result for the detected molecular structural formula” may be stored or visually reconstructed in conjunction with the layout information of the original document to be analyzed (800) and / or the layout information of the molecular structural formula within the original document to be analyzed (800).

[0177] Additionally, the document understanding unit (500) can analyze the relationships between components constituting at least one and / or at least one chemical reaction equation (e.g., reaction relationship, transformation relationship, causal relationship, structural relationship, correspondence relationship, etc.) of the detected at least one and / or at least one chemical reaction equation to derive the analysis results for the detected at least one and / or at least one chemical reaction equation.

[0178] In the present invention, “analyzing the relationships between the components constituting the chemical reaction equation to derive the analysis results for the chemical reaction equation” may mean a process of separating and identifying the reactants, reaction conditions, and products constituting the chemical reaction equation detected from the document to be analyzed (800) into individual components, and interpreting the structural and semantic associations between the components by comprehensively considering the spatial location information, connection direction, arrow relationship, and contextual information of each component.

[0179] At this time, the “analysis of relationships between the above components” may be performed to include the transformation relationship between reactants and products within a chemical reaction equation, the target reactant or reaction step to which specific reaction conditions are applied, the influence of reaction conditions on the reaction path or result, and a structure in which multiple reactants are combined into a single product or a single reactant branches into multiple products. That is, it may be performed by comprehensively considering the spatial arrangement and connection relationships of each component within the document (800) to be analyzed.

[0180] In addition, the “analysis process” may include a process of generating structured reaction information that clearly expresses what input substance is converted into what product under what conditions by not merely listing each detected component but organically connecting them as a single chemical reaction unit.

[0181] Additionally, the analysis may include determining whether each detected component belongs to the same reaction equation and defining a chemical reaction structure logically combined for each reaction unit, thereby converting the chemical reaction expressed in the document to be analyzed (800) into structured reaction data that can be processed by a machine (or computer). For example, structured reaction data may refer to a formalized data form that can be analyzed and processed by a machine (or computer) by identifying the reactants, products, and reaction conditions constituting the chemical reaction equation expressed in the document as individual elements and logically defining the conversion and application relationships between them.

[0182] The analysis results for the chemical reaction equations derived accordingly are provided in a form in which clear correspondence relationships between reactants, reaction conditions, and products constituting at least one and / or at least one chemical reaction equation are defined, and this can be utilized as basic data for understanding, searching, visualizing, and building databases of chemical reactions.

[0183] Furthermore, the document comprehension unit (500) can generate structured data related to the document to be analyzed (800) by performing data processing on at least one of the detection results of the data detected from the document to be analyzed (detection results of molecular structural formulas and detection results of chemical reaction formulas) and the analysis results of the detected data (analysis results of molecular structural formulas and analysis results of chemical reaction formulas). Since a detailed explanation of this has been described above, it will be omitted to avoid duplication of explanation.

[0184] Meanwhile, the analysis results of the molecular structure formula and the chemical reaction formula detected from the document to be analyzed (800) can be visualized and provided to the user.

[0185] In this case, the document understanding system (1000) can visualize the analysis results for at least one molecular structural formula and at least one chemical reaction formula detected within the document (800) to be analyzed on a user terminal based on (or based on) structured data. Here, “visualizing the analysis results” can be understood as visually providing (or displaying, outputting, etc.) the analysis results for the molecular structural formula and the analysis results for the chemical reaction formula using structured data on a user terminal (10, or a display unit, service page, etc. of the user terminal (10)) so that the user can intuitively recognize the analysis results.

[0186] Specifically, the document understanding system (1000) can provide the analysis result of at least one molecular structural formula and the analysis result of at least one chemical reaction formula detected from the document to be analyzed (800) through a service page output to the user terminal (10).

[0187] In this case, the service page may include at least one and / or at least one or more areas (i.e., multiple areas). As illustrated in FIG. 9, the service page (2000) may include at least one of a first area (2010) in which at least one graphic object corresponding to at least one page included in the document to be analyzed (800) is provided (or displayed), a second area (2020) in which the document to be analyzed (800) is provided (or displayed), and at least one of an analysis result for at least one molecular structural formula and an analysis result for at least one chemical reaction formula is provided (or displayed).

[0188] The visual provision of the above analysis results is not merely a simple display of information, but can be achieved through an interface that is dynamically generated based on the results of structurally processing the chemical domain data included in the document to be analyzed.

[0189] In the document understanding method according to the present invention, detection results of molecular structural formulas and chemical reaction formulas and analysis results thereof can be visually provided on the screen of a user terminal. For example, an area providing a page image included in the document to be analyzed and an area providing analysis results for chemical domain data detected from said page can be displayed together. Accordingly, the user can intuitively check the analysis results while maintaining the original context of the document.

[0190] In addition, the analysis results may be provided with rating information to identify reliability, accuracy, or quality, thereby allowing the user to determine the reliability of the analysis results.

[0191] First, the first area (2010) of the service page (2000) may include a plurality of graphic objects (2011a, 2012a, 2013a, 2014a) corresponding to each of the plurality of pages (e.g., pages 1 to 4) included in the document to be analyzed (800). At this time, the number of graphic objects included in the first area (2010) can be understood as being equal to the number of pages included in the document.

[0192] Next, at least one page included in the document to be analyzed (800) may be provided in the second area (2020) of the service page (2000).

[0193] Specifically, the second area (2020) of the service page (2000) may be provided with any one page of the document to be analyzed (800) that corresponds to any one graphic object selected in the first area (2010) among at least one page included in the document to be analyzed (800). For example, let us assume that among a plurality of graphic objects (2011a, 2012a, 2013a, 2014a) included in the first area (2010), a graphic object (2014a, or the fourth graphic object) corresponding to any one page (e.g., page 4) included in the document to be analyzed (800) is selected. In this case, any one page of the document to be analyzed (800) that corresponds to the selected graphic object (2014a) (2014, or a page image corresponding to any one page) may be provided in the second area (2020).

[0194] At this time, any one page (2014) provided in the second area (2020) may be a page image reconstructed by maintaining the layout of any one page included in the document to be analyzed (800).

[0195] More specifically, the phrase “providing at least one page included in the document to be analyzed (800)” in the present invention can also be understood as providing at least one page image corresponding to at least one page included in the document to be analyzed (800). In this case, the page image may be provided with the same layout as at least one page included in the document to be analyzed (800).

[0196] Such page images may also be understood as page images (or reconstructed images) that reflect detection results while maintaining the layout of any one page included in the document to be analyzed (800). For example, the page image may be an image that provides data (e.g., molecular structural formula, chemical reaction formula, text, table, etc.) detected in the document to be analyzed (800) while maintaining the page composition, arrangement relationships, location, flow, etc. of the document to be analyzed (800) as they are. In this specification, this may also be expressed as “page images corresponding to the layout of the document to be analyzed” or “page images corresponding to the layout of the pages included in the document to be analyzed.”

[0197] Next, in the third area (2030) of the service page (2000), an analysis result for at least one data (e.g., molecular structural formula or chemical reaction formula) related to a chemical domain detected from any one page (2014) provided in the second area (2020) may be provided.

[0198] In this regard, if at least one molecular structural formula is detected from any one page (2014) of the document (800) to be analyzed, the document understanding system (1000) may provide an analysis result for at least one molecular structural formula to the third area (2030). For example, the document understanding system (1000) may provide an analysis result (2031, 2032, 2033) for each of the plurality of molecular structural formulas (2021, 2022, 2023) detected from any one page (2014) to the third area (2030).

[0199] Here, the analysis result for the molecular structural formula may include at least one of the following: at least one molecular structural formula corresponding to the layout of at least one detected molecular structural formula (e.g., the molecular structural formula detected in the original document); at least one standardized molecular structural formula based on at least one molecular structural formula (e.g., the result of reconstructing a digital molecular structural formula by interpreting it with an artificial intelligence model); and at least one chemical structural representation format corresponding to at least one molecular structural formula (e.g., SMILES, InChI, Mol, etc.). In this case, “at least one molecular structural formula corresponding to the layout of at least one detected molecular structural formula” may mean a molecular structural formula that maintains the same layout as the detected molecular structural formula. This may mean maintaining (or preserving) the original arrangement, bonding position, direction, and relative distance of the molecular structural formula included in the document to be analyzed (800) without altering them. That is, it may be a reconstruction that is identical not only to the chemical connection relationships of the molecules but also to their visual arrangement.

[0200] For example, among the analysis results (2031, 2032, 2033) for each of the plurality of molecular structural formulas (2021, 2022, 2023), the analysis result (2031) for the first molecular structural formula (2021) may include at least one of a molecular structural formula (2031a, or molecular structural formula image) corresponding to the layout of the first molecular structural formula (2021), a molecular structural formula (2031b, or molecular structural formula image) standardized based on the first molecular structural formula (2021), and a chemical structure representation format (2031c) corresponding to the first molecular structural formula (2021). In this case, the chemical structure representation format (2031c) may be obtained by analyzing the atoms constituting the first molecular structure formula (2021) and the bonding relationships between the atoms in at least one module (the third module (530)) and converting the first molecular structure formula (2021) into a specific chemical structure representation format (e.g., SMILES).

[0201] At this time, the document understanding system (1000) may display the detection result of at least one molecular structural formula in at least one area containing (or corresponding to at least one molecular structural formula) of the document to be analyzed (800) provided in the second area, so that it is possible to identify (or recognize, confirm) that at least one molecular structural formula has been detected from the document to be analyzed (800). For example, the document understanding system (1000) may display the detection result (2021a) of said molecular structural formula (2021) in at least one area containing said molecular structural formula (2021) of any one page (2014) provided in the second area (2020), so that it is possible to identify that at least one molecular structural formula (2021) has been detected from any one page (2014) of the document to be analyzed (800). In this case, the detection result (2021a) can be displayed through a pre-set method (e.g., graphic object (e.g., square box), overlap, overlay, highlighting, highlighting object, etc.).

[0202] In this way, the document understanding system (1000) can provide the detection result of the molecular structural formula detected from the document (800) to be analyzed and the analysis result of the said molecular structural formula together.

[0203] Additionally, in the third region (2030), a reliability rating for the analysis result of at least one molecular structural formula may be displayed so as to identify the quality of the analysis result for at least one molecular structural formula. For example, in the third region (2030), reliability ratings (2031d, 2032d, 2033d) for the analysis results (2031, 2032, 2033) of each of the plurality of molecular structural formulas (2021, 2022, 2023) may be displayed.

[0204] These reliability grades may include at least one of a first grade (e.g., “High”), a second grade (e.g., “Medium”), and a third grade (e.g., “Low”). In one embodiment, if the reliability score (or quality score) of the analysis result for a molecular structural formula is a first threshold value (e.g., 80 points or higher), the first grade may be displayed on the analysis result for the molecular structural formula. On the other hand, if the reliability score of the analysis result for the molecular structural formula is a second threshold value (e.g., 50 points or higher, but less than 80 points), the second grade may be displayed on the analysis result for the molecular structural formula. On the other hand, if the reliability score of the analysis result for the molecular structural formula is a third threshold value (e.g., less than 50 points), the third grade may be displayed on the analysis result for the molecular structural formula.

[0205] The above reliability rating may be provided in the form of a score, probability value, level, or visual indicator, enabling the identification of the quality of the analysis results.

[0206] Furthermore, when a user input is received selecting one of the analysis results for at least one molecular structural formula, the document understanding system (1000) may provide at least one interface that displays in detail the analysis result for the selected molecular structural formula according to the user input.

[0207] Let us assume that, as illustrated in FIGS. 9 and 10, among the analysis results (2031, 2032, 2033) for each of the plurality of molecular structural formulas (2021, 2022, 2023), an analysis result (2031) for the first molecular structural formula (2021) is selected from the user terminal (10). The document understanding system (1000) may provide at least one interface (2050) that displays the analysis result (2031) for the first molecular structural formula (2021) in detail. Such a page (2050) may also be expressed as a “detail page,” “detail view,” or “detail screen,” etc.

[0208] In one embodiment, the page (2050) may provide a reconstruction result that maintains the same layout of the first molecular structural formula (2021). As one moves from left to right, the original image-based molecular structure (2051a) detected in the document to be analyzed (800), the molecular structure recognized by the artificial intelligence model (170) (2051b), the reconstruction result that maintains the layout (2051c), and the final refined molecular structure (2051d) may be provided in stages.

[0209] Meanwhile, if at least one chemical reaction formula is detected from at least one page of the document to be analyzed (800), the document understanding system (1000) may provide an analysis result for at least one chemical reaction formula in a third area of ​​the service page.

[0210] For example, as illustrated in FIG. 11, the document understanding system (1000) may provide the analysis result (2131) for a chemical reaction equation (2121) detected from any one page (2111) included in the document (800) to the third area (2130) of the service page (2100).

[0211] Here, the analysis result (2131) for the chemical reaction equation (2121) may include at least one chemical reaction information derived by analyzing the relationship between the components constituting at least one chemical reaction equation (e.g., reactant (2131a), reaction condition (2131b), product (2131c), etc.). The chemical reaction information is structured data obtained by automatically detecting and analyzing the key components constituting the chemical reaction equation, and may include the reactant (2131a), reaction condition (2131b), and product (2131c) participating in the reaction, as well as the reaction relationships between them. This is extracted along with location information within the document and is reconstructed and visualized while maintaining the layout of the original document to be analyzed (800), so that the user can intuitively understand and verify the accuracy and context of the reaction.

[0212] At this time, the document understanding system (1000) may display the detection result of at least one chemical reaction formula in at least one area containing (or corresponding to at least one chemical reaction formula) of the document to be analyzed (800) provided in the second area, so that it is possible to identify (or recognize, confirm) that at least one chemical reaction formula has been detected from the document to be analyzed (800). For example, the document understanding system (1000) may display the detection result (2121a, 2121b, 2121c) of the chemical reaction formula (2121) in at least one area containing the chemical reaction formula (2121) of any one page (2111) provided in the second area (2120), so that it is possible to identify that at least one chemical reaction formula (2121) has been detected from any one page (2111) of the document to be analyzed (800).

[0213] Here, the detection result of the chemical reaction equation (2121) may include the detection result of each of the components constituting the chemical reaction equation (2121). For example, the detection result of the chemical reaction equation (2121) may include the detection result of each of the reactant (2121a), reaction condition (2121b), and product (2121c) constituting the chemical reaction equation (2121). These detection results may be displayed with different visual appearances (e.g., reactants in a blue square box, reaction conditions in a red square box, products in a green square box, etc.). In this case, the detection results may be displayed with different visual appearances through a pre-set method (e.g., graphic object (e.g., square box), overlap, overlay, highlighting, highlighting object, etc.).

[0214] In this way, the document understanding system (1000) can provide the detection result of a chemical reaction formula detected from the document to be analyzed (800) and the analysis result for the said chemical reaction formula together.

[0215] Additionally, although not illustrated, a reliability grade for the analysis result of at least one chemical equation may be displayed in the third area (2130) so as to identify the quality of the analysis result for at least one chemical equation. Such reliability grades may include at least one of a first grade (e.g., “High”), a second grade (e.g., “Medium”), and a third grade (e.g., “Low”). In one embodiment, if the reliability score (or quality score) of the analysis result for the chemical equation is a first threshold value (e.g., 80 points or higher), the first grade may be displayed on the analysis result for the chemical equation. On the other hand, if the reliability score of the analysis result for the chemical equation is a second threshold value (e.g., 50 points or higher, but less than 80 points), the second grade may be displayed on the analysis result for the chemical equation. On the other hand, if the reliability score of the analysis result for a chemical equation is a third criterion value (e.g., less than 50 points), a third grade may be indicated on the analysis result for the chemical equation.

[0216] Furthermore, when a user input is received selecting one of the analysis results for at least one chemical reaction formula, the document understanding system (1000) may provide at least one interface that displays in detail the analysis result for the selected chemical reaction formula according to the user input.

[0217] Let us assume that an analysis result for a first chemical reaction equation (2121) is selected from a user terminal (10) as illustrated in FIGS. 11 and 12. The document understanding system (1000) may provide at least one interface (2150) that displays the analysis result (2131) for the first chemical reaction equation (2121) in detail. This interface (2150) may also be expressed as a “detail page,” “detail view,” or “detail screen,” etc.

[0218] In one embodiment, the interface (2150) may be provided with a reaction relationship between a reactant (2151a), reaction conditions (2151b), and a product (2151c) constituting a first chemical equation (2121). This may be a visual representation that “the first chemical equation (2121) is a reaction in which the reactant (2151a) is converted into the product (2151c) through the reaction conditions (2151b).”

[0219] Meanwhile, the service page described above can be updated based on (or based on) user input received from the user terminal (10).

[0220] When the document understanding system (1000) receives a service page update request from a user terminal (10), it can update the service page in response to the received request. Here, “service page update” can be understood as updating the analysis results of the document (or page of the document) provided (or output) on the service page and the data detected in the document (e.g., molecular structural formula, chemical reaction formula, etc.).

[0221] For example, as illustrated in FIG. 9, the document understanding system (1000) can update the service page (2000) based on receiving a user input selecting one graphic object (2011a) included in the first area (2010), while at least one page (2014) included in the document (800) to be analyzed and the analysis results (2031, 2032, 2033) for chemical data detected in said at least one page (2014) are respectively output to the second area (2020) and the third area (2030) of the service page (2000). Here, receiving a user input selecting one graphic object (2011a) included in the first area (2010) can be understood as receiving a request to update the service page. Alternatively, it can also be understood as receiving a user input requesting to switch said at least one page (2014) to at least one other page.

[0222] In one embodiment, the at least one page (2014) may be understood as the fourth page (e.g., page 4) among a plurality of pages included in the document (800) to be analyzed, and at least one other page may be understood as the first page different from the fourth page.

[0223] Accordingly, the document understanding system (1000) can provide an updated service page to the user terminal (10). For example, as illustrated in FIG. 11, the document understanding system (1000) can provide an updated service page (2100) to the user terminal (10) in response to a service page update request received from the user terminal (10). In this case, the updated service page (2100) may include at least one other page (2111) converted based on user input and an analysis result (2130) for chemical data detected in said at least one other page (2111).

[0224] That is, the document understanding system (1000) can update a service page so that at least one other page and an analysis result of chemical data detected on said at least one other page are provided to the user terminal (10) in accordance with an update request received from the user terminal (10).

[0225] Meanwhile, the service page described above can be implemented in various forms. According to one embodiment of the present invention, as illustrated in FIG. 13, the service page (2300) may include at least one of a first area (2310) in which at least one graphic object corresponding to at least one page included in the document to be analyzed is provided, a second area (2320) in which the original document to be analyzed is provided, a third area (2330) in which a detection result of at least one molecular structural formula or at least one chemical reaction formula detected in the document to be analyzed is displayed and provided, and a fourth area (2340) in which at least one of an analysis result for at least one molecular structural formula and an analysis result for at least one chemical reaction formula is provided.

[0226] Additionally, according to another embodiment of the present invention, a document understanding system (1000) may provide, based on the selection from a user terminal (10) of at least one area containing (or corresponding to at least one molecular structural formula) of a document to be analyzed (or a page (2421) of a document to be analyzed) provided in a second area (2420) of a service page (2400), an analysis result (2431) for a molecular structural formula (2421a) included in said at least one area and detailed information (2432) of said analysis result together.

[0227] Furthermore, according to another embodiment of the present invention, when a document understanding system (1000) receives a user request from a user terminal to edit an analysis result for at least one molecular structural formula, it may provide an editing interface to the user terminal (10) that provides an editing function for an analysis result for at least one molecular structural formula in response to the user request. For example, as illustrated in FIG. 15, the document understanding system (1000) may provide an editing interface (2512) that provides an editing function for an analysis result for a specific molecular structural formula in a region (2510) of a service page (2500) based on receiving a user request from a user terminal to edit an analysis result (2511) for a specific molecular structural formula. At this time, when editing of the analysis result for the specific molecular structural formula is performed through the editing interface (2512), the edited analysis result may be stored in a previously specified storage (e.g., storage unit (400) or memory, etc.).

[0228] According to one embodiment, a user may provide user input for editing analysis results regarding a molecular structural formula or a chemical reaction formula. For example, the user may modify atomic or bonding information of a detected or analyzed molecular structural formula, or edit components of a chemical reaction formula. Analysis results modified based on the editing input may be stored in a pre-configured repository and subsequently utilized for reuse, sharing, or further analysis of document understanding results.

[0229] In one embodiment, editing the analysis result (2511) for the specific molecular structural formula may be editing the specific molecular structural formula (2511a) included in the analysis result (2511). This may be selected according to user input (or user request). In this case, the editing interface (2512) may be provided with the specific molecular structural formula (2512a) selected by the user for editing.

[0230] Furthermore, although not illustrated, according to another embodiment of the present invention, when a document understanding system (1000) receives a user request from a user terminal to edit an analysis result for at least one chemical reaction formula, in response to the user request, it may provide an editing interface to the user terminal that provides an editing function for an analysis result for at least one chemical reaction formula. At this time, when editing of an analysis result for at least one chemical reaction formula is performed through the editing interface, the edited analysis result may be stored in a previously specified storage.

[0231] Each processing step described in the document understanding method according to the present invention may be performed by a single processing module or a plurality of processing modules, and these processing modules may be integrated within the same processing device or may operate in a distributed environment.

[0232] Furthermore, the document understanding method according to the present invention can be implemented in various processing environments including a server-based environment, a cloud environment, or a user terminal.

[0233] As described above, the document understanding method and system according to the present invention can automatically detect and analyze molecular structural formulas and chemical reaction information contained in a document, and visualize and provide this information to the user. Through this, the user can intuitively recognize the necessary information and understand it more quickly, thereby increasing the accuracy and efficiency of research. In other words, the user can receive the necessary information from the document quickly and accurately, thus reducing the time and cost required for research or development.

[0234] Furthermore, according to the document understanding method and system of the present invention, by providing the user with an image reconstructed with the same layout as the original document, the user is supported in easily comparing the original document with the analysis results and enhancing reliability. That is, when reconstructing the detection results of molecular structural formulas and chemical reaction information contained in a document, the present invention maintains (or preserves) the layout of the original document, thereby providing an environment in which the user can intuitively verify the analysis results.

[0235] Furthermore, according to the document understanding method and system of the present invention, an editing interface can be provided that allows for real-time modification of the analysis results of molecular structural formulas detected from a document and the analysis results of chemical reaction formulas detected from a document. That is, the present invention provides an intuitive editing environment based on the visual comparison of the analysis results and the original document, and can store data edited (or modified) by the user. Through this, the present invention enables the effective collection of user feedback and allows for continuous learning and performance optimization of the model.

[0236] Furthermore, according to the document understanding method and system of the present invention, by displaying and providing the reliability grade of the analysis results for molecular structural formulas detected from documents and the analysis results for chemical reaction formulas detected from documents, the user is supported in intuitively understanding the quality of each analysis result and making quick decisions.

[0237] Furthermore, according to the document understanding method and system of the present invention, by visually providing the analysis results of detected data in conjunction with the location of the original document, the user can easily verify the analysis results and efficiently modify them when necessary. In other words, the present invention can simultaneously improve the accuracy of analysis and productivity through a user-friendly interface.

[0238] Furthermore, the present invention can store edited analysis results, detected data, and location information within documents. This data can be utilized for document search and database construction, thereby providing high value in research and commercial applications.

[0239] Meanwhile, the present invention described above can be implemented based on a quantum computer. The present invention implemented based on a quantum computer may include a qubit-based quantum processor and quantum memory, and may include software and hardware interfaces optimized for quantum computation.

[0240] Quantum processors in quantum computers utilize qubits to efficiently process complex operations through parallel computation, quantum entanglement, and quantum superposition, which cannot be performed by the binary bits of classical computers. Quantum processors process data using quantum gates and can provide exponential speed improvements for specific problems.

[0241] Meanwhile, the present invention described above can be implemented as a program that is executed by one or more processes on a computer and can be stored on a computer-readable medium (or recording medium).

[0242] Furthermore, the present invention described above can be implemented as computer-readable code or instructions on a medium on which a program is recorded. That is, the present invention can be provided in the form of a program.

[0243] Meanwhile, computer-readable media include all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0244] Furthermore, the computer-readable medium may be a server or cloud storage that includes a storage and is accessible to an electronic device via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage via wired or wireless communication.

[0245] A computer program may reach the system (1000) through various suitable transmission mechanisms. The transmission mechanism may be, for example, a computer-readable storage medium, a computer program product, a memory device, a recording medium such as a CD-ROM or DVD, or a product that tangibly embodies the computer program. The transmission mechanism may be a signal configured to reliably transmit the computer program through air or an electrical connection. The system (1000) may propagate or transmit the computer program as a computer data signal.

[0246] Furthermore, references to 'computer-readable storage media,' 'computer program products,' 'computer programs embodied in a tangible form,' etc., or to 'controller,' 'computer,' 'processor,' etc., should be understood to include not only computers with various architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures, but also specialized circuits such as Field-Programmable Gate Arrays (FPGAs), Application Specific Circuits (ASICs), signal processing units, and other devices. References to computer programs, instructions, code, etc., should be understood to include software for programmable processors or firmware, such as programmable content for hardware devices, whether it is instructions for a processor or configuration settings for a fixed-function device, gate array, or programmable logic device.

[0247] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, namely a CPU (Central Processing Unit), and no special limitations are placed on its type.

[0248] Meanwhile, the above detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.

Claims

1. Regarding methods performed by a computer, A step of specifying at least one document to be analyzed; A step of processing the above-mentioned document to be analyzed to detect at least one molecular structural formula from the above-mentioned document to be analyzed; A step of processing the above-mentioned document to be analyzed and detecting at least one chemical reaction formula from the above-mentioned document to be analyzed; A step of analyzing the atoms constituting at least one detected molecular structural formula and the bonding relationships between the atoms to derive an analysis result for at least one detected molecular structural formula; A step of analyzing the relationships between the components constituting at least one chemical reaction equation detected above to derive an analysis result for at least one chemical reaction equation detected above; and A document understanding method characterized by including the step of providing to the user terminal the analysis result of at least one detected molecular structural formula and the analysis result of at least one detected chemical reaction formula.

2. In Paragraph 1, The analysis result of at least one detected molecular structural formula and the analysis result of at least one detected chemical reaction formula are provided through a service page output to the user terminal, and The above service page is, A first area provided with at least one graphic object corresponding to at least one page included in the above-mentioned document to be analyzed, The second area where the above-mentioned analysis target document is provided and A method for understanding a document characterized by including at least one of a third region in which at least one of the analysis result of at least one detected molecular structural formula and the analysis result of at least one detected chemical reaction formula is provided.

3. In Paragraph 2, A document understanding method characterized by displaying the detection result of the at least one molecular structural formula and the detection result of the at least one chemical reaction formula in at least one area containing the at least one molecular structural formula and at least one chemical reaction formula of the document to be analyzed, provided in the second area, so as to be able to identify that the at least one molecular structural formula and the at least one chemical reaction formula have been detected from the document to be analyzed.

4. In Paragraph 2, In the second area above, at least one page included in the document to be analyzed is provided, and In the aforementioned third area, A document understanding method characterized by providing an analysis result for at least one data related to a chemical domain detected from at least one page provided in the second area.

5. In Paragraph 4, If at least one molecular structural formula is detected from the above at least one page, In order to make it possible to identify that the at least one molecular structural formula has been detected from the at least one page, the detection result of the at least one molecular structural formula is displayed in at least one area containing the at least one molecular structural formula of the at least one page provided in the second area, and A document understanding method characterized in that the third region above provides an analysis result for the at least one molecular structural formula detected from the at least one page.

6. In Paragraph 5, The analysis results for the above at least one molecular structural formula are, It includes at least one molecular structural formula corresponding to the layout of the at least one detected molecular structural formula, at least one molecular structural formula standardized based on the at least one detected molecular structural formula, and at least one chemical structural representation format corresponding to the at least one molecular structural formula, and The above at least one chemical structure representation format is, A document understanding method characterized by analyzing atoms constituting the at least one molecular structural formula and the bonding relationships between the atoms in at least one module, and converting the at least one molecular structural formula into the at least one chemical structure expression format.

7. In Paragraph 6, When user input is received selecting any one of the analysis results for at least one of the above molecular structural formulas, A document understanding method characterized by providing at least one interface that displays in detail the analysis results for any one of the molecular structural formulas selected according to the above user input.

8. In Paragraph 4, If at least one chemical reaction formula is detected from the above at least one page, In order to identify that the at least one chemical reaction formula has been detected from the at least one page, the detection result of the at least one chemical reaction formula is displayed in the at least one area containing the at least one chemical reaction formula of the at least one page provided in the second area, and A document understanding method characterized in that the third area above provides an analysis result for the at least one chemical reaction formula detected from the at least one page.

9. In Paragraph 8, The detection result of the above at least one chemical reaction equation includes the detection result of each of the components constituting the above at least one chemical reaction equation, and The detection result of each of the above components is displayed in the second area with a different visual appearance, and The above components are, A method for understanding documents characterized by including at least one of reactants, reaction conditions, and products.

10. In Paragraph 8, The analysis result for the above at least one chemical reaction equation is, A document understanding method characterized by including at least one chemical reaction information derived by analyzing the relationship between the components constituting the above at least one chemical reaction equation.

11. In Paragraph 10, When user input is received selecting any one of the analysis results for at least one of the above chemical reaction equations, A document understanding method characterized by providing at least one page that displays in detail the analysis results for any one of the chemical reaction equations selected according to the above user input.

12. In Paragraph 2, In the aforementioned third area, A method for understanding a document characterized by displaying a reliability grade for the analysis result of the at least one molecular structural formula so as to enable identification of the quality of the analysis result of the at least one molecular structural formula.

13. In Paragraph 2, In the aforementioned third area, A method for understanding a document characterized by displaying a reliability grade for the analysis result of at least one chemical reaction formula so as to enable identification of the quality of the analysis result for at least one chemical reaction formula.

14. In Paragraph 1, When a user request to edit the analysis result for the at least one molecular structural formula is received from the above user terminal, The method further includes the step of providing an editing interface to the user terminal in response to the above user request, which provides an editing function for the analysis result of the at least one molecular structural formula. A document understanding method characterized by the fact that when an analysis result of at least one molecular structural formula is edited through the above-mentioned editing interface, the edited analysis result is stored in a previously specified repository.

15. In Paragraph 1, When a user request to edit the analysis result for the at least one chemical reaction equation is received from the above user terminal, The method further includes the step of providing an editing interface to the user terminal in response to the above user request, which provides an editing function for the analysis result of the at least one chemical reaction equation. A document understanding method characterized by the fact that when an analysis result of at least one chemical reaction equation is edited through the above-mentioned editing interface, the edited analysis result is stored in a previously specified repository.

16. In Paragraph 1, The step of detecting at least one molecular structural formula from the above-mentioned document to be analyzed is, This is performed by a first processing module that performs molecular structural formula detection, and The first processing module outputs at least one of a label, location information, and score corresponding to at least one detected molecular structural formula, and The step of detecting at least one chemical reaction formula from the above-mentioned document to be analyzed is, This is performed by a second processing module that performs chemical reaction equation detection, and The second processing module outputs at least one of class and position information corresponding to each of the components constituting the at least one detected chemical reaction equation, and The above components are, A method for understanding documents characterized by including at least one of reactants, reaction conditions, and products.

17. In Paragraph 1, It further includes a step of converting the above-mentioned document to be analyzed into a specific pre-set format, and When the above-mentioned document subject to analysis is converted into the above-mentioned specific format, The analysis target document converted into the above specific format is input into at least one module to detect the at least one molecular structural formula, and A document understanding method characterized by inputting an analysis target document converted into the above-mentioned specific format into at least one other module to detect the above-mentioned at least one chemical reaction equation.

18. In Paragraph 2, In the second area mentioned above, At least one page image corresponding to at least one page included in the above-mentioned document to be analyzed is provided, and The above at least one page image is, A method for understanding a document characterized by being reconstructed with the same layout as at least one page included in the document to be analyzed.

19. A system comprising memory configured to store executable instructions and one or more processors configured to perform operations by executing one or more instructions, The above system is, Identify at least one document to be analyzed, and By processing the above-mentioned document to be analyzed, at least one molecular structural formula is detected from the above-mentioned document to be analyzed, and Processing the above-mentioned document to be analyzed to detect at least one chemical reaction formula from the above-mentioned document to be analyzed, and By analyzing the atoms constituting at least one detected molecular structural formula and the bonding relationships between the atoms, an analysis result for at least one detected molecular structural formula is derived. By analyzing the relationships between the components constituting at least one chemical reaction equation detected above, an analysis result for at least one chemical reaction equation detected above is derived, and A document understanding system characterized by providing to the user terminal above an analysis result for at least one detected molecular structural formula and an analysis result for at least one detected chemical reaction formula.

20. A program that is executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The above program is, A step of specifying at least one document to be analyzed; A step of processing the above-mentioned document to be analyzed to detect at least one molecular structural formula from the above-mentioned document to be analyzed; A step of processing the above-mentioned document to be analyzed and detecting at least one chemical reaction formula from the above-mentioned document to be analyzed; A step of analyzing the atoms constituting at least one detected molecular structural formula and the bonding relationships between the atoms to derive an analysis result for at least one detected molecular structural formula; A step of analyzing the relationships between the components constituting at least one chemical reaction equation detected above to derive an analysis result for at least one chemical reaction equation detected above; and A program stored on a computer-readable recording medium characterized by including instructions for performing the step of providing to the user terminal an analysis result for at least one detected molecular structural formula and an analysis result for at least one detected chemical reaction formula.