Method and system for document understanding
Patent Information
- Application Number
- PCT/KR2026/001830
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2026-01-06
- Filing Date
- 2026-01-30
- Publication Date
- 2026-08-27
Smart Images

Figure KR2026001830_27082026_PF_FP_ABST
Abstract
Description
Document Understanding Methods and Systems
[0001] The present invention relates to a document understanding method and system, and provides a document understanding method and system using a model specialized for chemical information processing.
[0002] Documents are recorded through characters or symbols (or codes) to preserve and transmit thoughts, ideas, intentions, and information; in a broad sense, they can include everything that contains meaning, such as pictures, photographs, and videos. This means that information is recorded and preserved in various forms and can serve as a means of communication.
[0003] In this regard, a large volume of electronic documents created via computers is being utilized in various fields in modern society. These documents contain not only text but also various forms of information, such as tables, images, and graphs. For example, documents used in industrial settings, such as papers, patents, technical reports, and manuals, include not only text but also diverse forms of information like tables, graphs, diagrams, images, mathematical formulas, and molecular structural formulas.
[0004] While humans can read and understand the various forms of information contained in such documents, computers cannot fully comprehend complex forms of information in the same way. Consequently, analyzing and utilizing document content requires the cumbersome task of converting it into a format recognizable by computers. However, as documents contain a mixture of diverse information, manually reviewing or analyzing them can consume a significant amount of time and cost.
[0005] To address this, various technologies have conventionally been proposed to understand and analyze various forms of information contained in documents and to provide optimized solutions in various research and industrial settings.
[0006] However, conventional technologies are often optimized for specific domains or types of information (e.g., text-centric, table-centric, or graph-centric), so they have limitations in that they are difficult to apply to documents containing information from different domains or complex forms.
[0007] Furthermore, conventional technologies were limited to processing a single piece of information individually. For example, conventional chemical information detection technology was limited to processing either molecular structure (or molecular formula) or chemical reaction diagram individually. In this case, it is difficult to provide integrated information because molecular formulas and chemical reaction diagrams cannot be processed simultaneously, and it is difficult to directly extract reaction information from page-level data within a document.
[0008] Therefore, there is still a need for methods to deeply understand various forms of documents and efficiently process chemical information from them.
[0009] The present invention is intended to provide a document understanding method and system capable of effectively understanding various forms of documents.
[0010] More specifically, the present invention aims to provide a document understanding method and system that can maximize the efficiency of document-based chemical information processing and expand the potential for utilization in research and industrial settings.
[0011] In addition, the present invention is intended to provide a model specialized for efficiently processing chemical information contained in documents.
[0012] In particular, the present invention is intended to provide a model capable of simultaneously processing molecular structural formulas and chemical reaction diagrams included in a document.
[0013] Furthermore, the present invention aims to provide a document understanding method and system capable of improving efficiency in various industrial or research fields and providing optimized solutions.
[0014] To solve the problem described above, a document understanding method according to the present invention, performed by a computer, may include the steps of: specifying at least one document to be analyzed; processing the document to be analyzed as input to a plurality of models specialized in a chemical domain; detecting at least one molecular structural formula included in the document to be analyzed in a first model among the plurality of models and outputting a detection result of the at least one molecular structural formula; recognizing at least one chemical reaction formula included in the document to be analyzed in a second model among the plurality of models and outputting a recognition result of the at least one chemical reaction formula; and generating at least one output for chemical data included in the document to be analyzed using the detection result of the at least one molecular structural formula and the recognition result of the at least one chemical reaction formula.
[0015] In an embodiment, the step of generating the at least one output may combine the detection result of the at least one molecular structure formula output from the first model and the recognition result of the at least one chemical reaction formula output from the second model to generate the at least one output for the chemical data included in the document to be analyzed.
[0016] In an embodiment, the first model may be configured to detect at least one molecular structural formula in at least one page included in the document to be analyzed, and to output the detection result of the at least one molecular structural formula detected in the at least one page.
[0017] In an embodiment, the detection of the at least one molecular structural formula may include a task of detecting the at least one molecular structural formula in the document to be analyzed, analyzing the atoms constituting the at least one molecular structural formula and the bonding relationships between the atoms, and converting the at least one molecular structural formula into structural data based on the analysis results.
[0018] In an embodiment, the detection result of the at least one molecular structural formula may include at least one of the detected at least one molecular structural formula, location information of the detected at least one molecular structural formula, and at least one region of the document to be analyzed that includes the detected at least one molecular structural formula.
[0019] In an embodiment, the second model may be configured to recognize at least one chemical reaction formula in at least one page included in the document to be analyzed, and to output the recognition result of the at least one chemical reaction formula recognized in the at least one page.
[0020] In an embodiment, the at least one chemical reaction equation recognition may include a task of extracting reactants, reaction conditions, and products constituting the at least one chemical reaction equation from the document to be analyzed, or extracting reactants, reaction conditions, and products from a chemical reaction diagram included in the document to be analyzed.
[0021] In an embodiment, the recognition result of the at least one chemical reaction equation may include at least one of information on the components constituting the at least one chemical reaction equation detected, location information of each of the components, at least one region of the document to be analyzed containing the at least one chemical reaction equation, and information on the reaction relationship between the components.
[0022] In an embodiment, the document to be analyzed is configured to include at least one page, and the step of generating the at least one output may combine the detection result of the at least one molecular structural formula detected on the at least one page and the recognition result of the at least one chemical reaction formula detected on the at least one page to generate the at least one output for the chemical data included in the document to be analyzed.
[0023] In the embodiment, the step of detecting the at least one molecular structural formula included in the document to be analyzed in the first model and outputting the detection result of the at least one molecular structural formula, and the step of recognizing the at least one chemical reaction formula included in the document to be analyzed in the second model and outputting the recognition result of the at least one chemical reaction formula, may be steps performed in parallel.
[0024] In an embodiment, the first model performs inference to detect the at least one molecular structural formula included in the document to be analyzed and outputs a first inference result related to the at least one molecular structural formula, and the second model performs inference to recognize the at least one chemical reaction formula included in the document to be analyzed and outputs a second inference result related to the at least one chemical reaction formula.
[0025] In an embodiment, the first inference result includes a detection result of the at least one molecular structural formula, and the second inference result may include a recognition result of the at least one chemical reaction formula.
[0026] In an embodiment, the first inference result and the second inference result can be combined to generate at least one output for chemical data included in the document to be analyzed.
[0027] In an embodiment, based on the structural association between the detection result of the at least one molecular structure formula and the recognition result of the at least one chemical reaction formula, the detection result of the at least one molecular structure formula and the recognition result of the at least one chemical reaction formula can be combined to generate the at least one output.
[0028] In an embodiment, the at least one output may include a first molecular structural formula not included in the at least one chemical reaction formula among the at least one molecular structural formulas, a second molecular structural formula included in the at least one chemical reaction formula among the at least one molecular structural formulas, and a chemical reaction formula including the second molecular structural formula.
[0029] In an embodiment, the method further includes the step of converting the document to be analyzed into a pre-set specific format, and when the document to be analyzed is converted into the specific format, the document to be analyzed converted into the specific format is processed as an input to the plurality of models, and in the first model among the plurality of models, at least one molecular structural formula is detected from the document to be analyzed converted into the specific format and a detection result of the at least one molecular structural formula is output, and in the second model among the plurality of models, at least one chemical reaction formula is recognized from the document to be analyzed converted into the specific format and a recognition result of the at least one chemical reaction formula is output, and the at least one output can be generated using the detection result of the at least one molecular structural formula and the recognition result of the at least one chemical reaction formula.
[0030] A document understanding system according to the present invention, comprising a memory configured to store executable instructions and one or more processors configured to perform operations by executing one or more instructions, wherein the system specifies at least one document to be analyzed and processes the document to be analyzed as input to a plurality of models specialized in a chemical domain, wherein in a first model among the plurality of models, at least one molecular structural formula included in the document to be analyzed is detected and outputs a detection result of the at least one molecular structural formula, wherein in a second model among the plurality of models, at least one chemical reaction formula included in the document to be analyzed is recognized and outputs a recognition result of the at least one chemical reaction formula, and wherein at least one output of chemical data included in the document to be analyzed is generated using the detection result of the at least one molecular structural formula and the recognition result of the at least one chemical reaction formula.
[0031] A program according to the present invention is a program that is executed by one or more processes in an electronic device and can be stored on a computer-readable recording medium, and may include instructions for performing the steps of: specifying at least one document to be analyzed; processing the document to be analyzed as input to a plurality of models specialized in a chemical domain; detecting at least one molecular structural formula included in the document to be analyzed in a first model among the plurality of models and outputting a detection result of the at least one molecular structural formula; recognizing at least one chemical reaction formula included in the document to be analyzed in a second model among the plurality of models and outputting a recognition result of the at least one chemical reaction formula; and generating at least one output for chemical data included in the document to be analyzed using the detection result of the at least one molecular structural formula and the recognition result of the at least one chemical reaction formula.
[0032] A document understanding method according to the present invention comprises: a step of specifying at least one document to be analyzed; a step of processing the document to be analyzed as input to at least one information processing model configured to process information related to a chemical domain; a step of generating a first processing result for molecular structure information included in the document to be analyzed and generating a second processing result for chemical reaction information included in the document to be analyzed using the at least one information processing model; and a step of generating output data for chemical information reflecting the mutual correlation between the molecular structure information and the chemical reaction information using the first processing result and the second processing result.
[0033] In an embodiment, the at least one information processing model comprises a first model for generating the first processing result for the molecular structure information included in the document to be analyzed and a second model for generating the second processing result for the chemical reaction information included in the document to be analyzed, and the output data is generated by combining the first processing result output from the first model and the second processing result output from the second model.
[0034] In an embodiment, the first model is characterized by detecting at least one molecular structural formula from at least one page included in the document to be analyzed and outputting the first processing result for the molecular structural information.
[0035] In an embodiment, the detection of the at least one molecular structural formula is characterized by including a task of detecting the at least one molecular structural formula in the document to be analyzed, analyzing the bonding relationships between the atoms constituting the at least one molecular structural formula and the atoms, and converting the at least one molecular structural formula into structural data based on the analysis results.
[0036] In an embodiment, the first processing result is characterized by including at least one of the at least one molecular structural formula, location information of the at least one molecular structural formula, and at least one region information containing the at least one molecular structural formula in the document to be analyzed.
[0037] In an embodiment, the method is characterized by recognizing at least one chemical reaction formula from at least one page included in the document to be analyzed, and outputting the second processing result for the chemical reaction information.
[0038] In an embodiment, the at least one chemical reaction equation recognition is characterized by including a task of extracting at least one of reactants, reaction conditions, and products constituting the at least one chemical reaction equation from the document to be analyzed, or extracting at least one of reactants, reaction conditions, and products from a chemical reaction diagram included in the document to be analyzed.
[0039] In an embodiment, the second processing result is characterized by including at least one of information on components constituting the at least one chemical reaction equation, location information of the components, information on at least one region containing the at least one chemical reaction equation in the document to be analyzed, and information on the reaction relationship between the components.
[0040] In an embodiment, the document to be analyzed is configured to include at least one page, and the output data is characterized by combining a detection result of at least one molecular structural formula corresponding to the first processing result detected on the at least one page and a recognition result of at least one chemical reaction formula corresponding to the second processing result detected on the at least one page to generate the output data for the chemical information included in the document to be analyzed.
[0041] In an embodiment, the step of processing the molecular structure information in the first model and outputting the first processing result, and the step of processing the chemical reaction information in the second model and outputting the second processing result are characterized as being performed in parallel.
[0042] In an embodiment, the first model performs inference for processing the molecular structure information and outputs a first inference result, and the second model performs inference for processing the chemical reaction information and outputs a second inference result.
[0043] In an embodiment, the first inference result includes the first processing result for the molecular structure information, and the second inference result includes the second processing result for the chemical reaction information.
[0044] In an embodiment, the output data is characterized by being generated by combining the first inference result and the second inference result.
[0045] In an embodiment, the output data is characterized by being generated based on the structural association between the first processing result of the molecular structure information and the second processing result of the chemical reaction information.
[0046] In an embodiment, the output data is characterized by including first molecular structure information not included in the chemical reaction formula, second molecular structure information included in the chemical reaction formula, and chemical reaction information including the second molecular structure information.
[0047] In an embodiment, the method further includes the step of converting the document to be analyzed into a pre-set specific format, and when the document to be analyzed is converted into the specific format, the document to be analyzed converted into the specific format is processed as an input to at least one model, and the document to be analyzed converted into the specific format is processed as an input to a first model and a second model included in the at least one model, and a first processing result regarding the molecular structure information is output from the first model, and a second processing result regarding the chemical reaction information is output from the second model, and the output data is generated by combining the first processing result and the second processing result.
[0048] A system comprising a memory configured to store executable instructions according to the present invention and one or more processors configured to perform operations by executing one or more instructions is characterized by specifying at least one document to be analyzed, processing said document to be analyzed as input to at least one information processing model configured to process information related to a chemical domain, using said at least one information processing model to generate a first processing result for molecular structure information included in said document to be analyzed and to generate a second processing result for chemical reaction information included in said document to be analyzed, and using said first processing result and said second processing result to generate output data for chemical information reflecting the mutual correlation between said molecular structure information and said chemical reaction information.
[0049] A program according to the present invention, which is executed by one or more processes in an electronic device and stored in a computer-readable recording medium, is characterized by comprising instructions for performing the steps of: specifying at least one document to be analyzed; processing the document to be analyzed as input to at least one information processing model configured to process information related to a chemical domain; generating a first processing result for molecular structure information included in the document to be analyzed and generating a second processing result for chemical reaction information included in the document to be analyzed using the at least one information processing model; and generating output data for chemical information reflecting the mutual correlation between the molecular structure information and the chemical reaction information using the first processing result and the second processing result.
[0050] As described above, the document understanding method and system according to the present invention utilizes a model specialized for molecular structure formula processing and a model specialized for chemical reaction formula processing to perform inference on a document in parallel, and combines the molecular structure formula detection results and chemical reaction formula recognition results in a post-processing step to generate an integrated result. Through this, the present invention connects molecular structure formula detection and chemical reaction formula recognition, which were previously processed individually, into a single workflow, thereby maximizing the efficiency of document-based chemical information processing and significantly expanding the potential for application in research and industrial settings.
[0051] Furthermore, according to the document understanding method and system of the present invention, molecular structural formulas and chemical reaction formulas contained in a document can be simultaneously detected and analyzed by utilizing a model specialized for processing molecular structural formulas and a model specialized for processing chemical reaction formulas. The results of such detection and analysis can be provided to a user, allowing the user to intuitively recognize the necessary information and understand it more quickly, thereby increasing the accuracy and efficiency of the research. In other words, the user can receive the necessary information from the document quickly and accurately, thus reducing the time and cost required for research or development.
[0052] Furthermore, according to the document understanding method and system of the present invention, by utilizing a model specialized for processing molecular structural formulas to detect molecular structural formulas in documents and converting them into structural data, they can be utilized for building chemical databases, searching for and analyzing chemical information, etc.
[0053] Furthermore, according to the document understanding method and system of the present invention, a chemical reaction equation composed of reactants, conditions, and products can be extracted from a page-level document by utilizing a model specialized for chemical reaction equation processing. In other words, the present invention enables the understanding of structural data of a chemical reaction, the visualization of the reaction pathway, and the automatic processing of reaction data.
[0054] This invention can also be applied to automatically process chemical information contained in academic papers, patent documents, or LaTeX-based technical documents.
[0055] The document understanding method and system according to the present invention can provide an improvement effect in at least one of processing accuracy, processing efficiency, level of automation, or data usability by automatically processing chemical information within a document. Such effects may vary depending on the system configuration, processing environment, or field of application, and are not limited to a specific performance level or effect.
[0056] FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented.
[0057] FIG. 2 illustrates an example of a block diagram of a computing device that may be included in a user computing device, a server computing system, and a training computing system, as an embodiment of a computing system in which the present invention can be implemented.
[0058] Figure 3 illustrates an example of a block diagram from another perspective of a computing device, which is one of the components of a computing system.
[0059] FIGS. 4, FIGS. 5, and FIGS. 6 are conceptual diagrams for explaining a document understanding system according to the present invention.
[0060] FIG. 7 is a flowchart illustrating a method for understanding documents according to the present invention.
[0061] FIGS. 8, FIGS. 9, FIGS. 10, FIGS. 11, FIGS. 12, FIGS. 13, and FIGS. 14 are conceptual diagrams for explaining a document understanding method according to the present invention.
[0062] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components are assigned the same reference number regardless of the drawing symbols, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention.
[0063] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0064] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0065] Singular expressions include plural expressions unless the context clearly indicates otherwise.
[0066] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0067] Meanwhile, FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented. In this regard, the document understanding system according to the present invention can be implemented through a computing device described below and can perform data processing related to the document understanding method described in this specification.
[0068] Referring to FIG. 1, a computing system (10000) that performs a method of simultaneously detecting a molecular structure formula and recognizing a chemical reaction formula in a document according to one embodiment of the present invention and combining the result of the molecular structure formula detection and the result of the chemical reaction formula recognition may include at least one computing device. At this time, the at least one computing device may be a single processor or a multiprocessor computing device.
[0069] The components of at least one computing device of the present invention may include various hardware components such as one or more processors, memory, other hardware, and a system bus (not shown) that connects various system components so that they can transmit and receive data to and from each other (e.g., telecommutatively connected, physically connected, electrically connected), and the components of at least one computing device are not limited thereto and may be very diverse.
[0070] Meanwhile, at least one computing device included in a computing system (10000) that simultaneously performs molecular structure formula detection and chemical reaction formula recognition in a document and performs a method of combining the molecular structure formula detection result and the chemical reaction formula recognition result may be connected to communicate via a network (1070). For example, at least one computing device included in the computing system (10000) may be clustered or may be part of a local area network (LAN). Additionally, at least one computing device may be part of a wide area network (WAN) or connected to at least one of a client-server network and a peer-to-peer network within the cloud.
[0071] Meanwhile, when at least one computing device is used in at least one of a network environment and a cloud computing environment, the at least one computing device may be connected to at least one of a public and private network through a network interface or adapter. In one embodiment, other communication connection devices, such as a modem, may be used to establish communication through the network. The modem may be at least one of an internal modem and an external modem, and may be connected to a system bus through a network interface or a specific mechanism, etc. A wireless network component consisting of an interface and an antenna may be coupled to the network through a device such as an access point, a peer computer, etc. In the present invention, the method of connecting at least one computing device to communicate through the network (1070) is not limited, and it may be connected to communicate in a manner different from the described example.
[0072] Furthermore, other computer-type devices and / or systems not shown in FIG. 1 may also interact technically with at least one computing device or other system through one or more connections to the network (1070) via a network interface. Here, the network interface may include network interface equipment such as a physical network interface controller (NIC) or a virtual network interface (VIF).
[0073] The network (1070) of the present invention may include various forms such as the Internet, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, Wireless USB (Wireless Universal Serial Bus), etc., and in the present invention, data transmission may be performed based on standard communication protocols such as TCP / IP, HTTP, SSL, etc.
[0074] A computing system (10000) that performs a method of simultaneously detecting a molecular structure formula and recognizing a chemical reaction formula in a document according to the present invention and combining the results of the molecular structure formula detection and the chemical reaction formula recognition may include at least one of a user computing device (1010), a training computing system (1050), and a server computing system (1030).
[0075] A user computing device (1010) according to the present invention may be understood as a computing device comprising at least one and / or at least one processor (1011) and memory (1012) that simultaneously perform molecular structure formula detection and chemical reaction formula recognition in a document and combine the molecular structure formula detection result and the chemical reaction formula recognition result. For example, the user computing device (1010) may include at least one computing device among a smartphone, a smart TV, a laptop computer, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, and a head-mounted display).
[0076] At least one and / or at least one processor (1011) constituting the user computing device (1010) may include one or more general-purpose processors and / or one or more special-purpose processors. For example, at least one and / or at least one processor (1011) constituting the user computing device (1010) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), an application integrated circuit, an application semiconductor (ASIC), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.
[0077] Furthermore, at least one and / or at least one processor (1011) may be configured to execute computer-readable instructions contained in memory (1012) and / or other instructions described herein.
[0078] The memory (1012) constituting the user computing device (1010) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media and / or other types of physically durable storage media.
[0079] For example, the memory (1012) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, and combinations thereof, and may include web storage of a server that performs the storage function of memory on the internet. This memory (1012) may store data and instructions necessary for the at least one and / or at least one processor (1011) to simultaneously perform molecular structure formula detection and chemical reaction formula recognition in a document and to perform the operation of an application for combining the molecular structure formula detection result and the chemical reaction formula recognition result.
[0080] A user computing device (1010) may include one or more user input components (1021) that detect user input. For example, the user input component (1021) may also be referred to as a user interface module. The user input component (1021) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of user input component (1021). In this case, the user input component (1021) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user. Meanwhile, the user of the present invention may refer to an automated agent, script, playback software, etc., that operates on behalf of one or more people.
[0081] A user can interact with a computing system (10000) including at least one computing device through input text, touch, voice, movement, computer vision, gestures and / or other forms of input / output using a user input component (1021). For example, the user input component (1021) may include one or more of a command line interface (CLI), a graphical user interface (GUI), a natural user interface (NUI), a voice command interface and / or other user interface (UI) representations.
[0082] Between the user input component (1021) and the user computing device (1010), one or more application programming interface (API) calls may be made based on user input received from the user interface and / or network.
[0083] Here, the expression "based on" may be interpreted to include cases where it is based on the use of a specific configuration, modified from, derived from, influenced by, dependent on, or otherwise derived from a specific configuration. In some embodiments, an API call may be configured for a specific API, which may be interpreted or converted into an API call configured for another API. Here, an API may refer to a defined interface or connection between computers or between computer programs.
[0084] In one embodiment, the user computing device (1010) may store at least one machine learning model (1020). For example, the user computing device (1010) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks) that simultaneously perform molecular structure detection and chemical reaction recognition in a document and combine the results of molecular structure detection and chemical reaction recognition, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.
[0085] According to an embodiment of the present invention, a user computing device (1010) may use a local or / and external machine learning model (1020) to simultaneously perform molecular structure detection and chemical reaction recognition in a document, and perform a method of combining the molecular structure detection result and the chemical reaction recognition result. Alternatively, the user computing device (1010) may use a machine learning model (1040) provided by a server to simultaneously perform molecular structure detection and chemical reaction recognition in a document, and perform a method of combining the molecular structure detection result and the chemical reaction recognition result.
[0086] Additionally, according to another embodiment of the present invention, a server computing system (1030) communicating with a user computing device (1010) may provide at least one output of chemical data included in a document to the user computing device (1010) on an application or / and the web in accordance with a request from a user received through the user computing device (1010).
[0087] In addition, according to another embodiment of the present invention, at least a portion of a user computing device (1010) and a server computing system (1030) are interconnected to simultaneously perform molecular structure formula detection and chemical reaction formula recognition in a document, and by performing a method of combining the molecular structure formula detection result and the chemical reaction formula recognition result, at least one output of chemical data included in the document can be provided to the user.
[0088] Additionally, according to various embodiments of the present invention, a user computing device (1010) and / or a server computing system (1030) can learn a machine learning model (1020, 1040) that is performed in a method of combining molecular structure detection results and chemical reaction recognition results, by simultaneously performing molecular structure detection and chemical reaction recognition in a document through interaction with a training computing system (1050) that is communicated via a network (1070). In this case, the training computing system (1050) may be a computing system separate from the server computing system (1030). Alternatively, in some embodiments, the training computing system (1050) may be part of the server computing system (1030) or part of the user computing device (1010).
[0089] Meanwhile, the server computing system (1030) may include at least one processor (1031) and memory (1032). Here, the processor (1031) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an application integrated circuit, an application semiconductor (ASIC), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions. For example, at least one processor (1031) may include a circuit and a transistor configured to execute instructions from memory (1032).
[0090] The memory (1032) constituting the server computing system (1030) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media, and / or other types of physically durable storage media. For example, the memory (1032) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs the storage function of memory over the internet. Additionally, the server computing system (1030) may further include a data storage (data store). For example, the data storage may be composed of at least one of a relational database, a NoSQL database, a data warehouse, and a local file system.
[0091] In the memory (1032) constituting the server computing system (1030) according to the present invention, data and instructions necessary for the at least one processor (1031) to simultaneously perform molecular structure formula detection and chemical reaction formula recognition in a document and to perform the operation of an application for combining the molecular structure formula detection result and the chemical reaction formula recognition result may be stored.
[0092] In one embodiment, the server computing system (1030) may be composed of a single device or a plurality of computing devices, and these may be configured to operate according to a sequential or parallel computing architecture. Additionally, a distributed processing system may be configured with a plurality of networked devices.
[0093] Meanwhile, the training computing system (1050) may include at least one processor (1051) and memory (1052). The model trainer (1060) is a logical component that executes the training of at least one machine learning model (1020, 1040) and may be implemented in the form of hardware, firmware, or software. For example, the model trainer (1060) may be executed by the processor (1051) after loading training data (1061) stored in a storage device into memory (1052). For example, the model trainer (1060) may be configured to execute one or more operations (e.g., model training, model reconstruction, model validation, model testing) on at least one machine learning model.
[0094] The machine learning model of the present invention may include at least one of a statistical model, an algorithm, a neural network (NN), a convolutional neural network (CNN), a generative neural network (GNN), a Word2Vec model, a Bag of Words model, a TF-IDF (document frequency-inverse document frequency) model, a GPT (Generative Pre-trained Transformer) model (or other autoregressive models), a PPO (Proximal Policy Optimization) model, a nearest neighbor model (e.g., a k-nearest neighbor model), a linear regression model, a K-means clustering model, a Q-learning model, a TD (Temporal Difference) model, a Deep Adversarial Network model, and all other types of models further described herein.
[0095] Specifically, the model trainer (1060) may execute operations to train a machine learning model, and said operations may include at least one of adding, removing, and modifying model parameters. At this time, the training of the machine learning model may be at least one of supervised learning, semi-supervised learning, and unsupervised learning. In one embodiment, the training of the machine learning model may include the step of repeatedly inputting training data (1061) based on epochs and repeatedly performing the machine learning model training process configured in this way. Here, an epoch may refer to a unit in which the entire set of training data (1061) undergoes forward and backpropagation processing once. In some implementations, different levels of training methods (e.g., supervised learning, semi-supervised learning, unsupervised learning) may be used for different epochs.
[0096] The training data (1061) of the present invention may include input data and / or data previously output from at least one machine learning model (e.g., recursive learning feedback).
[0097] At least one parameter of a machine learning model may include at least one of a seed value, a model node, a model layer, an algorithm, a function, connections between different machine learning models, connections between parameters, machine learning model constraints, and other digital components that influence the output of the machine learning model. In this case, model connections between different machine learning models may include or represent relationships between model parameters and / or models, which may be dependent or interdependent, hierarchical, and / or static or dynamic. The combinations and configurations of model parameters described herein may be too complex to be maintained or utilized by human cognitive abilities.
[0098] In the present invention, the machine learning parameters described according to the embodiments are not limited, and a single machine learning model may further include a plurality of model parameters.
[0099] Meanwhile, FIG. 2 illustrates an example of a block diagram of a computing device (1100) that may be included in a user computing device (1010), a server computing system (1030), and a training computing system (1050), as an embodiment of a computing system (10000) in which the present invention can be implemented.
[0100] As illustrated in FIG. 2, the computing device (1100) may include at least one application (e.g., Application 1 to Application N), and each of the at least one application may include a machine learning library and a model execution environment for performing a method of simultaneously detecting molecular structure formulas and recognizing chemical reaction formulas in machine learning-based documents and combining the results of molecular structure formula detection and chemical reaction formula recognition. The at least one application included in the computing device (1100) may communicate with the sensor, context manager, device state manager, or additional component(s) within the computing device (1100) via an Application Programming Interface (API). In one embodiment, the at least one application may interface with device components, such as receiving sensor data or state data or transmitting prediction results to an output device via a public or private API.
[0101] Meanwhile, FIG. 3 illustrates an example of a block diagram in another aspect of a computing device (1200), which is one of the components of a computing system (10000) that performs molecular structure formula detection and chemical reaction formula recognition simultaneously in a document according to an embodiment of the present invention, and combines the molecular structure formula detection result and the chemical reaction formula recognition result.
[0102] A computing device (1200) according to the present invention may include at least one application (e.g., Application 1 to Application N), and at least one application may communicate with a central intelligence layer (1210). Each application may interact with a shared model within the central intelligence layer (1210) through an API (e.g., a common API).
[0103] The central intelligence layer (1210) includes one or more machine learning models and may share them among multiple applications or provide them independently to each. In one embodiment, the central intelligence layer (1210) may be integrated as part of an operating system or implemented as a separate logical layer.
[0104] Additionally, the central intelligence layer (1210) can communicate with the central device data layer (1220). The central device data layer (1220) can integrate and store at least one and / or at least one document containing chemical data (e.g., molecular structure, chemical reaction, etc.) stored within the computing device (1200), and can provide this as input data necessary to simultaneously perform molecular structure detection and chemical reaction recognition in the document, and to combine the molecular structure detection result and the chemical reaction recognition result. Each device component (e.g., sensor, state manager, etc.) can communicate with the central device data layer (1220) through a private API, etc.
[0105] The technology described herein may be composed of a single or multiple computing devices, and a machine learning model that performs molecular structure detection and chemical reaction recognition in a document simultaneously, and combines the molecular structure detection results and the chemical reaction recognition results, may be executed sequentially or in parallel on one component or multiple distributed components. The data storage, machine learning model, and application may be distributed and operated locally or over a network, and these configurations can be flexibly applied to various system architectures.
[0106] Meanwhile, the present invention relates to a document understanding method and system capable of effectively understanding various types of documents. The document understanding system according to the present invention may be a system that provides a service and / or function for automatically converting multimodal information (e.g., layout, text, table, chart, image, graph, chemical molecular structural formula, chemical reaction formula, etc.) from various types of documents into data using Deep Document Understanding (DDU) technology.
[0107] In particular, the present invention aims (or seeks) to maximize the efficiency of document-based chemical information processing and to expand the possibilities of application in research and industrial sites. Below, we will examine the document understanding system according to the present invention in more detail together with the attached drawings. Figures 4, 5, and 6 are conceptual diagrams for explaining the document understanding system according to the present invention.
[0108] In this specification, the terms “chemical information” or “chemical data” refer to a concept including molecular structural formulas, chemical reaction formulas, reactants, products, reaction conditions, information on the relationships between them, or data derived therefrom, and are not limited to specific representation formats or data structures.
[0109] In this specification, “molecular structure information” refers to information related to a molecular structure included in a document subject to analysis, and refers to a concept comprising at least one of a molecular structural formula, an image or region corresponding to the molecular structural formula, atoms constituting the molecular structure and bonding relationships between atoms, structural data representing the molecular structure in the form of a graph, or a chemical structural representation format representing the molecular structure (e.g., SMILES, InChI, Mol, etc.). The molecular structure information is not limited to a specific representation method or data format.
[0110] In this specification, “chemical reaction information” refers to information related to a chemical reaction contained in a document subject to analysis, and includes a concept comprising chemical equations, chemical reaction diagrams, reactants, products, reaction conditions, reaction flow, or information on the relationships between these. The chemical reaction information may be expressed in the form of text, images, structural data, or a combination thereof, and is not limited to a specific method of expression or data format.
[0111] Meanwhile, as illustrated in FIG. 4, the document understanding system (1000) according to the present invention may include at least one of an input unit (100), an output unit (200), a communication unit (300), a storage unit (400), a document understanding unit (500), and a control unit (600). However, the components of the document understanding system (1000) according to the present invention are not limited thereto and may further include various hardware components that perform the same or similar roles as described in the description of the present specification.
[0112] Although not illustrated, the document understanding system (1000) according to the present invention may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processor, tensor processing unit (TPU), graphics processing unit (GPU), neural network processing unit (NPU), application integrated circuit, application semiconductor (ASIC), field programmable gate array (FPGA), quantum processing unit (or quantum processor, QPU), etc.). One or more processors may be configured to execute instructions, computer-readable instructions, and / or other instructions described herein that are stored (or included) in the storage unit (400). The document understanding method and system according to the present invention may perform data processing described below in cooperation with memory and at least one processor. The processor may perform a series of operations and data processing using data and information stored in memory. In this case, memory may be a component of the storage unit (400).
[0113] In addition, the document understanding system (1000) according to the present invention can perform data processing and computation processes using quantum gates, quantum entanglement, and quantum superposition states, taking into consideration implementation in a quantum computer environment. For example, the present invention can perform parallel computations based on qubits, and such quantum computations can operate complementarily with existing classical computers.
[0114] Such quantum computers may include parallel computation using qubits and high-speed data processing devices utilizing quantum entanglement, and hardware-based computational optimization using FPGAs and ASICs is possible. In addition, quantum computers may utilize quantum processors capable of qubit-based parallel computation, and data processing efficiency can be improved through a hybrid structure with existing classical computers.
[0115] Meanwhile, the input unit (100) can be configured in various ways as a means of data input. For example, the input unit (100) can be configured to receive user input. The input unit (100) can be configured to receive user input from a user terminal (10). Here, “receiving input” may mean receiving an input signal (or selection signal) corresponding to the user’s input based on input made by the user through the input unit configuration provided in the user terminal (10).
[0116] Here, the user terminal (10) may include at least one of a mobile phone, a smartphone, a notebook computer, a laptop computer, a slate PC, a tablet PC, an ultrabook, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, and a wearable device (e.g., a smartwatch, a smart glass, a head-mounted display).
[0117] In addition, the input unit (100) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user.
[0118] The input unit (100) may also be referred to as a user interface module. The input unit (100) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of input unit (100).
[0119] Here, user input may include documents, text, images (or videos), voice, etc. In this case, the document understanding system (1000) may further include a module that converts voice into text.
[0120] Next, the output unit (200) can output information through an output unit configuration (e.g., a display unit, a touch screen, a speaker, etc.) provided in a user terminal (10) linked to the document understanding system (1000) according to the present invention. For example, the output unit (200) can output at least one page (2000, or service page) linked to the document understanding system (1000) according to the present invention to the display unit of the user terminal (10). Additionally, the output unit (200) does not necessarily mean a hardware means, but can be understood as a channel for outputting results to a user.
[0121] Next, the communication unit (300) may be connected via a wireless or wired network to a user terminal (10), a server (e.g., a central server, an external server, etc.), a device, and at least one network, etc., to receive or transmit overall data and information necessary for the operation of the document understanding system (1000) according to the present invention.
[0122] The communication unit (300) can support various communication methods depending on the communication standard of the communicating device.
[0123] For example, the communication unit (300) may be configured to communicate with a communication target using at least one of the following technologies: WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ RFID (Radio Frequency Identification), Infrared Communication (Infrared Data Association; IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus).
[0124] Next, the storage unit (400, or memory) serves to store various data related to the present invention and may include one or more non-transient computer-readable storage media that can be read and / or accessed by at least one of one or more processors.
[0125] One or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the storage unit (400) may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage device), whereas in other examples, the storage unit (400) may be implemented using two or more physical devices.
[0126] The storage unit (400) may include computer-readable instructions and additional data. The storage unit (400) may include a storage necessary to perform at least some of the methods, scenarios, and techniques described herein and / or at least some of the functions of the device and network.
[0127] Furthermore, at least a portion of the storage unit (400) may be a cloud storage or a cloud server. At least a portion of the data corresponding to user input received from the input unit (100) and the training data may be stored in the storage unit (400).
[0128] Additionally, the storage unit (400) may store at least one document collected (or received) from various sources (e.g., a web corpus, a document corpus, a database (DB) website, an API, a server linked to the document understanding system (1000), a central server, an external server, cloud storage, a user terminal (10), a large dataset, etc.). For example, the storage unit (400) may store at least one and / or at least one document collected from at least one of the various sources (e.g., a user terminal (10)). In this case, the collected document may include documents related to at least one and / or at least one domain (or field). Alternatively, the collected document may include documents containing at least one and / or at least one content.
[0129] That is, the storage unit (400) is sufficient as a space where information necessary for the operation of the document understanding system (1000) according to the present invention is stored, and it can be understood that there are no restrictions on the physical space.
[0130] Furthermore, the storage unit (400) can store a computer program including computer program instructions. Furthermore, the storage unit (400) can store a computer program including computer program instructions that control the operation of the system (1000) or control the operation of the control unit (600) when loaded into the processor of the system (1000).
[0131] Next, the document understanding unit (500) may be configured to effectively understand complex forms of information contained in various types of documents. The document understanding unit (500) may be configured to recognize the structure of the document to be analyzed (20) and to recognize the relationships of the content (e.g., Image-Text, Image-Image, Text-Text) contained in the document to be analyzed (20) to perform the role of extracting information. In the present invention, the document understanding unit (500) may also be named a “deep document understanding model,” a “document understanding model,” or a “DDU model.”
[0132] The document comprehension unit (500) may be configured to extract various forms of content (e.g., layout, text, table, chart, image, graph, chemical molecular structural formula, chemical reaction formula, etc.) from at least one and / or at least one document (e.g., paper, book, patent document, report, etc.). Here, the various forms of content may also be understood as multimodal information.
[0133] More specifically, the document comprehension unit (500) may be a model trained to understand structured data, unstructured data, linguistic data (or linguistic elements) and non-linguistic data (or non-linguistic elements), etc., included in the document (20) to be analyzed, and to extract various content and / or knowledge based on the understood content.
[0134] In one embodiment, the document understanding unit (500) can understand the chemical structure of a molecular structure formula included in the document to be analyzed (20), and based on the result of understanding, convert the molecular structure formula into a SMILES string expression and extract it. Additionally, the document understanding unit (500) can understand the chemical structure of the molecular structure formula and perform a graph conversion corresponding to the molecular structure formula based on the result of understanding.
[0135] In another embodiment, the document understanding unit (500) can understand texts related to molecular structural formulas among the texts included in the document to be analyzed (20) and extract them as text data related to said molecular structural formulas.
[0136] In another embodiment, the document comprehension unit (500) can recognize rows and columns constituting a table associated with molecular structural formulas from the document to be analyzed (20), and convert them into structured data in a format such as HTML or Excel to extract them.
[0137] Additionally, the document understanding unit (500) can extract relationship information (or relationships) between molecular structures included in the document to be analyzed (20).
[0138] In one embodiment, the document understanding unit (500) can understand the relationship between the first molecular structure and the second molecular structure included in the document to be analyzed (20), and extract relationship information in which a third molecular structure is generated through a chemical reaction between the first molecular structure and the second molecular structure. In this case, the relationship information between the molecular structures can be extracted by understanding the text included in the document to be analyzed (20) or by understanding the non-verbal data included in the document to be analyzed (20).
[0139] In another embodiment, the document understanding unit (500) can understand the relationship between the first molecular structure and the second molecular structure through a symbol (e.g., plus sign, arrow, etc.) located in one region among a plurality of regions included in the document (20) to be analyzed, and can extract relationship information in which a third molecular structure is generated through a chemical reaction between the first molecular structure and the second molecular structure.
[0140] Furthermore, the document comprehension unit (500) can extract various forms of content satisfying established content criteria from at least one and / or at least one document. Here, the established content criteria can be set in various ways and can be determined according to the purpose or use of the document comprehension system (1000). For example, if the purpose of use of the document comprehension system (1000) is chemistry, bio, new materials, new substances, and new drug development, the document comprehension unit (500) can be trained to understand and extract content related to chemistry, bio, new materials, new substances, and new drug development from the document to be analyzed (20). In this case, the established content criteria may include content related to molecular structures related to at least one of chemistry, bio, new materials, new substances, and new drug development. Here, the document comprehension unit (500) can extract content related to chemistry, bio, new materials, new substances, and new drug development from the document to be analyzed (20) according to the established content criteria. However, this is merely one embodiment, and the established content standards in the present invention are not necessarily limited thereto.
[0141] Meanwhile, the document understanding unit (500) may include at least one and / or at least one artificial intelligence model (or model, module) to properly process chemical data (or chemical information) included in the document (20) to be analyzed.
[0142] For example, as illustrated in FIG. 5, the document comprehension unit (500) may include a chemical information processing unit (510) configured to perform the role of processing chemical data included in the document (20) to be analyzed.
[0143] The chemical information processing unit (510) may include at least one information processing model specialized for processing chemical data included in the document (20) to be analyzed. In the present invention, the “at least one information processing model” for processing chemical data may be implemented as a single integrated model included in the chemical information processing unit (510), or as a plurality of distinct models. For example, processing of molecular structure information and chemical reaction information may be performed through different processing modules, processing paths, or output heads within a single information processing model, but is not limited thereto. Additionally, a single information processing model may be configured to process molecular structure information and chemical reaction information sequentially or in parallel. Accordingly, the present invention is not limited by the number, internal structure, processing method, or implementation form of the information processing models for processing chemical data. In this specification, “information processing model” refers to a computing-based processing means configured to receive data included in the document to be analyzed as input, perform computation, inference, or analysis, and output a processing result. The above information processing model may be implemented as an artificial intelligence model, a machine learning model, a deep learning model, a rule-based processing module, or a combination thereof, and is not limited to a specific implementation method. In addition, “at least one information processing model” should be understood as a concept that includes one information processing model or multiple information processing models.
[0144] Accordingly, the term “at least one information processing model” in this specification should be understood as a concept that includes a single model, a plurality of processing modules included in a single model, or a plurality of individual models. That is, the information processing model may be implemented as a single configuration or may be implemented in a form including a plurality of components, and all such implementation forms are included within the scope of the present invention.
[0145] As previously mentioned, when the above-mentioned at least one information processing model is implemented as a single information processing model (or a single model), the information processing model may internally include different processing paths, functional modules, or output heads to process molecular structure information and chemical reaction information simultaneously or selectively. For example, the information processing model may be configured to be trained using a multitask learning method to process molecular structure information and chemical reaction information separately, or to process each piece of information through a common feature extraction unit and branched processing units. Such internal branching structures are not limited to a single one.
[0146] When the document comprehension unit (500) receives the document to be analyzed (20), it can process the received document to be analyzed (20) as input to the chemical information processing unit (510, or a plurality of models (511, 512)). At this time, the document comprehension unit (500) can check the extension of the document to be analyzed (e.g., PDF, PPTX, HWPX, PNG, JPG, etc.) and convert the document to be analyzed (20) into a specific format that has been set. For example, the specific format that has been set (or format, form, etc.) may include an image. This may be to convert the document to be analyzed (20) into a specific format to generate (or obtain) a document to be analyzed of a specific format (or a document to be analyzed corresponding to a specific format, a document to be analyzed having a specific format, an image of a document to be analyzed, an image of a document to be analyzed corresponding to a document to be analyzed, etc.). Alternatively, at least one and / or at least one page included in (or constituting) the document to be analyzed (20) may be converted into an image. In this case, it may be understood that at least one page image corresponding to at least one page included in the document to be analyzed (20) is generated. In the present invention, the document to be analyzed may include not only PDF and Word documents, but also documents written based on LaTeX or documents generated by compiling LaTeX documents. As such, the document to be analyzed is not limited to PDF, word processor documents, or image documents, but may include documents written based on a markup language such as LaTeX or documents generated by compiling such documents. LaTeX-based documents may explicitly include formulas, text, images, and document structure information, and particularly in the field of chemistry, molecular structures or chemical reactions are often expressed in formulas, diagrams, or structured forms.Accordingly, the document understanding method and system according to the present invention can also be applied to processing molecular structure information and chemical reaction information contained in LaTeX-based documents and generating output data that reflects the interrelationships between them.
[0147] The chemical information processing unit (510) can perform inference in parallel on a document to be analyzed (20, or a document to be analyzed converted into a specific format) using a plurality of models (511, 512). In the present invention, this can be expressed as “performing parallel inference.” That is, in the present invention, the plurality of models (511, 512) can operate to perform parallel inference on the document to be analyzed (20) and output inference results according to the results of the parallel inference.
[0148] In this specification, “parallel inference” or “parallel execution” means that multiple information processing tasks are performed such that they overlap in time, and this does not necessarily require them to start or end simultaneously. Furthermore, parallel inference is optionally interchangeable with sequential execution and is not limited to a specific execution method.
[0149] In the present invention, the processing of molecular structure information and the processing of chemical reaction information may be performed in parallel, but are not limited thereto and may also be performed sequentially. That is, the execution order or parallelism of each processing step may be selectively determined according to the system environment, implementation method, or processing efficiency, and is not limited to a specific execution method.
[0150] In this specification, the term “inference result” refers to data generated as a result of an information processing model performing inference on input data, and the inference result may include detection results, recognition results, classification results, location information, structural data, or a combination thereof. The inference result may be used as an intermediate processing result or a final processing result and is not limited to a specific processing step or result format.
[0151] To this end, the chemical information processing unit (510) can process an analysis target document converted into a specific format as input to a plurality of models (511, 512). For example, as illustrated in FIGS. 5 and 6, the chemical information processing unit (510) can input an analysis target document converted into a specific format (e.g., “Document Image”, 21) to a first model (e.g., “Molecule Detection”, 511) and a second model (e.g., “Reaction Parsing”, 512), respectively. However, in the present invention, the first model (511) and the second model (512) may also be implemented as separate configurations from the chemical information processing unit (510) (e.g., configuration of a document understanding system (1000) or configuration of a document understanding unit (500)).
[0152] First, the first model (511, or molecular structure formula detection model) may be a model trained to perform molecular detection by identifying molecular structure formulas contained in a document (or a page of a document) and determining the location of the molecular structure formulas. This first model (511) may be a model trained with at least one training data set (e.g., a first training data set specialized for molecular structure formula detection) so as to accurately detect molecular structure formulas (or regions corresponding to molecular structure formulas) from a document. In this way, when the converted document to be analyzed is provided as an input to the chemical information processing unit (510), the first model (511) can perform processing on the molecular structure information contained in the document to be analyzed. For example, the molecular structure contained in the LaTeX document may be expressed in the form of a formula, a diagram, or an image, and the first model can detect regions corresponding to molecular structure formulas from these forms of expression, and analyze the bonding relationships between atoms constituting the detected molecular structure to generate a first processing result for the molecular structure information.
[0153] The training data used for training the first and second models may be completely separate datasets, partially overlapping datasets, or trained based on the same dataset. Furthermore, the training data is not limited to a specific format or source and may include various forms of data related to molecular structure information and chemical reaction information. Accordingly, the present invention is not limited by the composition or scope of the training data.
[0154] In the parallel inference process, the first model (511) may perform inference (or first inference) to detect at least one and / or at least one molecular structure formula (or molecular structure) contained in the document (21) to be analyzed converted into a specific format. Then, the first model (511) may output a first inference result related to the at least one and / or at least one molecular structure formula detected from the document (21) to be analyzed converted into a specific format. In this specification, “performing inference to detect a molecular structure formula contained in the document to be analyzed converted into a specific format” may also be understood to mean “performing inference to detect a molecular structure formula contained in the document to be analyzed.”
[0155] In this case, the process of detecting a molecular structural formula performed by the first model (511) may include a task of detecting (or detecting) a molecular structural formula in a document (21) converted into a specific format. In this specification, “detecting a molecular structural formula in a document converted into a specific format” may also be understood to mean “detecting a molecular structural formula in a document to be analyzed.”
[0156] In addition, the process of detecting the molecular structural formula may include a task of analyzing the atoms constituting the detected molecular structural formula and the bonding relationships between the atoms, and converting at least one molecular structural formula into structural data based on the analysis results.
[0157] In one embodiment, the structural data may include data of a specific form (or format, etc.) in which atoms constituting the molecular structural formula are represented as nodes and the bonding relationships between said atoms are represented as edges. The data of the specific form may be a molecular graph for (or corresponding to) the molecular structural formula. Such a molecular graph may include the atoms constituting the molecular structural formula and the connection relationships of said atomic bonds. Alternatively, the structural data may be data obtained by converting (or expressing) a molecular graph for the molecular structural formula into at least one chemical structure representation format (e.g., SMILES, InChI, Mol, etc.).
[0158] Finally, the first inference result output from the first model (511) may include at least one and / or at least one detection result of a molecular structural formula. For example, the detection result of the molecular structural formula may include at least one of i) a detected molecular structural formula (e.g., all molecular structural formulas present on the page of the document to be analyzed (20), regardless of whether before or after the reaction), ii) location information of the detected molecular structural formula (e.g., coordinates (bounding box) of the molecular structural formula present (or included) within the page of the document to be analyzed (20)), and iii) at least one area of the document to be analyzed (20) containing the detected molecular structural formula (e.g., an area corresponding to the molecular structural formula including bond lines, atomic symbols, etc.). However, in addition to what is mentioned above, the detection result of a molecular structural formula may further include at least one of the following: a molecular graph (or corresponding to) the detected molecular structural formula (e.g., graphic structures such as bond lines, atomic symbols (C, N, O, etc.), ring structures, etc.); an identification ID or index of the detected molecular structural formula; a chemical structure representation format for the detected molecular structural formula; and at least one visual object separated into at least one molecular structural formula unit (individual molecular unit image). That is, the detection result of a molecular structural formula may include various information indicating what the molecular structural formula is that exists in the document and where it is located.
[0159] In this specification, the detection result, inference result, or combination thereof output by the first model (511) may be understood as an embodiment of the “first processing result” for molecular structure information defined in the claims. That is, the detection result of the molecular structure formula, location information, structural data, or combination thereof may all correspond to the processing result for molecular structure information and are not limited to a specific result form. Next, the second model (512, or chemical reaction formula recognition model) may be a model trained to perform reaction parsing (or reaction parsing) to extract and classify chemical reaction formulas contained in a document (or page of a document). The second model (512) may be a model trained with at least one training dataset (e.g., a second training dataset specialized for learning chemical reaction formula recognition) to extract and classify reaction roles within a chemical reaction formula (or chemical reaction diagram). In the parallel inference process, the second model (512) may perform inference (or second inference) to recognize at least one and / or at least one chemical reaction formula (Chemical Reaction Formula) included in the document (21) to be analyzed converted into a specific format. The second model (512) may output a second inference result related to at least one and / or at least one chemical reaction formula recognized from the document (21) to be analyzed converted into a specific format. In this specification, “performing inference to recognize a chemical reaction formula included in the document to be analyzed converted into a specific format” may also be understood to mean “performing inference to recognize a chemical reaction formula included in the document to be analyzed.”
[0160] In this way, the second model (512) can perform processing on the chemical reaction information included in the document to be analyzed. For example, the chemical reaction may be included in the LaTeX document in the form of a formula or a chemical reaction diagram, and the second model can recognize the reactants, products, or reaction conditions constituting the chemical reaction to generate a second processing result for the chemical reaction information.
[0161] In this case, the process of recognizing a chemical reaction equation performed by the second model (512) may include the task of recognizing (or extracting) the reactants, reaction conditions, and products constituting the chemical reaction equation from the document to be analyzed (21) converted into a specific format. In this specification, “recognizing a chemical reaction equation in a document converted into a specific format” may also be understood to mean “recognizing a chemical reaction equation in the document to be analyzed.”
[0162] Additionally, the process of recognizing the chemical reaction equation may include a task of extracting reactants, reaction conditions, and products from a chemical reaction diagram included in an analysis target document (21) converted into a specific format. This chemical reaction diagram is a diagrammatic representation of reactants, reaction conditions, and products using arrows and / or symbols, visually representing the transformation relationships between molecules and the reaction flow so that they can be understood at a glance. This is used to convey structural information about the chemical reaction in the document.
[0163] Finally, the second inference result output from the second model (512) may include at least one and / or at least one recognition result of a chemical reaction equation. For example, the recognition result of the chemical reaction equation may include at least one of i) information of the components constituting the recognized chemical reaction equation (e.g., reactants (molecular structures introduced into the reaction or regions of said molecular structures), products (molecular structures produced as a result of the reaction), reaction conditions (condition information in text form such as reagents, catalysts, solvents, temperature, time, etc.), ii) location information of each of the above components (e.g., coordinates (bounding box) of each of the reactants, reaction conditions, and products existing (or included) within the page of the document (20) to be analyzed), iii) at least one region of the document (20) to be analyzed containing the recognized chemical reaction equation (e.g., regions corresponding to the reactants, reaction conditions, products, etc.), and iv) information on the reaction relationship between the components (e.g., arrows, directions, steps, sequence relationships, or reaction flow information based on arrows, etc.). However, in addition to what is mentioned above, the result of recognizing the chemical equation may further include at least one of the following: a structured output of chemical reaction units (e.g., an output in which at least one chemical reaction is organized into a “reactant - reaction condition - product” structure), or a result of classifying reaction units (e.g., a result in which several independent reactions within the document (20) under analysis are each classified). That is, the result of recognizing the chemical equation may include various information indicating what reaction relationship (reactant - reaction condition - product) specific molecules form.
[0164] In this specification, the recognition result, inference result, or combination thereof output by the second model (512) may be understood as an embodiment of the “second processing result” for chemical reaction information defined in the claim. That is, component information, location information, reaction relationship information, or combinations thereof of a chemical reaction equation may all correspond to the processing result for chemical reaction information and are not limited to a specific result form. Furthermore, the chemical information processing unit (510) may combine (or integrate) the first inference result output from the first model (511) and the second inference result output from the second model (512) to generate at least one output (e.g., “Combined Result”, 22) for chemical data (e.g., molecular structure formula, chemical reaction formula, etc.) included in the document to be analyzed (20).
[0165] More specifically, the chemical information processing unit (510) can generate the at least one output (22) by combining the molecular structure formula detection result and the chemical reaction diagram recognition result. In the present invention, this can be expressed as a “post-processing process (or step).” That is, the post-processing process may include a process of generating at least one output (22) by combining the first inference result (molecular structure formula detection result) and the second inference result (chemical reaction formula recognition result). More specific details regarding this will be described later.
[0166] In this specification, “output data” refers to data generated based on the first processing result and the second processing result, in a form that can be stored, reprocessed, or used as input to other information processing steps by a computer. The output data is not limited to data for visual display and should be understood as machine-readable and processable data.
[0167] The interrelationship between the first processing result and the second processing result described above may be determined based on objectively identifiable data, such as positional relationships within the document, connectivity relationships on molecular graphs, correspondence relationships between reactants and products included in chemical equations, and reaction flow information expressed by arrows or symbols, rather than relying on subjective interpretation of meaning. Such criteria for determining correlation are not limited to a single criterion, and multiple criteria may be applied individually or in combination.
[0168] Accordingly, the combination of the first and second processing results may be performed not merely by simple merging or listing, but by considering the structural, functional, or semantic associations between molecular structure information and chemical reaction information. For example, methods may be applied to distinguish between molecular structural formulas included in a specific chemical reaction equation and those not included, or to combine processing results by reflecting the relationship between reactants and products, but are not limited thereto.
[0169] The determination of the correlation between the first processing result and the second processing result described above can be implemented through technical processing that can be performed by a computer. For example, at least one of the following may be applied: distance calculation based on the relative positional relationship between a molecular structural formula and a chemical reaction equation; graph matching between a molecular structure graph and a reaction equation component; analysis of the correspondence relationship between the reactants and products included in the chemical reaction equation and the molecular structural formula; or a method of scoring the correlation by assigning weights to multiple judgment criteria. Such a method for determining correlation is not limited to a single method, and multiple methods may be applied individually or in combination.
[0170] Meanwhile, in the present invention, the process of generating output data by combining a first processing result and a second processing result can be configured to enable the generation of output data even if some molecular structure information or chemical reaction information is missing or incomplete. For example, molecular structure information that does not correspond to a chemical reaction equation may be maintained as independent output data, and partial combination or individual output may be performed even if chemical reaction information corresponding to the molecular structure information does not exist. Furthermore, if conflicting information exists between the first processing result and the second processing result, output data may be generated according to a preset priority rule, reliability criterion, or subsequent verification procedure.
[0171] Meanwhile, in the present invention, the output generated as a result of the combination process is not limited to a result simply displayed on a screen, but can be generated in the form of “output data” that reflects chemical information contained in the document to be analyzed. Such output data may be in the form of structured data, unstructured data, or a combination thereof, and is not limited to a specific data format or method of representation.
[0172] That is, the output data generated by the above combination process is not limited to results merely displayed visually to the user, but can be generated in a data form that can be stored, reprocessed, retrieved, or utilized as input for other information processing stages by a computer. For example, the output data can be used for subsequent analysis, statistical processing, search indexing, or integration with other systems, and accordingly, can be understood as machine-readable and processable data.
[0173] The output data may include a plurality of data items composed of units of molecular structure information or chemical reaction information, and each data item may consist of information representing a molecular structural formula, a chemical reaction formula, or the relationship between them. For example, the output data may include at least one of a record of a molecular structure unit, a record of a chemical reaction unit, or mapping information representing the correspondence between a molecular structure and a chemical reaction. The internal configuration of such output data is not limited to a specific form and can be implemented in various data structures.
[0174] Meanwhile, in this specification, “first processing result” and “second processing result” are higher-level concepts referring to the results of processing molecular structure information and chemical reaction information, respectively, and encompass results generated by detection, recognition, inference, or a combination thereof performed by the first model or the second model. That is, detection results of molecular structural formulas, recognition results of chemical reaction formulas, inference results, location information, structural data, or combinations thereof may all be included in the first processing result or the second processing result, and are not limited to a specific processing method or result form. In this specification, “interrelationship” refers to a relationship existing between molecular structure information and chemical reaction information, and means a relationship determined based on objectively identifiable data, such as positional relationships within a document, correspondence relationships between reactants and products, chemical reaction flow, and connectivity relationships on graph structures. The determination of the interrelationship may be performed automatically according to predefined rules, formulas, or judgment logic. For example, the association between molecular structure information and chemical reaction information can be determined based on whether the result of calculating the distance between relative positions within a document satisfies a threshold, the correspondence between nodes or edges on a molecular graph, or whether there is a match between the reactants and products in a chemical equation and the molecular structure information. This judgment logic is not limited to a single one, and multiple logics may be applied sequentially or in parallel.
[0175] Furthermore, the aforementioned correlation determination is not limited to being performed once at a single point in time, but may be performed repeatedly as processing results regarding molecular structure information or chemical reaction information are sequentially generated or updated. For example, if new molecular structure information or chemical reaction information is additionally detected or recognized, the correlation determination may be re-performed based on previously generated output data to update the output data.
[0176] In this specification, “combination” does not mean simply merging or listing the first processing result and the second processing result, but rather means a processing process that generates new output data by considering the interrelationships between the processing results.
[0177] In this way, the chemical information processing unit (510) can determine the correlation between molecular structure information and chemical reaction information using the first processing result and the second processing result, and generate output data that reflects the correlation. For example, molecular structure information placed adjacent to a specific chemical reaction formula within a LaTeX document, molecular structure information included within the same formula environment or the same paragraph structure, can be determined as information associated with the corresponding chemical reaction.
[0178] Meanwhile, as described above, the present invention provides a document understanding system (1000) that automatically converts multimodal information from various types of documents into data using Deep Document Understanding (DDU) technology. In particular, the document understanding system (1000) according to the present invention can simultaneously perform molecular structure formula detection and chemical reaction formula recognition in a document, and combine the molecular structure formula detection result and the chemical reaction formula recognition result to generate a final output.
[0179] As previously discussed, the output data can be utilized for various subsequent processes, such as storage, retrieval, visualization, further analysis, and integration with other systems, and can be used in various forms depending on the purpose or method of utilization. The present invention is not limited by the purpose of utilization, the environment of utilization, or the field of application of the output data.
[0180] The above output data can be provided as input data to other information processing models or systems and utilized as base data for performing additional information processing, analysis, or inference. Accordingly, the output data according to the present invention is not limited to a simple result of providing information but can function as data that can be reused in subsequent processing steps.
[0181] The above output data is data reflecting chemical information contained in a LaTeX document, and can be stored, searched, or utilized for subsequent analysis in an associated form with molecular structure information and chemical reaction information. Accordingly, the document understanding method according to the present invention can also be applied to automatically process chemical information contained in LaTeX-based academic documents, technical documents, or patent documents.
[0182] In this regard, the document understanding method performed by the document understanding system (1000) according to the present invention will be examined in more detail below, together with the attached drawings. FIG. 7 is a flowchart for explaining the document understanding method according to the present invention, and FIGS. 8, FIGS. 9, FIGS. 10, FIGS. 11, FIGS. 12, FIGS. 13 and FIGS. 14 are conceptual diagrams for explaining the document understanding method according to the present invention.
[0183] Meanwhile, as illustrated in FIG. 7, the document understanding method according to the present invention may include a process of specifying at least one document to be analyzed (S710); a process of processing the document to be analyzed as input to a plurality of models specialized in a chemical domain (S720); a process of detecting at least one molecular structural formula included in the document to be analyzed in a first model among the plurality of models and outputting a detection result of at least one molecular structural formula (S730); a process of recognizing at least one chemical reaction formula included in the document to be analyzed in a second model among the plurality of models and outputting a recognition result of at least one chemical reaction formula (S740); and a process of generating at least one output for chemical data included in the document to be analyzed using the detection result of at least one molecular structural formula and the recognition result of at least one chemical reaction formula (S750).
[0184] In one embodiment, the document to be analyzed may be a document written based on a markup language such as LaTeX. For example, the document to be analyzed may be a LaTeX source file or a document generated by compiling a LaTeX source file, and the compiled document may include a document in PDF format or an image format. Such a LaTeX-based document may be composed in a form in which text, formulas, images, and document structure information are mixed.
[0185] The document understanding method according to the present invention may include a step of converting a document to be analyzed into a specific format that has been set, and the document converted into the specific format may subsequently be used as an input for processing molecular structure information and chemical reaction information. This format conversion step may be performed as part of the document understanding method, and subsequent processing steps may be performed using the converted document.
[0186] In this specification, “first processing result” and “second processing result” refer to the results of information processing performed on molecular structure information and chemical reaction information, respectively, and may include detection results, recognition results, inference results, location information, structural data, or combinations thereof. The processing result is not limited to a specific processing method or result format.
[0187] In this specification, the phrase “detecting at least one molecular structural formula” may be understood to mean “detecting one (i.e., singular (1)) molecular structural formula” or “detecting multiple (i.e., multiple (2, 3, 4 or more)) molecular structural formulas.” In this case, “detecting result of at least one molecular structural formula” may be understood as an expression including “detecting result of one molecular structural formula” or “detecting result of each of the multiple molecular structural formulas.”
[0188] Additionally, in this specification, the phrase “recognizing at least one chemical equation” may be understood to mean “recognizing one (i.e., singular (1)) chemical equation” or “recognizing multiple (i.e., multiple (2, 3, 4 or more)) chemical equations.” In this case, “result of recognition of at least one chemical equation” may be understood as an expression including “result of recognition of one chemical equation” or “result of recognition of each of multiple chemical equations.”
[0189] Furthermore, in the present invention, the process of detecting at least one molecular structural formula included in the document to be analyzed in the first model and outputting the detection result of at least one molecular structural formula (S730), and the process of recognizing at least one chemical reaction formula included in the document to be analyzed in the second model and outputting the recognition result of at least one chemical reaction formula (S740) may be processes performed in parallel. Here, "performed in parallel" may mean that the detection of a molecular structural formula using the first model and the recognition of a chemical reaction formula using the second model are performed independently of each other and simultaneously. This can also be understood as the "parallel inference" method discussed above.
[0190] Meanwhile, the document understanding system (1000) can identify at least one document to be analyzed. In the present invention, there may be various ways in which the document to be analyzed is identified.
[0191] For example, the document understanding system (1000) can identify at least one document received from the user terminal (10) as a document to be analyzed. As illustrated in FIG. 8, the document understanding system (1000) can receive a document (800) corresponding to user input based on the fact that at least one document (800) is entered into a document input area included in the service page (2000) from the user terminal (10). In this case, the document understanding system (1000) can identify the received document (800) as a document to be analyzed. Hereinafter, the received document (800) will be described by referring to it as the “document to be analyzed (800).”
[0192] Additionally, the document understanding system (1000) can process a specific document to be analyzed (800) as input to the document understanding unit (500). The document understanding unit (500) can check the extension of the document to be analyzed (e.g., PDF, PPTX, HWPX, PNG, JPG, etc.) and convert the document to be analyzed (800) into a specific format that is pre-set. For example, the specific format that is pre-set (or format, form, etc.) may include an image. This may involve converting the document to be analyzed (800) into a specific format to generate (or obtain) a document to be analyzed of a specific format (or a document to be analyzed corresponding to a specific format, a document to be analyzed having a specific format, an image of the document to be analyzed, an image of the document to be analyzed corresponding to the document to be analyzed, etc.).
[0193] Meanwhile, if the document to be analyzed is a LaTeX source file, the document comprehension unit (500) can convert the LaTeX source file into a pre-set specific format. For example, the specific format may include an image format, and the conversion process may include a process of converting a document generated by compiling the LaTeX source file into an image form. In this case, at least one page included in the LaTeX document may be converted into a corresponding page image.
[0194] Alternatively, at least one and / or at least one page included in (or constituting the document to be analyzed (800)) may be converted into an image. In this case, it may be understood that at least one page image corresponding to at least one page included in the document to be analyzed (800) is generated. For example, let us assume that the document to be analyzed (800) includes multiple pages. As illustrated in FIG. 9 (a) and (b), the document to be analyzed converted into a specific format may include multiple page images (810, 820) corresponding to each of the multiple pages included in the document to be analyzed (800). In this case, among the multiple page images (810, 820), the first page image (810) may correspond to the first page of the document to be analyzed (800), and the second page image (820) may correspond to the second page of the document to be analyzed (800).
[0195] Subsequently, the document comprehension unit (500) can process chemical data (e.g., molecular structural formulas, chemical reaction formulas, etc.) related to the chemical domain contained in the document to be analyzed (800). More specifically, the document comprehension unit (500) can input the document to be analyzed, converted into a specific format, into the chemical information processing unit (510) to process the chemical data contained in the document to be analyzed (800).
[0196] In this case, the chemical information processing unit (510) can process the document to be analyzed, converted into a specific format, as input to a plurality of models (511, 512). For example, the chemical information processing unit (510) can input a first page image (810) corresponding to the first page of the document to be analyzed (800) and a second page image (820) corresponding to the second page of the document to be analyzed (800) into the first model (511) and the second model (512).
[0197] Here, the first model (511) may be a model specialized in molecular structure formula processing (e.g., molecular structure formula detection), and the second model (512) may be a model specialized in chemical reaction formula processing (e.g., chemical reaction formula recognition). In the present invention, the molecular structure formula detection process using the first model (511) and the chemical reaction formula recognition process using the second model (512) can be performed in parallel.
[0198] Specifically, the first model (511) may be configured to analyze at least one page included in the document (800) to be analyzed to detect at least one molecular structural formula and to output the detection result of at least one molecular structural formula detected in at least one page. Here, “detecting a molecular structural formula” can also be understood as detecting a region corresponding to a molecular structural formula. Additionally, the phrase “detecting a molecular structural formula in a page image” in this specification may also be understood as detecting a molecular structural formula in a page. For example, as illustrated in FIG. 10 (a), the first model (511) can detect a plurality of molecular structural formulas (811, 812, 813, 814, 815, 816, 817, 818, 819) in a first page image (810), and output the detection results for each of the plurality of molecular structural formulas (811, 812, 813, 814, 815, 816, 817, 818, 819) detected in the first page image (810). In this case, the detection result of each of the plurality of molecular structural formulas (811, 812, 813, 814, 815, 816, 817, 818, 819) includes i) the plurality of molecular structural formulas (811, 812, 813, 814, 815, 816, 817, 818, 819) detected in the first page image (810), ii) the location information of each of the plurality of molecular structural formulas (811, 812, 813, 814, 815, 816, 817, 818, 819) detected in the first page image (810) (e.g., the coordinates (bounding box) of each of the plurality of molecular structural formulas existing within the first page image (810)), and iii) the plurality of molecular structural formulas (811, It may include at least one of a plurality of regions (e.g., regions corresponding to each of the plurality of molecular structural formulas including bond lines, atomic symbols, etc.) each including 812, 813, 814, 815, 816, 817, 818, 819).
[0199] Additionally, as illustrated in FIG. 10 (b), the first model (511) can detect a plurality of molecular structural formulas (821, 822, 823, 824, 825, 826, 827, 828, 829) in the second page image (820) and output a detection result for each of the plurality of molecular structural formulas (821, 822, 823, 824, 825, 826, 827, 828, 829) detected in the second page image (820). In this case, the detection result of each of the plurality of molecular structural formulas (821, 822, 823, 824, 825, 826, 827, 828, 829) includes i) location information (e.g., the coordinates of each of the plurality of molecular structural formulas existing within the second page image (820)) of each of the plurality of molecular structural formulas (821, 822, 823, 824, 825, 826, 827, 828, 829) detected in the second page image (820), and iii) the plurality of molecular structural formulas (821, 822, 823, 824, 825, 826, 827, 828, 829) detected in the second page image (820). It may include at least one of the regions each containing 824, 825, 826, 827, 828, 829 (e.g., regions corresponding to each of the plurality of molecular structural formulas including bond lines, atomic symbols, etc.).
[0200] However, the molecular structural formulas detected in each of the first page image (810) and the second page image (820) examined above may further include at least some molecular structural formulas that are not assigned reference numerals. That is, at least some molecular structural formulas that are not assigned reference numerals in (a) and (b) of FIG. 10 may also be understood to be included in the detected molecular structural formulas.
[0201] Next, the second model (512) may be configured to analyze at least one page included in the document to be analyzed (800) to recognize at least one chemical reaction equation and output the recognition result of at least one chemical reaction equation recognized on at least one page. In this specification, the phrase “recognizing a chemical reaction equation in a page image” may also be understood as recognizing a chemical reaction equation on a page.
[0202] For example, as illustrated in FIG. 11 (a), the second model (512) can recognize a plurality of chemical reaction equations (831, 832, 833) in the first page image (810) and output the recognition result of each of the plurality of chemical reaction equations (831, 832, 833) recognized in the first page image (810). In this case, the recognition result of each of the plurality of chemical reaction equations (831, 832, 833) includes i) information on the components constituting the plurality of chemical reaction equations (831, 832, 833) recognized in the first page image (810) (e.g., reactants (molecular structures introduced into the reaction or regions of said molecular structures), products (molecular structures produced as a result of the reaction), reaction conditions (condition information in text form such as reagents, catalysts, solvents, temperature, time, etc.), ii) location information on each of the above components (e.g., coordinates (bounding boxes) of each of the reactants, reaction conditions, and products constituting each of the plurality of chemical reaction equations included in the first page image (810), iii) multiple regions containing each of the plurality of chemical reaction equations (831, 832, 833) in the first page image (810) (e.g., regions corresponding to each of the reactants, reaction conditions, and products constituting each of the plurality of chemical reaction equations), and iv) information on the reaction relationship between the above components (e.g., arrows, directions, steps, It may include at least one of the following: sequential relationships or reaction flow information based on arrows.
[0203] In one embodiment, the recognition result of the first chemical reaction equation (831) among the plurality of chemical reaction equations (831, 832, 833) recognized in the first page image (810) may include at least one of: i) information of the components constituting the first chemical reaction equation (831) (e.g., first reactant (811), first reaction conditions (811a, 811b), first product (812), etc.); ii) location information of each of the components (811, 811a, 811b, 812); iii) a plurality of regions each containing the components (811, 811a, 811b, 812) of the document to be analyzed (800, or the first page of the document to be analyzed (800)); and iv) information on the reaction relationship between the components (811, 811a, 811b, 812).
[0204] As another embodiment, the recognition result of the second chemical reaction equation (832) among the plurality of chemical reaction equations (831, 832, 833) recognized in the first page image (810) may include at least one of: i) information of the components constituting the second chemical reaction equation (832) (e.g., second reactant (813), second reaction conditions (813a, 813b), second product (814), etc.); ii) location information of each of the components (813, 813a, 813b, 814); iii) a plurality of regions each containing the components (813, 813a, 813b, 814) of the document to be analyzed (800, or the first page of the document to be analyzed (800)); and iv) information on the reaction relationship between the components (813, 813a, 813b, 814).
[0205] As another embodiment, the recognition result of the third chemical reaction equation (833) among the plurality of chemical reaction equations (831, 832, 833) recognized in the first page image (810) may include at least one of: i) information of the components constituting the third chemical reaction equation (833) (e.g., second reactant (813), third reaction condition (815a), third product (815), etc.), ii) location information of each of the said components (813, 815a, 815), iii) a plurality of regions each containing said components (813, 815a, 815) of the document to be analyzed (800, or the first page of the document to be analyzed (800)), and iv) information on the reaction relationship between said components (813, 815a, 815).
[0206] Additionally, as illustrated in FIG. 11 (b), the second model (512) can recognize each of the plurality of chemical reaction equations (841, 842) in the second page image (810) and output the recognition result of each of the plurality of chemical reaction equations (841, 842) recognized in the second page image (820). In this case, the recognition result of each of the plurality of chemical reaction equations (841, 842) includes at least one of i) information on the components constituting the plurality of chemical reaction equations (841, 842) recognized in the second page image (820) (e.g., reactants (molecular structures introduced into the reaction or regions of said molecular structures), products (molecular structures produced as a result of the reaction), reaction conditions (condition information in text form such as reagents, catalysts, solvents, temperature, time, etc.), ii) location information on each of the said components (e.g., coordinates (bounding boxes) of each of the reactants, reaction conditions, and products constituting each of the plurality of chemical reaction equations included in the second page image (820), iii) multiple regions containing each of the plurality of chemical reaction equations (841, 842) in the second page image (820) (e.g., regions corresponding to each of the reactants, reaction conditions, and products constituting each of the plurality of chemical reaction equations), and iv) information on the reaction relationships between the said components (e.g., arrows, directions, steps, sequence relationships, or reaction flow information based on arrows, etc.). It can be included.
[0207] In one embodiment, the recognition result of the first chemical reaction equation (841) among the plurality of chemical reaction equations (841, 842) recognized in the second page image (820) may include at least one of: i) information of the components constituting the first chemical reaction equation (841) (e.g., first reactant (824), first reaction conditions (824a, 824b), first product (825), etc.); ii) location information of each of the components (824, 824a, 824b, 825); iii) a plurality of regions each containing the components (824, 824a, 824b, 825) of the document to be analyzed (800, or the second page of the document to be analyzed (800)); and iv) information on the reaction relationship between the components (824, 824a, 824b, 825).
[0208] In another embodiment, the recognition result of the second chemical reaction equation (842) among the plurality of chemical reaction equations (841, 842) recognized in the second page image (820) may include at least one of: i) information of the components constituting the second chemical reaction equation (842) (e.g., second reactant (828), second reaction condition (828a), second product (829), etc.); ii) location information of each of the components (828, 828a, 829); iii) a plurality of regions each containing the components (828, 828a, 829) of the document to be analyzed (800, or the second page of the document to be analyzed (800)); and iv) information on the reaction relationship between the components (828, 828a, 829).
[0209] However, the chemical reaction equations recognized in each of the first page image (810) and the second page image (820) examined above may further include at least some chemical reaction equations that are not assigned reference numbers. That is, at least some chemical reaction equations that are not assigned reference numbers in (a) and (b) of FIG. 11 may also be understood to be included in the recognized chemical reaction equations.
[0210] Furthermore, the chemical information processing unit (510) can combine the detection result of the molecular structure formula output from the first model (511) and the recognition result of the chemical reaction formula output from the second model (512) to generate at least one output of chemical data included in the document to be analyzed (800).
[0211] In the present invention, the process of generating at least one output can be understood as a process of combining molecular structural formula detection results and chemical reaction formula recognition results to map each molecular structural formula to reactants, reaction conditions, and products, and generating document-unit structured chemical reaction information (or chemical reaction data structured by reaction units). Alternatively, it can be understood as a process of structuring and generating chemical reaction information composed of reactants, reaction conditions, and products.
[0212] As seen above, the process of detecting molecular structural formulas and recognizing chemical reaction formulas from the document to be analyzed (800) can be processed page by page included in the document to be analyzed (800).
[0213] The chemical information processing unit (510) can combine the detection result of a molecular structural formula detected in at least one page included in the document to be analyzed (800) and the recognition result of a chemical reaction formula to generate at least one output for the chemical data included in the document to be analyzed (800).
[0214] More specifically, the chemical information processing unit (510) can generate at least one output by combining the detection result of the molecular structure formula and the recognition result of the chemical reaction formula based on the structural association between the detection result of the molecular structure formula and the recognition result of the chemical reaction formula.
[0215] Here, structural association refers to a relationship in which independently detected and / or recognized objects or elements correspond to, map, and / or are connected to one another based on objective and rule-based structural characteristics—such as form, components, spatial arrangement, structural role, and connectivity—rather than semantic similarity or simple co-occurrence. In other words, it refers to a relationship in which objects correspond to, map, and are connected to one another based on their structural characteristics. This indicates the position and function of objects (or elements) within the same structural system or functional unit, and how they are consistently connected and combined. For example, in the analysis of chemical domain documents and / or chemical reactions, it may refer to connecting the atoms, bonds, and graph structures of a molecular formula at the structural level to indicate their position and function as reactants, products, or intermediates within a chemical equation. Alternatively, the atoms, bonds, and graph structures of a molecular formula may correspond at the structural level to reactants, products, and intermediates within a chemical equation, and may be associated through structural changes and roles before and after the reaction. Alternatively, it includes relationships in which atoms, bonds, and graph structures of a molecular formula correspond, / or match, and / or transform with reactants, products, and intermediates within a chemical equation at the structural level, and can serve as key connection criteria for integrated molecular-reaction analysis and the generation of accurate combined results. Such structural associations can integrate different detection, recognition, and analysis results into a single consistent structural semantic framework, and serve as key connection criteria for accurately and systematically combining and interpreting complex information.
[0216] In this regard, in the present invention, structural association may refer to a chemical, structural, spatial, and contextual correspondence existing between a molecular structural formula detected in a document under analysis and a chemical reaction equation recognition result. More specifically, structural association may refer to a connection defined through chemical, structural, spatial, and contextual criteria that the molecular structural formula and the chemical reaction equation recognition result detected within the same document share identical or transformative chemical structures and play corresponding roles within the reaction flow.
[0217] In one embodiment, the structural association may include cases where the molecular structures of the reactants, products, and intermediates recognized in the chemical equation are structurally identical or partially identical to the individual molecules obtained from the molecular structure formula detection results, or where structural differences such as bond formation, cleavage, substitution, or functional group changes before and after the reaction can be explained through comparison between molecular structures.
[0218] In another embodiment, structural association may include cases that can be explained by the preservation or predictable change of the core framework, functional groups, or bonding patterns.
[0219] In another embodiment, the structural association may include structural transformation relationships in which structural differences, such as bond formation, cleavage, substitution, and functional group changes between molecules before and after the reaction, can be explained through a comparison between molecular structures.
[0220] In another embodiment, structural associations may include spatial and contextual associations that define the role (reactant, product, intermediate, etc.) a specific molecular formula plays within the reaction flow through reaction arrows, reaction steps, relative positions with condition text, and visual connection elements. Additionally, structural representations extracted from the chemical equation recognition results (e.g., chemical structure representation format, graph, etc.) may include relationships in which structural data from the molecular formula detection results can be matched at the graph level.
[0221] In other words, structural association refers not merely to visual coexistence or proximity, but to an integrated standard (core semantic connection) that links molecular structure information with reaction procedure information at the document level to explain, within a consistent chemical semantic structure, the roles molecules (each molecule) play and how they transform within the reaction flow. This can refer to an integrated chemical semantic relationship that enables the explanation of how molecules transform within the reaction flow using a consistent reaction graph. Alternatively, it may be a core chemical semantic connection for explaining consistent chemical reaction flows and structural transformations at the document level and constructing an integrated reaction graph.
[0222] The chemical information processing unit (510) can analyze the structural association between the detection result of the molecular structure formula and the recognition result of the chemical reaction formula, and combine the detection result of the molecular structure formula and the recognition result of the chemical reaction formula based on the analyzed structural association to generate at least one output. Here, “analyzing the structural association and combining based on the analysis result” can be understood to mean analyzing at least one relationship between the detection result of the molecular structure formula and the recognition result of the chemical reaction formula and combining based on the analysis result. For example, at least one relationship may include at least one of a spatial relationship and a semantic relationship (or a logical relationship).
[0223] Spatial relationships may refer to information indicating where each component (or object), such as molecular structural formulas, reaction arrows, and text objects detected in a document, is positioned and what geometric relationships (placement relationships) they have with one another. Spatial relationships are determined based on the relative positions, distances, directions, and alignment states between objects, and may include criteria for determining whether the objects belong to the same chemical reaction unit or the same reaction stage. For example, at least one of the following may be a factor for determining spatial relationships: ii) left / right or up / down position (or placement) relative to the reaction arrow; ii) distance, adjacency, alignment state, alignment on the same axis, or overlap or inclusion relationships between bounding boxes (position information). This can be used to determine whether objects belong to the same chemical reaction and / or positional distinctions within the reaction. In other words, spatial relationships may be relationship information that describes where an object is located within a document.
[0224] Furthermore, semantic relationships can be defined not by the physical location of an object, but based on the semantic role the object plays within a chemical reaction (or chemical equation), such as reactant, product, or condition, as well as the relationship of membership with the reaction step, direction of reaction, and reaction unit. For example, whether a specific molecule is a starting material or a product of a reaction, whether a text object corresponds to reaction conditions such as reagent, catalyst, temperature, or solvent, or which of multiple reactions within the same page a specific molecular formula is included in can be determined by semantic relationships. In other words, semantic relationships can be relational information that explains what role an object performs within a chemical equation. In the present invention, based on these semantic relationships, functional roles (e.g., reactant, product, reaction condition, etc.) within a chemical reaction can be determined (or assigned) for at least some of the detected objects (e.g., molecular formula or text object, etc.). This can be understood as specifying (or determining, assigning, etc.) the role performed by at least some of the molecular structural formulas in the chemical reaction when, among the detected molecular structural formulas, there are at least some of the molecular structural formulas included in the chemical reaction.
[0225] Furthermore, at least one output generated by combining based on such structural associations may include information that combines (or aligns) the spatial relationships between each molecular structural formula and chemical equation detected at the page level of a document, and explicitly links the role that the said molecular structural formula plays as a reactant, product, or reaction condition element. Additionally, the output may include location and structural information of each molecular structural formula detected within the page, or structurally represent reaction components such as reactants, products, and reaction conditions included in the chemical equation. In this case, the role that each molecular structural formula plays in a specific chemical reaction can be identified by analyzing the spatial relationship between the molecular structural formula and the chemical equation during the combining process. That is, semantic linkage information, such as whether a molecule is a reactant or a product, can be specified.
[0226] In other words, the output generated through this combination does not list molecular structural formulas and chemical equations (or chemical reaction information) independently, but rather expresses their interrelationships in a structured form. This allows for an understanding of the components and meanings of chemical reactions within a document through a consistent data structure, and enables the consistent interpretation of chemical information at the document level. This can be utilized as a final output for document-based chemical information extraction, reaction database construction, and automated analysis.
[0227] For example, as illustrated in (a) of FIG. 12, the chemical information processing unit (510) can combine the detection result of the molecular structure formula detected in the first page image (810) and the recognition result of the chemical reaction formula recognized in the first page image (810) based on the structural association between the detection result of the molecular structure formula detected in the first page image (810) and the recognition result of the chemical reaction formula recognized in the first page image (810) to generate an output (910) for chemical data included in the first page image (810).
[0228] In this case, the output (910) may include molecular structural formulas (816, 817, 818, 819) that are not included in the chemical reaction formulas (831, 832, 833) detected in the first page image (810), among the molecular structural formulas (811, 812, 813, 814, 815, 816, 817, 818, 819) detected in the first page image (810).
[0229] Additionally, the output (910) may include molecular structural formulas (811, 812, 813, 814, 815) included in at least some chemical reaction formulas (831, 832, 833) recognized in the first page image (810), among the molecular structural formulas (811, 812, 813, 814, 815, 816, 817, 818, 819) detected in the first page image (810).
[0230] In one embodiment, the output (910) may include a first molecular structural formula (811), reaction conditions (811a, 811b), and a second molecular structural formula (812) included in the first chemical reaction formula (831). In this case, the first molecular structural formula (811) may correspond to a reactant, and the second molecular structural formula (812) may correspond to a product.
[0231] In another embodiment, the output (910) may include a third molecular structural formula (813), reaction conditions (813a, 813b), and a fourth molecular structural formula (814) included in the second chemical reaction formula (832). In this case, the third molecular structural formula (813) may correspond to a reactant, and the fourth molecular structural formula (814) may correspond to a product.
[0232] Furthermore, the output (910) can be understood to include a chemical reaction formula including a molecular structure formula.
[0233] In one embodiment, the output (910) may include a first chemical reaction formula (831) composed of a first molecular structural formula (811), reaction conditions (811a, 811b), and a second molecular structural formula (812) detected in a first page image (810).
[0234] In another embodiment, the output (910) may include a second chemical reaction formula (832) composed of a third molecular structural formula (813), reaction conditions (813a, 813b), and a fourth molecular structural formula (814) detected in the first page image (810).
[0235] As another example, as illustrated in FIG. 12 (b), the chemical information processing unit (510) can combine the detection result of the molecular structure formula detected in the second page image (820) and the recognition result of the chemical reaction formula recognized in the second page image (820) based on the structural association between the detection result of the molecular structure formula detected in the second page image (820) and the recognition result of the chemical reaction formula recognized in the second page image (820) to generate an output (920) for chemical data included in the second page image (820).
[0236] In this case, the output (920) may include molecular structural formulas (821, 822, 823, 824, 825, 826, 827, 828, 829) detected in the second page image (820) that are not included in the chemical reaction formulas (841, 842) recognized in the second page image (820).
[0237] Additionally, the output (920) may include molecular structural formulas (824, 825, 828, 829) included in at least some of the chemical reaction formulas (841, 842) recognized in the second page image (820), among the molecular structural formulas (821, 822, 823, 824, 825, 826, 827, 828, 829) detected in the second page image (820).
[0238] In one embodiment, the output (920) may include a first molecular structural formula (824), reaction conditions (824a, 824b), and a second molecular structural formula (825) included in the first chemical reaction formula (841). In this case, the first molecular structural formula (824) may correspond to a reactant, and the second molecular structural formula (825) may correspond to a product.
[0239] In another embodiment, the output (920) may include a third molecular structural formula (828), reaction conditions (828a), and a fourth molecular structural formula (829) included in the second chemical reaction formula (842). In this case, the third molecular structural formula (828) may correspond to a reactant, and the fourth molecular structural formula (829) may correspond to a product.
[0240] Furthermore, the above output (920) can be understood to include a chemical reaction formula including a molecular structure formula.
[0241] In one embodiment, the output (920) may include a first chemical reaction formula (841) composed of a first molecular structural formula (824), reaction conditions (824a, 824b), and a second molecular structural formula (825) detected in the second page image (820).
[0242] In another embodiment, the output (920) may include a second chemical reaction formula (842) composed of a third molecular structural formula (828), reaction conditions (828a), and a fourth molecular structural formula (829) detected in the second page image (820).
[0243] In this way, the present invention can perform data processing that combines the molecular structure formula detection result and the chemical reaction formula recognition result to reconstruct integrated chemical reaction information at the document level.
[0244] In other words, the final output examined above may include, among all molecular structural formulas detected in the document, those not included in the chemical reaction equation and those included in the chemical reaction equation that function as reactants or products. Furthermore, the molecular structural formulas included in the chemical reaction equation are linked to correspond to the same reaction and are provided along with relationship information regarding the chemical reaction equation to which each molecular structural formula belongs. Accordingly, the final output does not list the molecular structural formulas and chemical reaction equations included in the document individually, but rather provides an integrated chemical information representation that reflects the reaction participation status and reaction structure of each molecular structural formula.
[0245] Furthermore, the structural associations discussed above may also be referred to as “structural association information.” In this case, structural association information may be information used to correspond and / or link molecular structural formula detection results and chemical reaction equation recognition results, which are detected and / or recognized independently within a document, to the same chemical object. Based on structural features such as the shape, components, bonding relationships, and spatial arrangement of the molecular structural formula, this structural association information may indicate the relationship of roles as reactants, products, or other reaction components within the equation. Additionally, structural association may include correspondences based on visual and / or structural elements, such as arrows, positions, connection structures, and structural change correspondences included in chemical reaction diagrams. In other words, structural association information is defined based on structural shapes, arrangements, and connection relationships, without relying on the semantic interpretation or context of the document; thereby, it can be utilized as reference information to combine multiple recognition results into integrated final information with a consistent structure.
[0246] Furthermore, in the present invention, structural association may also be referred to as a “structural relationship.” In this case, a structural relationship may refer to a relationship defined by comprehensively reflecting the spatial arrangement relationship, semantic role relationship, and logical connection relationship according to reaction flow that exist between the output results of different models within a document. Such a structural relationship may include the location of molecular structures, functional roles in chemical reactions (reactants, products, etc.), connection patterns between elements within a reaction diagram, and the stepwise reaction sequence.
[0247] Meanwhile, the present invention can generate and provide to the user an analysis result of chemical data based on an output generated by combining the molecular structure formula detection result and the chemical reaction formula recognition result.
[0248] In one embodiment, the document understanding system (1000) may visualize and provide to the user terminal (10) the analysis result of a molecular structural formula detected in the document to be analyzed and / or the analysis result of a chemical reaction formula recognized in the document to be analyzed. As illustrated in FIG. 13, the document understanding system (1000) may provide the analysis result of at least one chemical reaction formula (1301) recognized in the first page (1300) of the document to be analyzed to a part of the service page (2000) output on the user terminal (10). In this case, it may be understood that the analysis result of the document to be analyzed and the analysis result of the molecular structural formula and / or chemical reaction formula detected and / or recognized in the document to be analyzed are provided together.
[0249] In another embodiment, the document understanding system (1000) may integrate the analysis result of a molecular structural formula detected in the document to be analyzed and the recognition result of a chemical reaction formula detected in the document to be analyzed and provide them to the user terminal (10). In this case, the document understanding system (1000) may provide the analysis result for the selected target to the user when at least one of the components constituting the molecular structural formula and / or chemical reaction formula is selected from the user terminal (10). As illustrated in FIG. 14 (a) and (b), the document understanding system (1000) may provide a document (1400) containing the detection result of the molecular structural formula and the recognition result of the chemical reaction formula to a part of the service page (2000). At this time, the document understanding system (1000) may provide the analysis result (1401a) for the selected chemical reaction formula (1401) to the user terminal (10) based on receiving a user input selecting one of the chemical reaction formulas (1401) in the document (1400).
[0250] Meanwhile, according to another embodiment of the present invention, the document understanding system (1000) may include a function that allows a user to directly upload a molecular image in order to perform molecular structure recognition.
[0251] For example, the molecular image may be a paper, patent document, experimental document, research report, image file (e.g., JPG, PNG, TIFF, etc.), or an image extracted from a document, and may include a single molecular structural formula or multiple molecular structural formulas.
[0252] In one embodiment, the user may select (or input) a molecular image to be analyzed through a user terminal or upload it using a drag-and-drop method. In this case, the document understanding system (1000) may perform preprocessing steps such as resolution normalization, noise removal, contrast correction, and binarization on the uploaded image. These preprocessing steps are intended to improve the accuracy of molecular structure recognition performed thereafter and can contribute to minimizing recognition deviations due to image quality or shooting environment.
[0253] Furthermore, the molecular image upload step described above can be performed on a user-selected area basis rather than the entire document, thereby allowing only specific molecular structural formulas within the document to be selectively designated as analysis targets. In this way, the present invention provides the effect of improving analysis efficiency by enabling the rapid processing of only the molecular structures required by the user.
[0254] Meanwhile, according to another embodiment of the present invention, the document understanding system (1000) can perform artificial intelligence-based molecular structure recognition using an uploaded molecular image as input.
[0255] In this case, the artificial intelligence model performing molecular structure recognition is a model specialized in detecting molecular structural formulas and can automatically identify atomic symbols, bond lines, ring structures, directional bonds, substituents, etc. included in the image.
[0256] In one embodiment, the document understanding system (1000) can identify individual atoms from an uploaded molecular image and analyze bonding relationships between atoms to convert the molecular structure into structural data in the form of a graph. At this time, each atom can be represented as a node and a bond as an edge, and the type of bond (single bond, double bond, directional bond, etc.) can also be defined.
[0257] Furthermore, the above molecular structure recognition process can convert molecular structural formulas into chemical structure representation formats such as SMILES, InChI, and Mol files. Through this, the recognized molecular structures can be utilized for database storage, retrieval, comparison, and further analysis.
[0258] Meanwhile, according to another embodiment of the present invention, the document understanding system (1000) can identify areas where there is a possibility of error among the results of artificial intelligence-based molecular structure recognition and provide them to the user by visually highlighting them.
[0259] In one embodiment, the document understanding system (1000) can calculate a confidence score for each recognized atom, bond, or structural unit, and the confidence score can be calculated based on the model's output probability, chemical rule fit, consistency with adjacent structures, etc. At this time, if the confidence score is below a certain threshold, the area may be determined to be an area with a high probability of error.
[0260] Areas where such potential errors exist can be represented on the user interface through methods such as color highlighting, dotted lines, icon attachment, and warning message display. For example, if the recognition of an atomic symbol is uncertain, a highlight is provided around the atom, and if the bond order is ambiguous, the corresponding bond can be displayed separately. This allows users to intuitively identify parts of the system's recognized results that are unreliable and enables them to focus on reviewing and correcting those areas in subsequent steps.
[0261] Meanwhile, according to another embodiment of the present invention, the document understanding system (1000) can provide the user with a function to directly modify the molecular structure recognition result.
[0262] In one embodiment, the user can modify the corresponding atom or bond by clicking or selecting a highlighted error area, and can perform editing operations such as changing the atom type, adding or deleting bonds, changing the bond order, or redefining the ring structure. Additionally, the user can intuitively readjust the molecular structure through interactions such as dragging, rotating, and zooming in / out.
[0263] Such user modifications enable the generation of a more accurate final structure by reflecting the user's expertise and judgment, rather than simply accepting the AI recognition results. In particular, when complex substituent structures or stereochemical information are involved, user intervention can significantly improve structural accuracy.
[0264] Meanwhile, according to another embodiment of the present invention, during the process in which a molecular structure is modified by a user, the document understanding system (1000) can perform chemical rule verification in real time.
[0265] In one embodiment, the document understanding system (1000) can automatically verify chemical rules such as valence, bond order, charge balance, chirality, and ring closure whenever a user modifies an atom or bond.
[0266] In this case, if the modified structure is chemically unreasonable, the system can immediately output a warning message or visually indicate the problem area. For example, if the bonding order of a carbon atom exceeds the allowable range, the atom may be marked as an error, and if the chiral center is mismatched, the user may be notified of the need for correction. This real-time verification function ensures the chemical validity of the final result by enabling the user to immediately recognize and correct structural errors.
[0267] Meanwhile, according to another embodiment of the present invention, the document understanding system (1000) may include a function to automatically save all molecular structure modification history performed by the user.
[0268] In one embodiment, each modification step may be recorded along with a timestamp, modification details, and structural information before and after the modification, and the user may rollback to a previous version or compare different modification versions. Additionally, the modification history may be shared with other users or researchers for research collaboration and may be provided in the form of a link, file, or database integration.
[0269] Through this, the present invention can be utilized in collaborative chemical research environments beyond a single-user environment, and can ensure transparency and reproducibility in the molecular structure recognition and modification process.
[0270] As described above, the document understanding method and system according to the present invention utilizes a model specialized for molecular structure formula processing and a model specialized for chemical reaction formula processing to perform inference on a document in parallel, and combines the molecular structure formula detection results and chemical reaction formula recognition results in a post-processing step to generate an integrated result. Through this, the present invention connects molecular structure formula detection and chemical reaction formula recognition, which were previously processed individually, into a single workflow, thereby maximizing the efficiency of document-based chemical information processing and significantly expanding the potential for application in research and industrial settings.
[0271] Furthermore, according to the document understanding method and system of the present invention, molecular structural formulas and chemical reaction formulas contained in a document can be simultaneously detected and analyzed by utilizing a model specialized for processing molecular structural formulas and a model specialized for processing chemical reaction formulas. The results of such detection and analysis can be provided to a user, allowing the user to intuitively recognize the necessary information and understand it more quickly, thereby increasing the accuracy and efficiency of the research. In other words, the user can receive the necessary information from the document quickly and accurately, thus reducing the time and cost required for research or development.
[0272] Furthermore, according to the document understanding method and system of the present invention, by utilizing a model specialized for processing molecular structural formulas to detect molecular structural formulas in documents and converting them into structural data, they can be utilized for building chemical databases, searching for and analyzing chemical information, etc.
[0273] Furthermore, according to the document understanding method and system of the present invention, a chemical reaction equation composed of reactants, conditions, and products can be extracted from a page-level document by utilizing a model specialized for chemical reaction equation processing. In other words, the present invention enables the understanding of structural data of a chemical reaction, the visualization of the reaction pathway, and the automatic processing of reaction data.
[0274] Meanwhile, the present invention described above can be implemented based on a quantum computer. The present invention implemented based on a quantum computer may include a qubit-based quantum processor and quantum memory, and may include software and hardware interfaces optimized for quantum computation.
[0275] Quantum processors in quantum computers utilize qubits to efficiently process complex operations through parallel computation, quantum entanglement, and quantum superposition, which cannot be performed by the binary bits of classical computers. Quantum processors process data using quantum gates and can provide exponential speed improvements for specific problems.
[0276] Meanwhile, the present invention described above can be implemented as a program that is executed by one or more processes on a computer and can be stored on a computer-readable medium (or recording medium).
[0277] Furthermore, the present invention described above can be implemented as computer-readable code or instructions on a medium on which a program is recorded. That is, the present invention can be provided in the form of a program.
[0278] Meanwhile, computer-readable media include all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0279] Furthermore, the computer-readable medium may be a server or cloud storage that includes a storage and is accessible to an electronic device via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage via wired or wireless communication.
[0280] A computer program may reach the system (1000) through various suitable transmission mechanisms. The transmission mechanism may be, for example, a computer-readable storage medium, a computer program product, a memory device, a recording medium such as a CD-ROM or DVD, or a product that tangibly embodies the computer program. The transmission mechanism may be a signal configured to reliably transmit the computer program through air or an electrical connection. The system (1000) may propagate or transmit the computer program as a computer data signal.
[0281] Furthermore, references to 'computer-readable storage media,' 'computer program products,' 'computer programs embodied in a tangible form,' etc., or to 'controller,' 'computer,' 'processor,' etc., should be understood to include not only computers with various architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures, but also specialized circuits such as Field-Programmable Gate Arrays (FPGAs), Application Specific Circuits (ASICs), signal processing units, and other devices. References to computer programs, instructions, code, etc., should be understood to include software for programmable processors or firmware, such as programmable content for hardware devices, whether it is instructions for a processor or configuration settings for a fixed-function device, gate array, or programmable logic device.
[0282] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, namely a CPU (Central Processing Unit), and no special limitations are placed on its type.
[0283] Meanwhile, the above detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.
Claims
1. Regarding methods performed by a computer, A step of specifying at least one document to be analyzed; A step of processing the above-mentioned document to be analyzed as input to at least one information processing model configured to process information related to a chemical domain; Using the above-mentioned at least one information processing model, a step of generating a first processing result for molecular structure information included in the document to be analyzed and generating a second processing result for chemical reaction information included in the document to be analyzed; and A document understanding method characterized by including the step of generating output data for chemical information reflecting the mutual correlation of the molecular structure information and the chemical reaction information using the first processing result and the second processing result.
2. In Paragraph 1, The above at least one information processing model is, A first model for generating the first processing result for the molecular structure information included in the above-mentioned document to be analyzed, and It includes a second model for generating the second processing result for the chemical reaction information included in the above-mentioned document to be analyzed, and A document understanding method characterized by the fact that the above output data is generated by combining the first processing result output from the first model and the second processing result output from the second model.
3. In Paragraph 2, The above first model is, A document understanding method characterized by detecting at least one molecular structural formula from at least one page included in the above-mentioned document to be analyzed, and outputting the first processing result for the above-mentioned molecular structural information.
4. In Paragraph 3, The detection of at least one molecular structural formula above is, Detecting at least one molecular structural formula in the above-mentioned analysis target document, and By analyzing the bonding relationships between atoms constituting at least one molecular structural formula, A document understanding method characterized by including a task of converting at least one molecular structural formula into structural data based on the above analysis results.
5. In Paragraph 3, The above first processing result is, The above at least one molecular structural formula, The position information of at least one molecular structural formula and A document understanding method characterized by including at least one of at least one region information containing at least one molecular structural formula in the above-mentioned document to be analyzed.
6. In Paragraph 2, The above second model is, A document understanding method characterized by recognizing at least one chemical reaction formula from at least one page included in the above-mentioned document to be analyzed, and outputting the second processing result for the above-mentioned chemical reaction information.
7. In Paragraph 6, The above at least one chemical reaction equation recognition is, Extract at least one of the reactants, reaction conditions, and products constituting the at least one chemical reaction equation from the above-mentioned document subject to analysis, or A method for understanding a document characterized by including a task of extracting at least one of reactants, reaction conditions, and products from a chemical reaction diagram included in the document to be analyzed.
8. In Paragraph 6, The above second processing result is, A document understanding method characterized by including at least one of information on components constituting at least one chemical reaction equation, location information of said components, information on at least one region containing said at least one chemical reaction equation in said document to be analyzed, and information on reaction relationships between said components.
9. In Paragraph 2, The above-mentioned document subject to analysis is configured to include at least one page, and The above output data is, A document understanding method characterized by combining a detection result of at least one molecular structural formula corresponding to the first processing result detected in at least one page and a recognition result of at least one chemical reaction formula corresponding to the second processing result detected in at least one page to generate output data for the chemical information included in the document to be analyzed.
10. In Paragraph 2, A step of processing the molecular structure information in the first model and outputting the first processing result, and A document understanding method characterized in that the step of processing the chemical reaction information in the second model and outputting the second processing result is a step performed in parallel.
11. In Paragraph 10, The above first model is, Performing inference to process the above molecular structure information and outputting a first inference result, The above second model is, A document understanding method characterized by performing inference to process the above chemical reaction information and outputting a second inference result.
12. In Paragraph 11, The first inference result above includes the first processing result for the molecular structure information, and A document understanding method characterized in that the above second inference result includes the above second processing result for the above chemical reaction information.
13. In Paragraph 12, A document understanding method characterized by the output data being generated by combining the first inference result and the second inference result.
14. In Paragraph 9, The above output data is, A document understanding method characterized by being generated based on the structural association between the first processing result of the molecular structure information and the second processing result of the chemical reaction information.
15. In Paragraph 14, The above output data is, First molecular structure information not included in the above chemical reaction formula, Second molecular structure information included in the above chemical reaction formula and A document understanding method characterized by including the chemical reaction information including the second molecular structure information.
16. In Paragraph 1, It further includes a step of converting the above-mentioned document to be analyzed into a specific pre-set format, and When the above-mentioned document subject to analysis is converted into the above-mentioned specific format, Processing the analysis target document converted into the above specific format as input to the above at least one model, and The analysis target document converted into the above specific format is processed as input to the first model and the second model included in the above at least one model, and Output the first processing result for the molecular structure information from the first model, and Outputs the second processing result for the chemical reaction information from the second model above, and A document understanding method characterized by the output data being generated by combining the first processing result and the second processing result.
17. A system comprising memory configured to store executable instructions and one or more processors configured to perform operations by executing one or more instructions, The above system is, Identify at least one document to be analyzed, and The above-mentioned document to be analyzed is processed as input to at least one information processing model configured to process information related to the chemical domain, and Using the above at least one information processing model, a first processing result for molecular structure information included in the document to be analyzed is generated, and a second processing result for chemical reaction information included in the document to be analyzed is generated. A document understanding system characterized by generating output data for chemical information that reflects the mutual correlation between the molecular structure information and the chemical reaction information, using the first processing result and the second processing result.
18. A program that is executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The above program is, A step of specifying at least one document to be analyzed; A step of processing the above-mentioned document to be analyzed as input to at least one information processing model configured to process information related to a chemical domain; Using the above-mentioned at least one information processing model, a step of generating a first processing result for molecular structure information included in the document to be analyzed and generating a second processing result for chemical reaction information included in the document to be analyzed; and A program stored on a computer-readable recording medium characterized by including instructions for performing a step of generating output data for chemical information reflecting the mutual correlation of the molecular structure information and the chemical reaction information using the first processing result and the second processing result.