Method and system for parsing document
Patent Information
- Application Number
- PCT/KR2026/004608
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-23
- Publication Date
- 2026-10-01
Smart Images

Figure KR2026004608_01102026_PF_FP_ABST
Abstract
Description
Document Parsing Methods and Systems
[0001] The present invention relates to a method for parsing documents related to manufacturing and a system for the same.
[0002] Training generative artificial neural networks specialized for the manufacturing sector requires securing a certain amount of manufacturing-related documents. Since these documents encompass various types, a document parsing process is necessary. Furthermore, because manufacturing-related documents contain data in diverse forms such as text, formulas, tables, charts, diagrams, and images, extracting this data through general text mining processes is challenging.
[0003] Since conventional parsing tools operate based on fixed rules, problems such as errors occurring or some information being omitted frequently arise during the parsing process when the document structure is complex or contains unexpected data formats.
[0004] In particular, the Retrieval Augmented Generation (RAG) technique, recently applied to enhance the inference performance of generative artificial neural networks, requires meta-information to be utilized for inference when searching for graphs, diagrams, tables, and images. However, since meta-information is often absent in documents, there are significant instances where graphs and diagrams cannot be used for inference. Furthermore, when the generative artificial neural network generates an answer after inference, if formulas or tables are not formed in a specific format, the generated answer may differ from the intended result.
[0005] The information described above disclosed in the background technology of this invention is intended only to enhance understanding of the background of the present invention and may therefore include information that does not constitute prior art.
[0006] The present disclosure provides an effective parsing method, system, and computer program stored on a computer-readable recording medium for components included in a document related to manufacturing in order to solve the above-mentioned problems.
[0007] The present disclosure may be implemented in various ways, including a method, an apparatus (system), or a computer program stored on a readable storage medium.
[0008] A document parsing method performed by at least one processor according to one embodiment of the present disclosure for solving a technical problem may include the steps of receiving an original document, inputting the received original document into a layout analysis model to detect one or more components included in the original document, executing a parsing logic based on the type of the detected components, and processing the original document into output data of a preset data format based on the result of executing the parsing logic and the placement information of the detected components within the original document.
[0009] According to one embodiment of the present disclosure, the detected component may include at least one of text, formula, table, graph, chart, or diagram.
[0010] According to one embodiment of the present disclosure, a layout analysis model may be a model pre-trained to detect one or more components included in a received original document based on an element pool in which labeling is performed for each component of a document data set, and to determine the type of the one or more detected components.
[0011] According to one embodiment of the present disclosure, a layout analysis model can output placement information of components within an original document based on layout information of a document data set.
[0012] According to one embodiment of the present disclosure, the step of executing parsing logic includes the step of recognizing text included in a detected component based on Optical Character Recognition (OCR), and the text may include at least one of a character, syllable, word, or sentence included in one or more languages.
[0013] According to one embodiment of the present disclosure, the step of executing parsing logic may include the step of recognizing a formula part in a detected formula using a formula recognition model, and the step of converting the detected formula into a specific formula format based on the recognized formula part and the recognized text.
[0014] According to one embodiment of the present disclosure, the step of executing parsing logic may include recognizing the structure of a detected table using a table recognition model, and converting the detected table into a specific table format based on at least one of the structure of the recognized table, the recognized text, or the recognized formula part.
[0015] According to one embodiment of the present disclosure, the step of executing parsing logic includes the step of inputting at least one of a detected graph, chart, or diagram into a Vision-Language Model to generate meta-information, wherein the Vision-Language Model may be a model pre-trained to generate meta-information including at least one of descriptive information, type information, or analysis information associated with at least one of the detected graph, chart, or diagram.
[0016] According to one embodiment of the present disclosure, the document parsing method may further include the step of generating a converted document corresponding to an original document based on processed output data.
[0017] According to one embodiment of the present disclosure, the document parsing method may further include the step of outputting the generated converted document to a user interface.
[0018] According to one embodiment of the present disclosure, the document parsing method may further include the step of receiving user input to modify output data through a user interface.
[0019] A computer program stored on a computer-readable recording medium may be provided to execute a method according to one embodiment of the present disclosure on a computer.
[0020] An information processing system according to one embodiment of the present disclosure comprises a memory and a processor connected to the memory and configured to execute at least one computer-readable program included in the memory, and the at least one program may include instructions for receiving an original document, inputting the received original document into a layout analysis model to detect one or more components included in the original document, executing parsing logic based on the type of the detected components, and processing the original document into output data of a preset data format based on the result of executing the parsing logic and the placement information of the detected components within the original document.
[0021] According to one embodiment of the present disclosure, the detected component may include at least one of text, formula, table, graph, chart, or diagram.
[0022] According to one embodiment of the present disclosure, a layout analysis model may be a model pre-trained to detect one or more components included in a received original document based on an element pool in which labeling is performed for each component of a document data set, and to determine the type of the one or more detected components.
[0023] According to one embodiment of the present disclosure, a layout analysis model can output placement information of components within an original document based on layout information of a document data set.
[0024] According to one embodiment of the present disclosure, at least one program includes instructions for recognizing text included in a detected component based on Optical Character Recognition (OCR), and the text may include at least one of a character, syllable, word, or sentence included in one or more languages.
[0025] According to one embodiment of the present disclosure, at least one program may include instructions for recognizing a formula part in a detected formula using a formula recognition model, and converting the detected formula into a specific formula format based on the recognized formula part and the recognized text.
[0026] According to one embodiment of the present disclosure, at least one program may include instructions for recognizing the structure of a detected table using a table recognition model, and converting the detected table into a specific table format based on at least one of the structure of the recognized table, the recognized text, or the recognized formula part.
[0027] According to one embodiment of the present disclosure, at least one program includes instructions for generating meta-information by inputting at least one of a detected graph, chart, or diagram into a vision-language model, and the vision-language model may be a model pre-trained to generate meta-information including at least one of descriptive information, type information, or analysis information associated with at least one of the detected graph, chart, or diagram.
[0028] According to various embodiments of the present disclosure, various components included in documents related to manufacturing can be effectively structured.
[0029] According to various embodiments of the present disclosure, meta-information regarding unstructured components included in documents related to manufacturing can be generated and thus effectively utilized for inference by generative artificial neural networks.
[0030] According to various embodiments of the present disclosure, formulas or tables included in documents related to manufacturing can be structured in a specific format, so they can be effectively utilized when generating answers for a generative artificial neural network.
[0031] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art to which the present disclosure pertains (referred to as "person skilled in the art") from the description in the claims.
[0032] Embodiments of the present disclosure will be described with reference to the accompanying drawings described below, wherein similar reference numerals indicate similar elements, but are not limited thereto.
[0033] FIG. 1 is a diagram schematically illustrating a document parsing system according to one embodiment of the present disclosure.
[0034] FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to communicate with a plurality of user terminals to provide a parsing service for a document associated with manufacturing according to one embodiment of the present disclosure.
[0035] FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure.
[0036] FIG. 4 is a drawing for explaining a layout analysis model according to one embodiment of the present disclosure.
[0037] FIG. 5 is a diagram illustrating the learning process of a layout analysis model according to one embodiment of the present disclosure.
[0038] FIG. 6 is a diagram illustrating a method for parsing text detected in an original document according to one embodiment of the present disclosure.
[0039] FIG. 7 is a diagram illustrating a method for parsing a formula detected in an original document according to one embodiment of the present disclosure.
[0040] FIG. 8 is a diagram illustrating a method for parsing a graph detected in an original document according to one embodiment of the present disclosure.
[0041] FIG. 9 is a diagram illustrating a method for parsing a table detected in an original document according to one embodiment of the present disclosure.
[0042] FIG. 10 is a drawing illustrating a user interface for modifying output data of a preset format for a detected component included in an original document according to one embodiment of the present disclosure.
[0043] FIG. 11 is a flowchart illustrating a document parsing method according to one embodiment of the present disclosure.
[0044] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions regarding widely known functions or configurations will be omitted if there is a risk that the gist of the present disclosure may be unnecessarily obscured.
[0045] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Additionally, in the description of the following embodiments, the description of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.
[0046] The advantages and features of the disclosed embodiments and the methods for achieving them will become clear by referring to the embodiments described below in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms, and the embodiments provided are merely to make the present disclosure complete and to fully inform those skilled in the art of the scope of the invention.
[0047] The terms used in this specification will be briefly explained, and the disclosed embodiments will be described in detail. The terms used in this specification have been selected to be as generally used as possible, taking into account their functions in this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.
[0048] In this specification, singular expressions include plural expressions unless the context clearly specifies them as singular. Additionally, plural expressions include singular expressions unless the context clearly specifies them as plural. Throughout the specification, when a part is described as including a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0049] Additionally, the terms 'module' or 'part' as used in the specification refer to software or hardware components, and the 'module' or 'part' performs certain roles. However, the meaning of 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside in an addressable storage medium or configured to run on one or more processors. Thus, as an example, the 'module' or 'part' may include components such as software components, object-oriented software components, class components, and task components, and at least one of processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The components and the functions provided within the 'module' or 'part' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.
[0050] According to one embodiment of the present disclosure, a ‘module’ or ‘part’ may be implemented as a processor and memory. The term ‘processor’ should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, the term ‘processor’ may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. The term ‘processor’ may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a DSP core, or any other combination of such configurations. Additionally, the term ‘memory’ should be broadly interpreted to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as Random Access Memory (RAM), Read-Only Memory (ROM), Non-Volatile Random Access Memory (NVRAM), Programmable Read-Only Memory (PROM), Erasable-Programmable Read-Only Memory (EPROM), Electrically Erasable PROM (EEPROM), Flash Memory, Magnetic or Optical Data Storage Devices, Registers, etc. If a processor can read information from memory and / or write information to memory, the memory is said to be in an electronic communication state with the processor. Memory integrated into a processor is in an electronic communication state with the processor.
[0051] In the present disclosure, the 'system' may include at least one of a server device and a cloud device, but is not limited thereto. For example, the system may be composed of one or more server devices. As another example, the system may be composed of one or more cloud devices. As yet another example, the system may be configured and operated with both a server device and a cloud device.
[0052] In the present disclosure, 'display' may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided by the computing device.
[0053] In the present disclosure, 'each of a plurality of A' or 'each of a plurality of A' may refer to each of all components included in a plurality of A, or each of some components included in a plurality of A.
[0054] Various embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The size or location of display screens, images, buttons, etc., as illustrated and described in the drawings are exemplary and are not limited thereto. For example, some buttons may be added or omitted, or configured with sizes and locations different from those illustrated. Furthermore, the flowcharts and descriptions illustrated in the drawings are merely examples and may be implemented differently in some embodiments. For example, one or more steps may be omitted, the order of each step may be changed, one or more steps may be performed in overlap, or one or more steps may be performed repeatedly.
[0055] FIG. 1 is a diagram for schematically illustrating a document parsing system (100) according to one embodiment of the present disclosure.
[0056] Referring to FIG. 1, a document parsing system (100) may receive an original document (10) related to manufacturing. The original document (10) related to manufacturing may be a document related to manufacturing fields such as forging, injection molding, PLC (Programmable Logic Controller), press, welding, etc., but is not limited thereto. Additionally, the original document (10) related to manufacturing may include various documents related to at least one of product design, research, development, production, quality control, certification, maintenance, operation, laws, or regulations. Additionally, the original document (10) related to manufacturing may be a PDF (Portable Document Format) based document, but is not limited thereto.
[0057] The document parsing system (100) can detect at least one component among text, formulas, tables, graphs, charts, diagrams, or images included in the original document (10) associated with the received manufacturing.
[0058] In one embodiment, the document parsing system (100) inputs the received original document (10) into a layout analysis model to be described later to detect one or more components included in the original document (10) and determine the type of the detected components. Additionally, the layout analysis model can determine the placement information of the components within the original document (10).
[0059] The document parsing system (100) can execute parsing logic based on the type of detected component. The document parsing system (100) can analyze data and determine the meaning for each type of detected component to perform a customized structuring process.
[0060] A document parsing system (100) can process and output the original document (10) into output data (20) in a preset data format (e.g., JSON) based on the execution result of the parsing logic and the placement information of the detected components within the original document (10). Additionally, the document parsing system (100) can generate a converted document corresponding to the original document (10) based on the output data (20).
[0061] In the present disclosure, the parsing process of a document parsing system (100) for documents related to manufacturing is mainly described, but the documents are not limited to the manufacturing field.
[0062] FIG. 2 is a schematic diagram showing a configuration in which an information processing system (230) is connected to communicate with a plurality of user terminals (210_1, 210_2, 210_3) to provide a parsing service for a document associated with manufacturing according to one embodiment of the present disclosure.
[0063] Referring to FIG. 2, the information processing system (230) may include system(s) that perform a parsing process for documents related to manufacturing. In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer-executable programs (e.g., downloadable applications) and data related to document parsing services, or one or more distributed computing devices and / or distributed databases based on cloud computing services. For example, the information processing system (230) may include separate systems (e.g., servers) for document parsing services.
[0064] The document parsing service provided by the information processing system (230) can be provided to the user through an application installed on each of the multiple user terminals (210_1, 210_2, 210_3).
[0065] Multiple user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) through a network (220). The network (220) can be configured to enable communication between the multiple user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) may be configured as a wired network such as Ethernet, Power Line Communication, telephone line communication devices and RS-serial communication, a mobile communication network, a Wireless LAN (WLAN), Wi-Fi, Bluetooth and ZigBee, or a combination thereof. The communication method is not limited and may include not only communication methods utilizing communication networks that the network (220) may include (e.g., mobile communication network, wired internet, wireless internet, broadcasting network, satellite network, etc.) but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).
[0066] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are illustrated as examples of user terminals, but are not limited thereto, and the user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication and capable of installing and running applications, etc. For example, user terminals may include smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, etc. Additionally, FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with an information processing system (230) through a network (220), but is not limited thereto, and may be configured so that a different number of user terminals communicate with an information processing system (230) through a network (220).
[0067] In one embodiment, each of the user terminals (210_1, 210_2, 210_3) can receive information or data from another user terminal or transmit it to another user terminal through the network (220).
[0068] In FIG. 2, the information processing system (230) is shown as an independent device separated from the user terminals (210_1, 210_2, 210_3), but is not limited thereto, and the information processing system (230) may be implemented in an integrated manner with the user terminals (210_1, 210_2, 210_3).
[0069] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure.
[0070] The user terminal (210) may refer to any computing device capable of running applications, web browsers, etc., and capable of wired / wireless communication, and may include, for example, the mobile phone terminal (210_1), tablet terminal (210_2), PC terminal (210_3) of FIG. 2. Referring to FIG. 3, the user terminal (210) may include memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data through a network (220) using their respective communication modules (316, 336). Additionally, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) through the input / output interface (318).
[0071] The memory (312, 332) may include any non-transient computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as ROM (read-only memory), a disk drive, a solid-state drive (SSD), or flash memory. As another example, a permanent mass storage device such as ROM, an SSD, flash memory, or a disk drive may be included in the user terminal (210) or information processing system (230) as a separate permanent storage device distinct from the memory. Additionally, an operating system and at least one program code may be stored in the memory (312, 332).
[0072] These software components may be loaded from a computer-readable recording medium separate from memory (312, 332). This separate computer-readable recording medium may include a recording medium that can be directly connected to the user terminal (210) and the information processing system (230), for example, a computer-readable recording medium such as a floppy drive, disk, tape, DVD / CD-ROM drive, or memory card. As another example, the software components may be loaded into memory (312, 332) via a communication module (316, 336) rather than a computer-readable recording medium. For example, at least one program may be loaded into memory (312, 332) based on a computer program installed by files provided through a network (220) by developers or a file distribution system that distributes installation files for the application.
[0073] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a recording device such as memory (312, 332).
[0074] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system). For example, a request or data generated by the processor (314) of the user terminal (210) according to program code stored in a recording device such as memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, a control signal or command provided under the control of the processor (334) of the information processing system (230) may be received by the user terminal (210) via the communication module (316) of the user terminal (210) through the communication module (336) and the network (220).
[0075] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, or a mouse, and the output device may include a device such as a display, a speaker, or a haptic feedback device. As another example, the input / output interface (318) may be a means for interfacing with a device in which the configuration or function for performing input and output is integrated into one, such as a touchscreen. For example, when the processor (314) of the user terminal (210) processes instructions of a computer program loaded in memory (312), a service screen configured using information and / or data provided by an information processing system (230) or another user terminal may be displayed on a display through the input / output interface (318). In FIG. 3, the input / output device (320) is depicted as not being included in the user terminal (210), but is not limited thereto and may be configured as a single device with the user terminal (210). Additionally, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (230) or that the information processing system (230) may include. In FIG. 3, the input / output interface (318, 338) is shown as an element configured separately from the processor (314, 334), but is not limited thereto, and the input / output interface (318, 338) may be configured to be included in the processor (314, 334).
[0076] The user terminal (210) and the information processing system (230) may include more components than those of FIG. 3. However, it is not necessary to clearly illustrate most of the prior art components. In one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. Additionally, the user terminal (210) may further include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, a database, etc. For example, if the user terminal (210) is a smartphone, it may include components that are generally included in a smartphone, and may be implemented to include various components such as an accelerometer, a gyroscope, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.
[0077] While a program for an application including a document parsing service is running, the processor (314) can receive text, images, video, voice and / or actions, etc. that are input or selected through an input device such as a touch screen, keyboard, audio sensor and / or image sensor, camera, microphone, etc. connected to an input / output interface (318), and can store the received text, images, video, voice and / or actions, etc. in memory (312) or provide them to an information processing system (230) through a communication module (316) and a network (220).
[0078] The processor (314) of the user terminal (210) may be configured to manage, process, and / or store information and / or data received from an input / output device (320), another user terminal, an information processing system (230), and / or a plurality of external systems. The information and / or data processed by the processor (314) may be provided to the information processing system (230) through a communication module (316) and a network (220). The processor (314) of the user terminal (210) may transmit information and / or data to the input / output device (320) through an input / output interface (318) and output it. For example, the processor (314) may display the received information and / or data on the screen of the user terminal (210).
[0079] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from a plurality of user terminals (210) and / or a plurality of external systems. The information and / or data processed by the processor (334) may be provided to the user terminals (210) through a communication module (336) and a network (220). Hereinafter, the processor may refer to the processor of the information processing system or the user terminal.
[0080] FIG. 4 is a drawing for explaining a layout analysis model (420) according to one embodiment of the present disclosure.
[0081] Referring to FIG. 4, the layout analysis model (420) may be a pre-trained model that, upon receiving an original document (410), detects one or more components (430) included in the original document and determines the type of the detected components. Here, the components (430) may include at least one of text (431), formulas (432), tables (433), graphs (434), diagrams (435), or charts (436), but are not limited thereto.
[0082] The layout analysis model (420) can be pre-trained to recognize the placement information (440) of each component within the original document (410) using a pre-obtained document data set. Additionally, the layout analysis model (420) can output the placement information (440) of the components within the original document (410) based on the layout information of the document data set.
[0083] FIG. 5 is a diagram illustrating the learning process (500) of a layout analysis model according to one embodiment of the present disclosure.
[0084] Referring to FIG. 5, the processor may acquire a document dataset (510). The document dataset (510) may include papers, manuals, various documents from companies / institutions, etc. related to the manufacturing field, but the present disclosure is not limited thereto. The document dataset (510) may include various components that can be placed within the documents. The document dataset (510) may include English documents and / or Korean documents, but the present disclosure is not limited thereto.
[0085] The processor may perform element labeling (520) for a document data set (510). In one embodiment, the processor may perform bounding box-based labeling for each document data included in the document data set (510) to detect at least one of a header, graph, chart, diagram, image, formula, footnote, caption, table, paragraph, or code. This labeling may be performed by the processor or manually by a user, and is not limited to the method.
[0086] The processor may create an element pool (530) based on the components that have been labeled. In one embodiment, the processor may group and store each labeled component in the element pool (530), but the present disclosure is not limited thereto.
[0087] The processor can augment components of the component pool (530) through the process of synthesis-based component augmentation (540). In one embodiment, the processor can augment data of the component pool (530) by processing components within the component pool (530).
[0088] The processor may perform a data augmentation process (550) to augment data in the component pool (530) and the document dataset (510). The processor may extract components (552) from the component pool (530) (Step 1), divide sections of empty documents (554) (Step 2), match the extracted components with the divided sections (556) (Step 3), and fill the matched layout with the corresponding components (558, Step 4). The processor may repeat Steps 1 through 4. Based on the data augmentation process (550), the processor may generate a synthetic document dataset (560). The processor may add the synthetic document dataset to the document dataset (510) and add the synthesized components to the component pool (530). Accordingly, a layout analysis model may be generated, and manufacturing-related documents that are difficult to obtain due to security maintenance by companies or institutions may be effectively constructed.
[0089] The generated layout analysis model can detect one or more components included in the received original document and determine the type of the detected one or more components. In addition, the layout analysis model can recognize and output placement information of components within the original document based on the layout information of the document data set.
[0090] The processor can execute parsing logic based on the type of detected component, which will be explained together with reference to FIGS. 6 to 9 below.
[0091] FIG. 6 is a diagram illustrating a method for parsing text detected in an original document (610) according to one embodiment of the present disclosure.
[0092] Referring to FIG. 6, the processor can parse text (e.g., 611, 612, 614, 616) contained in components detected in the original document (610) based, for example, on Optical Character Recognition (OCR). The processor can include the parsed information (622, 624) in output data (620). Here, the text may include at least one of a character, syllable, word, or sentence contained in one or more languages.
[0093] The processor can recognize and output (622) text (612, 614, 616) placed in a specific area (AA) of the original document (610) using a pre-secured document dataset. The pre-secured document dataset may include document data of various languages, including Korean document data.
[0094] Additionally, the processor can parse handwritten type text (611). The processor can recognize and output various types of text included in the original document (610) using a pre-secured document dataset.
[0095] FIG. 7 is a diagram illustrating a method for parsing a formula detected in an original document (710) according to one embodiment of the present disclosure.
[0096] Referring to FIG. 7, the processor can recognize text (712) contained in a component detected in the original document (710) based on, for example, OCR (Optical Character Recognition), and can extract text based on OCR and include it in output data (720) (722). In one embodiment, the processor can detect text (712) by masking the formula (714) of the original document (610).
[0097] After detecting a formula (714) included in the original document (710), the processor can recognize the formula portion using a formula recognition model (730). Based on the recognized formula portion and the recognized text, the processor can convert the detected formula into a specific formula format. The specific formula format may be configured as a Latex type, but the present disclosure is not limited thereto. The processor can include (724) the converted formula portion in the output data (720).
[0098] FIG. 8 is a diagram illustrating a method for parsing a graph (810) detected in an original document according to one embodiment of the present disclosure.
[0099] Referring to FIG. 8, the processor can input the detected graph (810) into a Vision-Language Model (VLM) (830). The Vision-Language Model (830) can generate meta-information including at least one of descriptive information, type information, or analysis information associated with the detected graph (810) and include it in the output data (820).
[0100] Here, the descriptive information may be based in English or Korean, but the present disclosure is not limited thereto. The descriptive information may include at least one of the parameters of the X-axis of the graph (810), the parameters of the Y-axis, or trend information of the X / Y values for each item. The type information may include matching information of Key and Value (e.g., Step 1 is matched to 2015–2019). The analysis information may be generated through the Visual Question Answering (VQA) function of the Vision-Language Model (830). The Vision-Language Model (830) may be trained to perform answers to questions related to objects, scenes, text, actions, or reasoning based on image data. Additionally, the Vision-Language Model (830) may be updated through prompt engineering or fine-tuning and may generate meta-information optimized for Retrieval Augmented Generation (RAG).
[0101] In one embodiment, the vision-language model (830) may be a model pre-trained to generate meta-information including at least one of descriptive information, type information, or analysis information associated with at least one of the detected graph, chart, or diagram when receiving at least one of the detected graph, chart, or diagram. The generated meta-information may include at least one of identification information, structural information (nodes, edges, layers, etc.), color information, shape information, and size information of each graph, chart, or diagram, but the present disclosure is not limited thereto.
[0102] FIG. 9 is a drawing for explaining a method of parsing a table (910) detected in an original document according to one embodiment of the present disclosure.
[0103] Referring to FIG. 9, the processor can recognize the structure of a detected table (910) using a table recognition model. The table recognition model is trained based on a pre-built table dataset to recognize at least one of the number of rows, the number of columns, information associated with merged rows, or information associated with merged columns of the detected table (910). The pre-built table dataset may include tables associated with manufacturing documents, tables with watermarks inserted, etc., but the present disclosure is not limited thereto.
[0104] The processor can convert the detected table into a specific table format (e.g., Markdown) based on at least one of the structure of the recognized table, the recognized text, or the recognized formula part. To do this, the processor can first convert the detected table (910) into a markup-based language (e.g., HTML) (920).
[0105] The processor can generate a specific table format (930, e.g., Markdown) corresponding to a detected table (910) based on a document (920) converted into a markup-based language. At this time, the processor may include unmerged cells to facilitate searching in the RAG. The processor may include the specific table format (930) converted corresponding to the detected table (910) in the output data of a preset data format (JSON).
[0106] FIG. 10 is a drawing illustrating a user interface (1020) for modifying output data of a preset format for a detected component (1010) included in an original document according to one embodiment of the present disclosure.
[0107] Referring to FIG. 10, the detected component (1010) included in the original document may be a table, but the present disclosure is not limited thereto. Based on the execution result of the parsing logic and the placement information of the detected component (1010) within the original document, the processor may process the original document into output data in a preset data format (e.g., JSON). Additionally, based on the processed output data, the processor may generate a converted document (e.g., a PDF document) corresponding to the original document.
[0108] The processor can output the generated converted document or output data in a preset data format to the user interface (1020). When the output data is output to the user interface (1020), the processor can receive user input (1022) to modify the output data through the user interface (1020). Accordingly, the user can edit the output data by comparing the original document with the output data.
[0109] FIG. 11 is a flowchart illustrating a document parsing method (1100) according to one embodiment of the present disclosure.
[0110] In step S1110, the processor can receive the original document.
[0111] In step S1120, the processor inputs the received original document into a layout analysis model to detect one or more components included in the original document.
[0112] In step S1120, the processor can execute parsing logic based on the type of detected component.
[0113] In step S1140, the processor can process the original document into output data of a preset data format based on the execution result of the parsing logic and the placement information of the detected components within the original document.
[0114] Here, the detected component may include at least one of text, formula, table, graph, chart, or diagram.
[0115] The layout analysis model may be a pre-trained model that detects one or more components included in a received original document based on an element pool in which labeling is performed for each component of a document dataset, and determines the type of the one or more detected components.
[0116] The layout analysis model can output placement information of components within the original document based on the layout information of the document dataset.
[0117] To explain step S1130 more specifically, the processor can recognize text contained in the detected component based on Optical Character Recognition (OCR). The text may include at least one of a character, syllable, word, or sentence contained in one or more languages.
[0118] The processor can use a formula recognition model to recognize formula parts in a detected formula, and convert the detected formula into a specific formula format based on the recognized formula parts and recognized text.
[0119] The processor can use a table recognition model to recognize the structure of a detected table and convert the detected table into a specific table format based on at least one of the structure of the recognized table, the recognized text, or the recognized formula part.
[0120] The processor can generate meta-information by inputting at least one of the detected graphs, charts, or diagrams into a vision-language model. The vision-language model may be a model pre-trained to generate meta-information including at least one of descriptive information, type information, or analysis information associated with at least one of the detected graphs, charts, or diagrams.
[0121] Based on the processed output data, the processor can generate a converted document corresponding to the original document.
[0122] The processor may further include a step of outputting the generated converted document to a user interface.
[0123] The processor may further include a step of receiving user input to modify output data through a user interface.
[0124] The method described above may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may continuously store a program executable by a computer, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or multiple hardware components combined, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Furthermore, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0125] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will understand that the various exemplary logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein may be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functional aspects. Whether such functions are implemented in hardware or in software depends on the design requirements imposed on the specific application and the overall system. Those skilled in the art may implement the functions described in various ways for each specific application, but such implementations should not be construed as departing from the scope of the present disclosure.
[0126] In a hardware implementation, the processing units used to perform the techniques may be implemented in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or a combination thereof.
[0127] Accordingly, the various exemplary logic blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors coupled with a DSP core, or any other combination of configurations.
[0128] In firmware and / or software implementations, techniques may be implemented as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage devices, etc. The instructions may be executable by one or more processors, and the processor(s) may be enabled to perform specific aspects of the functions described in this disclosure.
[0129] Where implemented in software, techniques may be stored on a computer-readable medium as one or more instructions or code, or transmitted through a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible by a computer. As a non-limiting example, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium accessible by a computer that can be used to transfer or store desired program code in the form of instructions or data structures. Additionally, any connection is appropriately referred to as a computer-readable medium.
[0130] For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of a medium. As used herein, disk and disc include CD, laser disc, optical disc, DVD (digital versatile disc), floppy disk, and Blu-ray disc, wherein disks usually play data magnetically, whereas discs play data optically using a laser. The above combinations should also be included within the scope of computer-readable media.
[0131] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other known form of storage medium. An exemplary storage medium may be connected to a processor so that the processor can read information from the storage medium or write information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and the storage medium may exist within an ASIC. The ASIC may exist within a user terminal. Alternatively, the processor and the storage medium may exist as separate components within the user terminal.
[0132] Although the embodiments described above have been described as utilizing aspects of the subject matter disclosed herein in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or a distributed computing environment. Furthermore, aspects of the subject matter in the present disclosure may be implemented in a plurality of processing chips or devices, and storage may be similarly affected across a plurality of devices. Such devices may include PCs, network servers, and portable devices.
[0133] Although the present disclosure has been described in relation to some embodiments, various modifications and changes may be made without departing from the scope of the present disclosure as understood by a person skilled in the art to which the invention of the present disclosure pertains. Furthermore, such modifications and changes should be considered to fall within the scope of the claims appended to this specification.
Claims
1. A document parsing method performed by at least one processor, Step of receiving the original document; A step of inputting the received original document into a layout analysis model to detect one or more components included in the original document; A step of executing parsing logic based on the type of the detected component; and A step of processing the original document into output data of a preset data format based on the execution result of the above parsing logic and the placement information of the detected components within the original document. including, Document parsing method.
2. In Paragraph 1, The above-detected component is, including at least one of text, formulas, tables, graphs, charts, or diagrams, Document parsing method.
3. In Paragraph 1, The above layout analysis model is, A pre-trained model for detecting one or more components included in the received original document based on an element pool in which labeling is performed for each component of the document dataset, and for determining the type of the one or more detected components, Document parsing method.
4. In Paragraph 3, The above layout analysis model is, Outputting placement information of the components within the original document based on the layout information of the document data set above, Document parsing method.
5. In Paragraph 2, The step of executing the above parsing logic is, A step of recognizing text contained in the above-mentioned detected components based on OCR (Optical Character Recognition). Includes, The above text is, including at least one of a letter, syllable, word, or sentence included in one or more languages, Document parsing method.
6. In Paragraph 5, The step of executing the above parsing logic is, A step of recognizing a formula part in the detected formula using a formula recognition model; and A step of converting the detected formula into a specific formula format based on the recognized formula part and the recognized text. including, Document parsing method.
7. In Paragraph 6, The step of executing the above parsing logic is, A step of recognizing the structure of the detected table using a table recognition model; and A step of converting the detected table into a specific table format based on at least one of the structure of the recognized table, the recognized text, or the recognized formula part. including, Document parsing method.
8. In Paragraph 6, The step of executing the above parsing logic is, A step of generating meta-information by inputting at least one of the above-detected graphs, charts, or diagrams into a Vision-Language Model. Includes, The above vision-language model is, A model pre-trained to generate meta-information including at least one of descriptive information, type information, or analysis information associated with at least one of the detected graph, chart, or diagram, Document parsing method.
9. In Paragraph 1, A step of generating a converted document corresponding to the original document based on the processed output data above. including, Document parsing method.
10. In Paragraph 9, A further step comprising outputting the generated converted document to a user interface, Document parsing method.
11. In Paragraph 10, A method further comprising the step of receiving user input to modify the output data through the user interface. Document parsing method.
12. A computer-readable, non-transient recording medium storing instructions for executing the method according to paragraph 1 on a computer.
13. In information processing systems, Memory; and A processor connected to the memory and configured to execute at least one computer-readable program contained in the memory. Includes, The above at least one program is, Receive the original document, and The received original document is input into a layout analysis model to detect one or more components included in the original document, and Execute parsing logic based on the type of the detected component above, and Commands for processing the original document into output data of a preset data format based on the execution result of the above parsing logic and the placement information of the above detected components within the original document, Information processing system.
14. In Paragraph 13, The above-detected component is, including at least one of text, formulas, tables, graphs, charts, or diagrams, Information processing system.
15. In Paragraph 13, The above layout analysis model is, A pre-trained model for detecting one or more components included in the received original document based on an element pool in which labeling is performed for each component of the document dataset, and for determining the type of the one or more detected components, Information processing system.
16. In Paragraph 15, The above layout analysis model is, Outputting placement information of the components within the original document based on the layout information of the document data set above, Information processing system.
17. In Paragraph 14, The above at least one program is, It includes commands for recognizing text contained in the above-mentioned detected components based on OCR (Optical Character Recognition), and The above text is, including at least one of a letter, syllable, word, or sentence included in one or more languages, Information processing system.
18. In Paragraph 17, The above at least one program is, Using a formula recognition model, the formula part of the above-detected formula is recognized, and Commands for converting the detected formula into a specific formula format based on the recognized formula part and the recognized text, Information processing system.
19. In Paragraph 18, The above at least one program is, Using a table recognition model, the structure of the detected table is recognized, and Commands for converting the detected table into a specific table format based on at least one of the structure of the recognized table, the recognized text, or the recognized formula part, Information processing system.
20. In Paragraph 18, The above at least one program is, It includes commands for generating meta-information by inputting at least one of the above-detected graph, chart, or diagram into a vision-language model, and The above vision-language model is, A model pre-trained to generate meta-information including at least one of descriptive information, type information, or analysis information associated with at least one of the detected graph, chart, or diagram, Information processing system.