Information processing device, information processing method, and information processing program

The information processing apparatus addresses the challenge of generating accurate structured data from documents by using OCR and VLM to process visual components within documents, ensuring high accuracy and excluding noise from unrelated images.

JP7679538B1Active Publication Date: 2025-05-19SOFTBANK CORPORATION

Patent Information

Application Number
JP2024226270
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-19
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing methods for generating structured data from documents, such as those using Optical Character Recognition (OCR) and Vision-Language Models (VLM), face challenges in accurately processing components with visual information, leading to decreased accuracy and inclusion of noise from unrelated images like logos and illustrations.

Method used

An information processing apparatus that acquires a target document, uses OCR to generate text and attribute information, determines components with visual information, generates input images, calculates similarity with pre-acquired image data using a structured target determination model, and employs a VLM to generate structured data for visual components, while excluding unnecessary information.

Benefits of technology

This approach enables the generation of highly accurate structured data by effectively handling visual components and excluding irrelevant image information, thereby improving the overall accuracy and relevance of the structured data produced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679538000001_ABST
    Figure 0007679538000001_ABST
Patent Text Reader

Abstract

An information processing device, an information processing method, and an information processing program capable of generating structured data with high accuracy are provided. [Solution] The system includes an acquisition unit that acquires a target document, a string recognition unit that generates text information, attribute information, and coordinate information of components in the target document, a first determination unit that determines whether the components contain visual information, an image generation unit that cuts out the components and uses them as an input image, a calculation unit that calculates the similarity between the input image and previously acquired image data by image recognition using a structured object determination model, a second determination unit that determines whether the image is an input image based on the similarity, an output unit that generates structured data of the input image using a visual language model, and an integration unit that integrates the text information generated by the string recognition unit and the structured data generated by the output unit, and the structured object determination model is constructed using one or more previously acquired image data to which tags related to image information have been added.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Conventionally, a technique for generating structured data from a target document has been disclosed (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0004] An information processing apparatus according to an embodiment of the present invention includes an acquisition unit that acquires a target document, a character string recognition unit that generates text information, attribute information, and coordinate information for each of one or more components in the target document, a first determination unit that determines, based on the text information, the attribute information, and the coordinate information, whether each of the one or more components includes visual information, an image generation unit that uses the coordinate information to set one or more components determined to be components to be cut out by the first determination unit as one or more input images, a calculation unit that calculates the similarity between each of the one or more input images and one or more image data acquired in advance by image recognition using a structured target determination model, a second determination unit that determines, based on the similarity, whether each of the one or more input images is an input image for generating structured data, an output unit that generates structured data of the input image by a vision-language model, an integration unit that integrates the text information generated by the character string recognition unit and the structured data generated by the output unit, and a storage unit that stores the structured target determination model, wherein the structured target determination model is constructed using one or more image data acquired in advance to which tags related to image information are assigned.

[0005] In an information processing apparatus according to an embodiment of the present invention, the similarity may be calculated from a comparison between the contours of each of one or more input images and the feature amounts of the contours of each of one or more image data acquired in advance.

[0006] In an information processing apparatus according to an embodiment of the present invention, a vision language model may be used as a structured object determination model.

[0007] In an information processing apparatus according to an embodiment of the present invention, the one or more image data acquired in advance may include at least any one of a graph, a box plot, a tree diagram, a table, or a flowchart.

[0008] In an information processing apparatus according to an embodiment of the present invention, the first determination unit may determine that a component whose attribute information is either "figure" or "table" is a component to be cut out.

[0009] In an information processing apparatus according to an embodiment of the present invention, an input unit may be further provided, and the input unit may receive an input of the coordinate information of a desired component among one or more components.

[0010] An information processing method according to an embodiment of the present invention is an information processing apparatus including a storage unit that stores a structured object determination model constructed using one or more pre-acquired image data to which tags related to image information are attached. The information processing method includes: an acquisition step of acquiring a target document; a character string recognition step of generating text information, attribute information, and coordinate information for one or more components in the target document; a first determination step of determining, for each of the one or more components, whether the component includes visual information based on the text information, the attribute information, and the coordinate information; an image generation step of using the one or more components determined to be components to be cut out in the first determination step as one or more input images based on the coordinate information; a calculation step of calculating a similarity between each of the one or more input images and the one or more pre-acquired image data by image recognition using the structured object determination model; a second determination step of determining, for each of the one or more input images, whether the input image is an input image for generating structured data based on the similarity; an output step of generating structured data of the input image by a visual language model; and an integration step of integrating the text information generated in the character string recognition step and the structured data generated by the output unit.

[0011] An information processing program according to an embodiment of the present invention causes a computer including a storage unit that stores a structured target determination model constructed using one or more pre-acquired image data to which tags related to image information are assigned, to have an acquisition function for acquiring a target document, a character string recognition function for generating text information, attribute information, and coordinate information for one or more components in the target document, a first determination function for determining whether each of the one or more components includes visual information based on the text information, the attribute information, and the coordinate information, an image generation function for using the one or more components determined to be components to be cut out in the first determination function as one or more input images based on the coordinate information, a calculation function for calculating the similarity between each of the one or more input images and the one or more pre-acquired image data by image recognition using the structured target determination model, a second determination function for determining whether each of the one or more input images is an input image for generating structured data based on the similarity, an output function for generating structured data of the input image by a visual language model, and an integration function for integrating the text information generated by the character string recognition function and the structured data generated by the output unit.

Effects of the Invention

[0012] According to the present invention, it is possible to provide an information processing apparatus, an information processing method, and an information processing program capable of generating highly accurate structured data.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0014] Next, embodiments of the present invention will be described with reference to the drawings. In the description of the drawings according to the embodiments, the same or similar parts are denoted by the same or similar reference numerals. Of course, the drawings also include parts having different relationships with each other.

[0015] Also, the embodiments illustrate an apparatus and a method for embodying the technical idea of the present invention, and the technical idea of the present invention does not specify the configuration of each component as follows. The technical idea of the present invention can be variously modified within the technical scope defined by the claims described in the claims.

[0016] (First Embodiment) The information processing apparatus according to the present embodiment generates structured data for a target document using optical character recognition (OCR), and then uses a vision language model (VLM) to generate structured data for data having an attribute that is not a character string, such as a chart, by OCR, and replaces the structured data corresponding to the data having an attribute that is not a character string in the structured data generated by OCR. The information processing apparatus according to the present embodiment generates structured data from a target document such as a form by a method combining OCR and VLM. When generating structured data using an LLM for data having an attribute that is not a character string, the information processing apparatus according to the present embodiment excludes images such as logos and illustrations from the structuring target and generates structured data.

[0017] As methods for generating structured data from documents such as forms, methods using OCR, methods using LLM, and methods combining these are generally known. OCR is a method of capturing a document with printed or handwritten characters as image data by an optical method and converting the characters in the document into text data.

[0018] In the generation of structured data using OCR for a document, the character strings in the document are detected, and the analysis of the attribute information of the detected character strings and the analysis of the coordinate information of the detected character strings are performed. However, in the generation of structured data using OCR for charts in a document, the visual meaning of the charts cannot be structured. Thus, a method combining OCR and LLM is known, in which structured data is generated using OCR for the character strings in the document and structured data is generated using LLM for the data that is not a character string in the document. In the generation of structured data using a method combining OCR and LLM for a document, when generating structured data using an LLM for data having an attribute that is not a character string, the accuracy of structured data generation may decrease, such as when images such as logos and illustrations that are not directly related to the content of the document become noise. To address these problems, the information processing apparatus according to the present embodiment enables the provision of an information processing apparatus, an information processing method, and an information processing program capable of generating highly accurate structured data.

[0019] <Configuration> FIG. 1 shows an example of the configuration and functional units of the information processing apparatus 10 according to the present embodiment. The information processing apparatus 10 shown in FIG. 1 includes a CPU 101 for executing various operations, a ROM 102 for storing programs for processing, a RAM 103 for storing data and the like, a storage unit 104 for storing various data and operation results and the like, an I / O (input / output interface) 105, a display unit 106, an input unit 107, and the like.

[0020] The I / O 105 is an interface for communication (transmission and reception), a buffer, and the like.

[0021] In addition, a keyboard, a mouse, or the like for input may be connected to the information processing apparatus 10 according to the present embodiment.

[0022] The information processing apparatus 10 is various electronic computers (computing resources) such as a mobile terminal, a personal computer (PC), a mainframe, a workstation, and a cloud computing system.

[0023] Furthermore, the block diagram of FIG. 1 shows the functional units within the CPU 101. When each functional unit of the CPU 101 is realized by software, the CPU 101 is realized by executing the instructions of a program that is software for realizing each function. Specifically, it includes an acquisition unit 108, a first generation unit 109, a first determination unit 110, an image generation unit 111, a calculation unit 112, a second determination unit 113, a second generation unit 114, an integration unit 115, and the like. The storage unit 104 includes a structured object determination model storage unit 116.

[0024] The acquisition unit 108 acquires a target document.

[0025] The first generation unit 109 uses optical character recognition (OCR) to generate text information, attribute information, and coordinate information of at least one or more components in the target document, and then generates structured data based on the generated text information, attribute information, and coordinate information.

[0026] The first determination unit 110 determines whether to generate image data of a component as an input image based on the text information, attribute information, and coordinate information.

[0027] The image generation unit 111 generates at least one or more input images of the determined components determined by the first determination unit to be components for which image data is to be generated.

[0028] The calculation unit 112 calculates the similarity between the input image and at least one or more image data acquired in advance by image recognition using a structured target determination model.

[0029] The second determination unit 113 determines whether the input image is an input image for which structured data is to be generated based on the similarity.

[0030] The second generation unit 114 generates structured data of the input image using a vision-language model.

[0031] The integration unit 115 integrates the structured data generated by the first generation unit 109 and the structured data generated by the second generation unit 114.

[0032] The structured target determination model storage unit 116 stores a structured target determination model.

[0033] The structured target determination model is constructed using one or more pre-acquired image data to which tags related to the pre-acquired image data are assigned.

[0034] The operations of each functional unit within the CPU 101 will be described in detail below. The target document acquired by the acquisition unit 108 is a document in which strings, images, numbers, graphs, tables, etc. are described, and may be, for example, a form, a calculation sheet, a mathematical formula, a data file, etc. FIG. 2 shows an example of the target document acquired by the acquisition unit 108. The target document 20 shown in FIG. 2 is composed of a title 21 described as "Title", a plurality of strings 22, a table 23, a graph 24, etc. The plurality of strings 22 are composed of the strings "1. Chapter 1", "This chapter is about...", "2. Purpose", "2.1. Table", "Examples of tables are shown...".

[0035] The acquisition unit 108 acquires the target document from an external device (not shown), for example, via the I / O 105. The external device may be, for example, a server and a user terminal, etc. The user terminal is a terminal used by the user of the information processing apparatus 10, and as an example, may be a desktop, a laptop, a tablet, a smartphone, etc. Also, when the target document is recorded in an external memory (not shown) and the external memory is connected to the interface of the information processing apparatus 10, the acquisition unit 108 may acquire the target document from the external memory. Further, when there is article information (article file) including the target document and the target document is selected via the input unit 121 and the user terminal, etc. (when the area where the target document is described is specified), the acquisition unit 108 may acquire the selected target document (using the specified area as the target document).

[0036] The first generation unit 109 structures the target document, that is, generates text information of at least one or more components within the target document, attribute information of the components, and coordinate information of the components, and generates structured data based on the generated at least one or more text information, attribute information, and coordinate information.

[0037] In the present embodiment, at least one or more components within the target document refer to the elements described in the target document, such as strings, images, numbers, graphs, tables, etc. For example, the document shown in FIG. 2 is composed of components such as the title 21, a plurality of strings 22, a table 23, a graph 24, etc.

[0038] The text information generated by the first generation unit 109 is a character string included in the constituent element. For example, when the constituent element is a graph that does not include a character string, or a logo or illustration, etc., the constituent element does not necessarily include a character string. When the constituent element does not include a character string, the first generation unit 109 generates attribute information and coordinate information. When a predetermined number of constituent elements among at least one or more constituent elements in the target document do not include a character string, the first generation unit 109 generates text information of the constituent element including the character string, attribute information of the constituent element, and coordinate information of the constituent element, and further generates attribute information of the constituent element not including the character string, and coordinate information of the constituent element. After that, structured data is generated based on the generated text information, attribute information, and coordinate information.

[0039] The attribute information generated by the first generation unit 109 is the type when classifying the constituent element. For example, the attribute of the title 21 is title, and the attribute of the character string 22 is "chapter title" for the character string "Chapter 1", "main text" for the character string "This chapter is about...", "chapter title" for the character string "2. Purpose", "section title" for the character string "2.1. Table", "main text" for the character string "Show an example of the table...", and the attribute of the table 23 is "table".

[0040] The coordinate information generated by the first generation unit 109 is information for specifying the position of the constituent element in the target document 20, and may include, for example, coordinates indicating the position within the page of the target document 20, the shape of the constituent element, information such as the vertical and horizontal lengths.

[0041] FIG. 3 shows the text information, attribute information, and coordinate information generated by structuring the title 21 described in the target document 20 shown in FIG. 2 by the first generation unit 109. In FIG. 3, as the text information generated by the first generation unit 109, "character string: \"title\"", as the attribute information, "attribute: title", and as the coordinate information, "coordinate information: 3.45, 1.23..." are shown.

[0042] FIG. 4 shows an example of the structured data generated based on the text information, attribute information, and coordinate information of each component in the target document 20 shown in FIG. 2 by the first generation unit 109. Based on the coordinate information of each component generated by the first generation unit 109, the content described from the upper left to the lower right within the page of the target document 20 is, in the structured data 40 shown in FIG. 4, in order from the top, in a format based on the attribute information of each component, with a character string described based on the text information of each component. For example, character strings with attributes of "title", "body text", "chapter title", and "section title" have corresponding character strings displayed, and character strings with an attribute of "table" are displayed in table format.

[0043] The structured data shown in FIG. 4 is an example where the first generation unit 109 can execute the structuring of the target document without problems. However, since OCR cannot structure visual information such as graphs, illustrations, logos, etc., the first generation unit 109 cannot generate accurate structured data for components including visual information.

[0044] FIG. 5 shows an example when the first generation unit 109 generates structured data for a component including visual information. FIG. 5(a) is a bar graph, FIG. 5(b) is the text information, attribute information, and coordinate information generated by the first generation unit 109 with the bar graph shown in FIG. 5(a) as a component of the target document, and FIG. 5(c) is the structured data generated by the first generation unit 109 based on the text information, attribute information, and coordinate information. Since the first generation unit 109 cannot structure visual information, it cannot distinguish the scale of the y-axis in the graph from the values of the bar graph. In the structured data shown in FIG. 5(c), the values of the scale of the y-axis in the graph and the values of the bar graph are displayed mixedly. The visual information possessed by the bar graph shown in FIG. 5(a) is missing in the structured data shown in FIG. 5(c), and information obtained from the bar graph shown in FIG. 5(a), such as the values of each bar graph for each month, etc., cannot be obtained from the structured data shown in FIG. 5(c).

[0045] In contrast, the information processing apparatus 10 according to the present embodiment generates structured data for components including visual information using a vision language model (VLM). At this time, if the attribute information of the component including visual information is "figure" or "table", it is possible to accurately generate structured data from the component using an LLM. However, if the attribute information of the component including visual information is not "figure" or "table" and the component is information unnecessary for generating structured data such as a logo or an illustration, it is not possible to accurately generate structured data. Therefore, as described below, the information processing apparatus 10 according to the present embodiment removes unnecessary information for generating structured data from the component including visual information and then generates structured data using a VLM.

[0046] The first determination unit 110 determines whether the component includes visual information based on the text information, attribute information, and coordinate information generated by the first generation unit 109. Here, the visual information is information that is not expressed by a character string but is visually expressed by the information that the component has. In the present embodiment, the first determination unit 110 may determine that the component includes visual information when the attribute information is "figure" or "table". Note that "figure" attribute information is assigned to an image.

[0047] The image generation unit 111 generates, as an input image for inputting the image data of the determined component to the calculation unit 112, based on the coordinate information of the determined component for the determined component determined by the first determination unit 110 to include visual information. In the example shown in FIG. 5, the image generation unit 111 uses the entire bar graph in FIG. 5(a) as the input image. The image generation unit 111 transmits the generated input image to the calculation unit 112.

[0048] The calculation unit 112 compares the input image received from the image generation unit 111 with at least one or more pieces of image data acquired in advance by image recognition using the structured object determination model, and calculates the similarity between the input image and the at least one or more pieces of image data acquired in advance. The one or more pieces of image data acquired in advance can be components of the target document 20 and can be data of images that can be structured, and can be, for example, at least any one of a graph, a box diagram, a tree diagram, a table, or a flowchart.

[0049] The structured object determination model is constructed using one or more pieces of image data acquired in advance, with information regarding the image data attached as tags. The information regarding the image data may include details and shapes of the images, such as the type of the image, e.g., graph, box diagram, tree diagram, table, or flowchart, and for example, if the image data is a graph, the details and shape of the graph. Note that a visual language model (VLM) may be used as the structured object determination model.

[0050] When the calculation unit 112 compares the input image with at least one or more pieces of image data acquired in advance by image recognition using the structured object determination model, it may remove the color information of the input image and the at least one or more pieces of image data acquired in advance, and compare the feature amounts of the contours of the input image with the feature amounts of the contours of the at least one or more pieces of image data acquired in advance. Thereby, the calculation unit 112 can perform the comparison more accurately. At this time, the similarity is calculated from the comparison of the feature amounts of the contours of each of the one or more input images with the feature amounts of the contours of each of the one or more pieces of image data acquired in advance.

[0051] The second determination unit 113 determines whether the input image is an input image for generating structured data based on the similarity calculated by the calculation unit 112. If one or more pieces of pre-acquired image data used when constructing the structured object determination model can be components of the target document 20 and can be the target of structuring, the second determination unit 113 may determine that an input image with a similarity higher than a predetermined threshold is an input image for generating structured data. Thereby, images that are not targets for structuring, such as logos and illustrations, can be removed as input images for generating structured data.

[0052] The second determination unit 113 may use, as the similarity used for determining whether the input image is an input image for generating structured data, the highest similarity among the similarities between the input image calculated by the calculation unit 112 and each of at least one or more pieces of pre-acquired image data.

[0053] The second generation unit 114 generates structured data for the input image determined by the second determination unit 113 to be an input image for generating structured data using the vision-language model. Fig. 6 shows an example of the structured data generated by the second generation unit 114 when the bar graph in Fig. 5(a) is used as the input image. The structured data shown in Fig. 6 indicates the values of the bar graph with respect to the x-axis. The visual information included in Fig. 5(a), that is, the y values represented by the bar graphs for the respective x values shown in Fig. 5(a), is shown in Fig. 6.

[0054] The integration unit 115 integrates the structured data generated by the first generation unit 109 and the structured data generated by the second generation unit 114 by replacing the structured data corresponding to the components determined by the second determination unit 113 to be input images for generating structured data among the structured data generated by the first generation unit 109 with the structured data generated by the second generation unit 114 after the first determination unit 110 determines that the components include visual information.

[0055] Note that, among the structured data generated by the integration unit 115, for the structured data corresponding to an image that is not a structured target, such as a logo or an illustration, which the second determination unit 113 determines is not the input image for generating the structured data, the integration unit 115 may delete it from the structured data generated by the first generation unit 109.

[0056] The information processing method according to this embodiment will be described with reference to the flowchart of FIG. 7.

[0057] In step S701, the acquisition unit 108 acquires a target document (acquisition step).

[0058] In step S702, the first generation unit 109 generates text information of at least one or more components in the target document, attribute information of the components, and coordinate information of the components (first generation step).

[0059] In step S703, the first determination unit 110 determines whether to generate image data of the component as input image based on the text information, attribute information, and coordinate information (first determination step).

[0060] In step S704, the image generation unit 111 generates at least one or more input images of the determined components determined by the first determination unit to be components for which image data is to be generated (image generation step).

[0061] In step S705, the calculation unit 112 calculates the similarity between the input image and at least one or more image data acquired in advance by image recognition using the structured target determination model (calculation step).

[0062] In step S706, the second determination unit 113 determines whether the input image is an input image for generating structured data based on the similarity (second determination step).

[0063] In step S707, the second generation unit 114 generates structured data of the input image using a visual language model (second generation step).

[0064] In step S708, the integration unit 115 integrates the text information generated by the first generation unit 109 and the structured data generated by the second generation unit 114 (integration step).

[0065] (Second Embodiment) In the information processing apparatus 10 according to the first embodiment, the image generation unit 111 generates, as an input image for inputting the image data of the determined component to the calculation unit 112, the image data of the determined component based on the coordinate information of the determined component for the determined component determined by the first determination unit 110 to be a component including visual information. At this time, when generating the image data of the determined component as the input image, the range of the image data may not be set to a sufficiently correct range that allows the calculation unit 112, the second determination unit 113, and the second generation unit 114 to perform accurate operations.

[0066] Specifically, as shown in FIG. 8(a), there is a case where a part of the graph 81 is set as the range 82 of the image data. When the image of the area of the range 82 shown in FIG. 8(a) is used as the input image, the information of the graph 81 may not be accurately reflected when the second generation unit 114 generates the structured data. Further, as shown in FIG. 8(b), when there are a plurality of consecutive graphs in the target document and only one graph has a legend displayed among the plurality of consecutive graphs and the other graphs do not have a legend displayed, the image generation unit 111 may generate independent input images for each of the plurality of graphs. When the second generation unit 114 generates the structured data, the structured data of the graph without a legend may not be accurately generated.

[0067] In contrast, the information processing apparatus according to the present embodiment may further include an input unit. The input unit may receive input of coordinate information of a desired component among one or more components. A user can input the coordinate information of a desired component to the input unit. In the case shown in FIG. 8(a), coordinate information can be input such that the entire graph 81 is set as the range 83 of the image data. In the case shown in FIG. 8(b), for example, coordinate information can be input such that all consecutive graphs are set as one input image.

[0068] After receiving the input of the coordinate information of the desired component, the input unit transmits the received coordinate information of the desired component to the image generation unit 111. The image generation unit 111 generates an input image of the component based on the received coordinate information of the component.

[0069] Needless to say, the present invention includes various embodiments and the like not described herein. Therefore, the technical scope of the present invention is defined only by the invention specific matters according to the reasonable claims based on the above description.

[0070] The program of each embodiment of the present disclosure may be provided in a state stored in a storage medium readable by the information processing apparatus. The storage medium can store the program in a "non-transitory tangible medium". The program includes, for example, a software program and a control program. When each functional unit of the information processing apparatus is realized by software, the information processing apparatus functions as the acquisition unit 108, the first generation unit 109, the first determination unit 110, the image generation unit 111, the calculation unit 112, the second determination unit 113, the second generation unit 114, the integration unit 115, and the structured object determination model storage unit 116 by the processor executing the program loaded on the memory.

[0071] The memory medium, where appropriate, can include one or more semiconductor-based, or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs), application-specific ICs (ASICs), etc.), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable memory medium, or a suitable combination of two or more of these. The memory medium, where appropriate, can be volatile, non-volatile, or a combination of volatile and non-volatile.

[0072] Also, the program of the present disclosure may be provided to the information processing apparatus via any transmission medium (such as a communication network or a broadcast wave) capable of transmitting the program.

[0073] Also, each embodiment of the present disclosure can also be realized in the form of a data signal embedded in a carrier wave in which the program is embodied by electronic transmission. Note that the program of the present disclosure may be implemented using, for example, script languages such as JavaScript (registered trademark), Python (registered trademark), etc., C language, Go language, Swift (registered trademark), Koltin (registered trademark), Java (registered trademark), etc.

[0074] According to each aspect of the present disclosure described above, by providing an information processing apparatus capable of generating highly accurate structured data, it is possible to contribute to the achievement of Goal 9, "Build the infrastructure for industry and innovation," of the Sustainable Development Goals (SDGs).

Explanation of Reference Numerals

[0075] 10 Information processing apparatus 101 CPU 102 ROM 103 RAM 104 Storage unit 105 I / O (Input / Output Interface) 106 Display unit 107 Input unit 108 Acquisition unit 109 First generation unit 110 First determination unit 111 Image generation unit 112 Calculation unit 113 Second determination unit 114 Second generation unit 115 Integration unit 116 Structured object determination model storage unit 20 Target document 21 Title 22 Multiple strings 23 Table 81 Graph 82, 83 Range

Claims

1. an acquisition unit for acquiring a target document; a first generation unit that generates text information, attribute information, and coordinate information for each of one or more components in the target document, and then generates structured data based on the generated text information, attribute information, and coordinate information; a first determination unit that determines, for each of the one or more components, whether or not the component includes visual information based on the text information, the attribute information, and the coordinate information; an image generating unit that cuts out one or more components determined by the first determining unit to be components including visual information based on the coordinate information to generate one or more input images; a calculation unit that calculates a similarity between each of the one or more input images and one or more image data acquired in advance by image recognition using a structured object determination model; a second determination unit that determines, based on the similarity, whether each of the one or more input images is an input image for generating structured data; a second generation unit that generates structured data of the input image using a visual language model; an integration unit that integrates the structured data generated by the first generation unit and the structured data generated by the second generation unit; A storage unit that stores the structured object determination model; wherein the structured target determination model is constructed using the one or more pieces of image data that have been acquired in advance and to which tags related to image information have been added.

2. 2 . The information processing apparatus according to claim 1 , wherein the degree of similarity is calculated by comparing a feature amount of each contour of the one or more input images with a feature amount of each contour of one or more image data acquired in advance.

3. 2. The information processing apparatus according to claim 1, wherein the visual language model is used as the structured object determination model.

4. 2 . The information processing apparatus according to claim 1 , wherein the one or more pieces of pre-acquired image data include at least one of a graph, a box diagram, a tree diagram, a table, and a flow chart.

5. The information processing device according to claim 1 , wherein the first determination unit determines that the one or more components whose attribute information is any one of "figure", "table", or "image" are components to be cut out.

6. The information processing apparatus according to claim 1 , further comprising an input unit, said input unit receiving input of coordinate information of a desired component among said one or more components.

7. An information processing device including a storage unit that stores a structured object determination model constructed using one or more image data items acquired in advance and to which tags related to image information are added, An acquisition step of acquiring a target document; a first generation step of generating text information, attribute information, and coordinate information for one or more components in the target document, and then generating structured data based on the generated text information, attribute information, and coordinate information; a first determination step of determining, for each of the one or more components, whether or not the component includes visual information based on the text information, the attribute information, and the coordinate information; an image generating step of generating one or more input images based on the coordinate information for one or more components determined in the first determination step to be components including visual information; a calculation step of calculating a similarity between each of the one or more input images and one or more image data acquired in advance by image recognition using the structured object determination model; a second determination step of determining, based on the similarity, whether each of the one or more input images is an input image for generating structured data; a second generation step of generating structured data of the input image using a visual language model; an integration step of integrating the structured data generated in the first generation step and the structured data generated in the second generation step; An information processing method comprising:

8. A computer including a storage unit that stores a structured object determination model constructed using one or more pieces of image data previously acquired to which tags related to image information are added, An acquisition function for acquiring a target document; a first generation function that generates text information, attribute information, and coordinate information for one or more components in the target document, and then generates structured data based on the generated text information, attribute information, and coordinate information; a first determination function for determining, for each of the one or more components, whether or not the component includes visual information based on the text information, the attribute information, and the coordinate information; an image generating function that generates one or more input images based on the coordinate information from one or more components determined to be components including visual information in the first determining function; a calculation function for calculating a similarity between each of the one or more input images and one or more image data acquired in advance by image recognition using the structured object determination model; a second determination function for determining, based on the similarity, whether each of the one or more input images is an input image for generating structured data; a second generating function for generating structured data of the input image by a visual language model; an integration function that integrates the structured data generated by the first generation function and the structured data generated by the second generation function; An information processing program characterized by realizing the above.

Citation Information

Patent Citations

  • Visual question and answer processing method

    CN118898240A

  • Deep document processing with self-supervised learning

    EP4002296A1

  • Method and device for recognizing character

    JP1997251511A

  • Document digitalization architecture by multi-model deep learning and document image processing program

    JP2022104411A

Cited By

  • Information processing device, information processing method, and program

    JP7795032B1

  • Information processing device, information processing method, and program

    JP7795033B1