Information processing device and information processing program

The information processing device and program address the challenge of ordering and extracting attribute values from split tables by grouping similar tables and using header row attributes to align and maintain continuity, ensuring accurate reconstruction of the original table structure.

JP7739895B2Active Publication Date: 2025-09-17FUJIFILM BUSINESS INNOVATION CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021156133
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-24
Publication Date
2025-09-17
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

When a table is split across multiple pages due to a page break, the subsequent tables lack a header row, making it difficult to determine the order and extract attribute values correctly from the split tables.

Method used

An information processing device and program that groups tables with similar structures and uses attribute values from the header row of one table to align and associate subsequent tables, ensuring continuity and consistency in the order of the tables, even if they lack a header row.

Benefits of technology

Enables the correct arrangement and extraction of attribute values from tables without header rows, maintaining order and consistency, thus accurately reconstructing the original table structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007739895000001
    Figure 0007739895000001
  • Figure 0007739895000002
    Figure 0007739895000002
  • Figure 0007739895000003
    Figure 0007739895000003
Patent Text Reader

Abstract

To acquire an attribute value of an attribute included in a heading line from a table not added with the heading line after sequentially arranging the plurality of tables even when there is the table not added with the heading line expressing the attribute of the attribute value.SOLUTION: An information processing device 10 extracts a table group made of a heading table 4A having common structural information and at least one or more succeeding tables 4B from a plurality of imaged tables 4, associates each table 4 with each other such that the arrangement order of the tables 4 included in the table group becomes continuous by using an attribute value of an order attribute expressing continuity of the heading table 4A and the succeeding tables 4B, and acquires an attribute value corresponding to each attribute included in a heading line 6 from a series of tables 4 having the continuity.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device and an information processing program. [Background technology]

[0002] Patent document 1 discloses a table information reading device that includes a document input unit that accepts input of document information and extracts table configuration information that represents the configuration of a table included in the document information, a table structure estimation unit that estimates table structural information that represents whether each cell that makes up the table corresponds to a heading element, a content entry element, or other element based on the table configuration information, and a table element relationship determination unit that determines the relationship between the information included in the heading element and the information included in the content entry element based on the positional relationship in the table between the cell that corresponds to the heading element and the cell that corresponds to the content entry element. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-155054 Summary of the Invention [Problem to be solved by the invention]

[0004] If a table does not fit on one page and, for example, a page break causes a single table to be split into multiple tables, only the first table will contain a header row indicating the attributes of the contents of the table, while the subsequent tables will not contain header rows. Therefore, if multiple tables of different types are split and each of the split tables is printed on a separate page of paper, and the paper is then scanned or otherwise imaged in a manner that does not allow the sheets to be arranged in the order in which they were split, it may be difficult to tell which tables constituted the same type of table before the split.

[0005] In such cases, conventionally, the table structure, such as the number of columns, column width, type of borders, background color, and type of font, is referenced to construct the pre-split table from each of the split tables, and attribute values ​​corresponding to the attributes contained in the header row are extracted from the table.

[0006] However, when the pre-split table is constructed from each post-split table by referencing the table structure, the order of the tables cannot be determined from the table structure, so it is not guaranteed that the tables constructed by associating tables with the same structure will be arranged in the same order as the pre-split table.

[0007] The present invention aims to provide an information processing device and an information processing program that can arrange multiple tables in order and acquire attribute values ​​of attributes included in a header row from a table without a header row, even if the multiple tables include a table without a header row that indicates the attributes of the attribute value. [Means for solving the problem]

[0008] An information processing device according to a first aspect includes a processor, which extracts from a plurality of imaged tables a group of tables consisting of a first table having a header row representing attributes of attribute values ​​in the table, and at least one or more second tables without the header row, which have common characteristics regarding the table structure, associates each table included in the group of tables so that the order of the tables is continuous using an attribute value of an attribute included in the header row that represents the continuity of the first table and the second table, and obtains attribute values ​​corresponding to each attribute included in the header row from a series of tables consisting of the first table and the second table with consecutive order.

[0009] In the information processing device of the second aspect, in the information processing device of the first aspect, the processor detects whether there is an excess or deficiency in the tables that make up the series of tables using attribute values ​​of attributes included in the header row that represent consistency regarding excess or deficiency in the tables that make up the series of tables, and obtains attribute values ​​corresponding to each attribute included in the header row from the series of tables in which no excess or deficiency was detected.

[0010] In the information processing device of the third aspect, in the information processing device of the second aspect, the attribute representing consistency is an attribute in which the sum of each attribute value in the series of tables corresponding to any attribute included in the header row is an attribute that is pre-displayed in an image including any of the tables that make up the series of tables, and the processor detects that there are no excesses or deficiencies in the tables that make up the series of tables if the sum of each attribute value corresponding to the attribute representing consistency in each of the tables that make up the series of tables matches the sum that is pre-displayed in an image including any of the tables that make up the series of tables.

[0011] An information processing device according to a fourth aspect is an information processing device according to any one of the first to third aspects, wherein the attribute representing the continuity is an attribute representing the order of attribute values ​​in the group of tables, and the processor detects that one table and the other table are contiguous when the attribute value of the attribute representing the continuity in the last row of one of two tables selected from the group of tables and the attribute value of the attribute representing the continuity in the first row of the other table follow a predetermined regularity as the state in which attribute values ​​are arranged.

[0012] In the information processing device of the fifth aspect, in the information processing device of the first aspect, if the header row does not include an attribute representing the continuity, the processor detects whether there is an excess or deficiency in the tables that make up the table group using the attribute values ​​of the attributes included in the header row that represent consistency regarding excess or deficiency in the tables that make up the table group, and obtains attribute values ​​corresponding to each attribute included in the header row from the table group in which no excess or deficiency was detected.

[0013] In a sixth aspect of the information processing device, in the information processing device of the fifth aspect, if the header row does not include an attribute representing the consistency, the processor uses external information, which is information existing outside the first table and the second table and represents the continuity of the first table and the second table, to associate each table included in the group of tables so that the order of the tables is continuous, and obtains attribute values ​​corresponding to each attribute included in the header row from a series of tables formed by the first table and the second table that are continuous in order.

[0014] In an information processing device according to a seventh aspect, in the information processing device according to the sixth aspect, the processor uses the page number of the page containing either the first table or the second table, or a table number representing the order of the first table and the second table, as the external information, and associates each table so that the order of each table included in the group of tables is continuous.

[0015] An information processing device according to an eighth aspect includes a processor that extracts from a plurality of imaged tables a group of tables consisting of a first table with a header row representing the attributes of the attribute values ​​in the table, and at least one or more second tables without the header row, which have common characteristics regarding the table structure, and detects whether there is a surplus or deficiency in the tables that make up the group of tables using attribute values ​​of attributes contained in the header row that represent consistency regarding surpluses or deficiencies in the tables that make up the group of tables, and obtains attribute values ​​corresponding to each attribute contained in the header row from the group of tables in which no surplus or deficiency was detected.

[0016] An information processing program according to a ninth aspect is a program for causing a computer to execute a process of extracting, from a plurality of tables that have been imaged, a group of tables consisting of a first table having a header row representing attributes of attribute values ​​in the table and having common characteristics regarding the table structure, and at least one or more second tables that do not have the header row, associating each table in the group of tables so that the order of the tables is continuous using an attribute value of an attribute contained in the header row that represents the continuity of the first table and the second table, and obtaining attribute values ​​corresponding to each attribute contained in the header row from a series of tables consisting of the first table and the second table that are continuous in order.

[0017] An information processing program according to a tenth aspect is a program for causing a computer to execute a process of extracting from a plurality of images of tables a group of tables consisting of a first table with a header row representing the attributes of the attribute values ​​in the table, which have common characteristics regarding the table structure, and at least one or more second tables without the header row, detecting whether there is a surplus or deficiency in the tables that make up the group of tables using attribute values ​​of attributes contained in the header row that represent consistency regarding surpluses or deficiencies in the tables that make up the group of tables, and obtaining attribute values ​​corresponding to each attribute contained in the header row from the group of tables in which no surplus or deficiency was detected. [Effects of the Invention]

[0018] According to the first, eighth, ninth, and tenth aspects, even if there is a table among multiple tables that does not have a header row that represents the attribute of the attribute value, it is possible to arrange the multiple tables in order and obtain the attribute value of the attribute included in the header row from the table that does not have a header row.

[0019] According to the second aspect, it is possible to obtain an effect that attribute values ​​corresponding to each attribute included in the header row can be obtained from a series of tables that are arranged in order without excess or deficiency.

[0020] According to the third aspect, it is possible to detect whether or not a series of tables are consistent from the sum of each attribute value corresponding to any attribute included in the header row.

[0021] According to the fourth aspect, it is possible to detect whether there is continuity between previous and next tables by using the attribute values ​​of the attributes in the table that represents the order.

[0022] According to the fifth aspect, even if the header row does not contain an attribute that indicates the continuity of the table, it is possible to construct the table before division and obtain attribute values ​​corresponding to each attribute contained in the header row.

[0023] According to the sixth aspect, even if the tables do not contain any attributes that represent continuity or consistency, it is possible to arrange and associate a plurality of tables in order.

[0024] According to the seventh aspect, there is an effect that external information can be used as information for arranging a plurality of tables in order. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 10 is a diagram showing an example of a heading table. [Figure 2] FIG. 10 is a diagram illustrating an example of a subsequent table. [Figure 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of an information processing device. [Figure 4] FIG. 1 is a diagram illustrating an example of the configuration of a main part of an electrical system in an information processing device configured using a computer. [Figure 5] 10 is a flowchart showing an example of the flow of an extraction process according to the first embodiment. [Figure 6] 10 is a flowchart showing an example of the flow of an extraction process according to the second embodiment. [Figure 7] 11 is a flowchart showing an example of the flow of an extraction process according to the third embodiment. [Figure 8]13 is a flowchart showing an example of the flow of an extraction process according to the fourth embodiment. [Figure 9] FIG. 10 is a diagram showing an example of a page number. [Figure 10] FIG. 10 is a diagram showing an example of a table number. [Figure 11] FIG. 10 is a diagram showing an example of a table divided within a page. DETAILED DESCRIPTION OF THE INVENTION

[0026] The present embodiment will be described below with reference to the drawings. Note that the same components and processes are given the same reference numerals throughout the drawings, and redundant description will be omitted.

[0027] The information processing device 10 of this embodiment obtains attribute values ​​corresponding to each attribute contained in the header row 6 of the table 4 from an image 2 generated by optically reading a document including the table 4 formed on a recording medium such as paper.

[0028] 1 is a diagram showing an example of an imaged document accepted by information processing device 10. As long as document image 2 includes table 4, there are no restrictions on the type of imaged document accepted by information processing device 10, and any type of document may be accepted, such as a purchase order, invoice, estimate, contract, blueprint, or assembly drawing. Naturally, the imaged document may include figures other than table 4, such as graphics and photographs.

[0029] Here, an "attribute" refers to an item used to identify the information you want to obtain from Image 2. An attribute value for an attribute refers to the content of each attribute represented by characters contained in Image 2.

[0030] To explain the relationship between attributes and attribute values ​​in Table 4 using purchase order image 2 shown in Figure 1 as an example, "No.", "Product Name," "Quantity," "Unit Price," and "Amount" listed in header row 6 of Table 4 represent items for identifying the information listed in Table 4 and are therefore examples of attributes in Table 4. Furthermore, for cells 3 in which each attribute of Table 4 included in header row 6 is listed, each character listed in cells 3 in each row other than header row 6 arranged in the column direction of Table 4 is an attribute value for each attribute of Table 4 included in header row 6. For example, for cell 3A representing the attribute "Amount," "100,000," "2,300," "17,200," "10,000," "2,400," "65,000," and "28,000" arranged in the column direction of Table 4 represent specific values ​​for the attribute "Amount," and are therefore examples of attribute values ​​for the attribute "Amount."

[0031] "Cell 3" refers to each area separated by the borders of Table 4. However, Cell 3 does not necessarily have to be separated by the borders of Table 4, and even if there are no borders of Table 4, the position of each character will be recognized as Cell 3 of Table 4 as long as the characters are arranged in such a way that the correspondence between rows and columns can be uniquely identified.

[0032] The information processing device 10 defines a two-dimensional coordinate system for image 2 with a specific position within image 2 as the origin P1, and represents positions within image 2 by coordinate values ​​in the two-dimensional coordinate system. In the example of image 2 of the order table shown in FIG. 1, the upper left vertex of image 2 is set as the origin P1, the X axis is set along the horizontal direction of image 2, and the Y axis is set along the vertical direction of image 2. In this embodiment, the row direction of table 4 represents the X axis direction of the two-dimensional coordinate system, and the column direction of table 4 represents the Y axis direction of the two-dimensional coordinate system.

[0033] The information processing device 10 also identifies the position of each column of Table 4 in the image 2 according to a two-dimensional coordinate system. The column position is represented, for example, as the X coordinate value of a frame line separating each column. In the example of Table 4 shown in FIG. 1, the column position is represented by the X coordinate value of the top left vertex P2 of each cell 3 in the first row of Table 4. Note that in FIG. 1 and in FIGS. 2 and 11, which will be described later, the column position is shown for only the leftmost column as a representative example.

[0034] It should be noted that Table 4 shown in FIG. 1 has a page break in the middle of it, for example, because the entire table cannot fit on one page or for layout reasons.

[0035] Figure 2 is a diagram showing an example of Table 4 following Table 4 shown in Figure 1. Three more Tables 4 shown in Figures 2(A) to 2(C) follow Table 4 shown in Figure 1, and the four Tables 4 shown in Figures 1 and 2(A) to 2(C) constitute Table 4 before division. In other words, each of Tables 4 shown in Figure 1 and 2(A) to 2(C) is originally one Table 4 divided into four, so each of them is called a divided Table 4.

[0036] A user does not necessarily create a complete table 4 so that it fits on one page. Therefore, the information processing device 10 may receive an image 2 including a divided table 4 that spans multiple pages (four pages in the example of table 4 shown in FIG. 1 and FIGS. 2(A) to 2(C)).

[0037] Since the header row 6 is placed in the first row of the table 4, the header row 6 will not be included in each of the tables 4 shown in Figures 2(A) to 2(C) that follow the table 4 shown in Figure 1 unless the user intentionally inserts the header row 6 in the first row of each subsequent table 4 after division.

[0038] Hereinafter, for convenience of explanation, a Table 4 that includes a header row 6, such as Table 4 shown in FIG. 1, will be referred to as "Header Table 4A" or "Table 4A," and a Table 4 that follows header Table 4A and does not include a header row 6, such as those shown in FIGS. 2(A) to 2(C), will be referred to as "Subsequent Table 4B" or "Table 4B." Furthermore, when there is no particular need to distinguish between header Table 4A and subsequent Table 4B, they will simply be referred to as "Table 4." Header Table 4A is an example of a "first Table" according to this embodiment, and subsequent Table 4B is an example of a "second Table" according to this embodiment.

[0039] On the other hand, image 2 may contain information representing the total for at least one addable attribute, such as amount and quantity, contained in header row 6. In the example of table 4 shown in Figures 1 and 2, the total amount is shown in area 8 of Figure 1 under the heading "Total Amount," and also in cell 3B of table 4B in Figure 2(C) under the heading "Total Amount."

[0040] 3 is a block diagram showing an example of the functional configuration of an information processing device 10 according to this embodiment. The information processing device 10 includes functional units, such as an image receiving unit 11, a user interface (UI) unit 12, an image processing unit 13, a control unit 14, and an output unit 15, storage areas for an OCR (Optical Character Recognition) result DB (Database) 16 and an extraction result DB 17, and association rules 18.

[0041] The image receiving unit 11 receives an image 2 containing a table 4 from an optical device, such as a scanner, that optically reads the contents of a document to generate an image 2 of the document, and transfers the received image 2 to the image processing unit 13. The images 2 received by the image receiving unit 11 include images 2 in which the table 4 fits on a single page, as well as images 2 containing a table 4 split across multiple pages. The order in which the images 2 are received by the image receiving unit 11 is not necessarily arranged in page order; pages may be out of order, pages may be missing, or pages of documents unrelated to the document for which attribute values ​​are being acquired may be mixed in. For example, the order of the images 2 containing the table 4 shown in Figures 1 and 2 is originally Figure 1, Figure 2(A), Figure 2(B), and Figure 2(C) from the beginning. However, if an order form not arranged in page order is scanned by a scanner, the images 2 may be received in the order of Figure 1, Figure 2(B), Figure 2(A), and Figure 2(C) from the beginning.

[0042] The UI unit 12 accepts instructions from a user who is using the information processing device 10 to obtain attribute values ​​of attributes in Table 4 in image 2, such as an instruction to start accepting image 2 by the image accepting unit 11, and notifies the user of various information such as the operation and status of the information processing device 10.

[0043] The image processing unit 13 extracts character information from the image 2 accepted by the image accepting unit 11, and also extracts, from the extracted character information, attribute values ​​for predetermined attributes in Table 4. For this purpose, the image processing unit 13 includes an OCR processing unit 13A and an extraction unit 13B.

[0044] The OCR processing unit 13A performs known image recognition on the received image 2 and converts into character codes portions of the image 2 that correspond to characters. That is, the OCR processing unit 13A treats portions of the image 2 that correspond to characters as character information, allowing copying and searching of characters.

[0045] Furthermore, the OCR processing unit 13A recognizes the borders of Table 4 and identifies the structure of Table 4. If Table 4 does not have a border, the OCR processing unit 13A may identify the structure of Table 4 by using, for example, the fact that the characters in Table 4 have a feature regarding the arrangement of characters, that is, the characters are arranged in a grid pattern along the row and column directions.

[0046] The structure of Table 4 refers to the external characteristics of Table 4, such as the number of columns, column width, row width, type of borders in Table 4, column position, background color of Cell 3, and font type and size of each character written in Cell 3. The structure of Table 4 can be identified using a known identification method such as that shown in Patent Document 1. Information representing the structure of Table 4 is referred to as "structural information of Table 4."

[0047] Hereinafter, the character information and the structural information of Table 4 obtained from Image 2 by OCR processing unit 13A will be referred to as the “OCR result.” OCR processing unit 13A stores the OCR result in OCR result DB 16.

[0048] The extraction unit 13B refers to the OCR results stored in the OCR result DB 16 to extract all Tables 4 included in the image 2 received by the image receiving unit 11, and classifies each of the extracted Tables 4 into Tables 4 that have a common structural feature. As an example, the OCR processing unit 13A refers to the column width WD (see FIG. 1) in each of the extracted Tables 4, and classifies Tables 4 that have the same column width WD as Tables 4 that have a common structural feature. Naturally, the extraction unit 13B may combine the column width WD with structural information about the Table 4 other than the column width WD to classify each of the extracted Tables 4 into Tables 4 that have a common structural feature.

[0049] Tables 4 split from the same Table 4 tend to have the same structure, so split Tables 4 that share common structural features are more likely to be Tables 4 that make up the same Table 4 than other split Tables 4. In this way, a set of split Tables 4 that are thought to make up the same Table 4 is called a "table group."

[0050] Naturally, in the case of an unsplit table 4, the table group includes one unsplit table 4. In the case of a split table 4, because the table 4 before splitting only contains a header row 6 at the beginning of the table 4, the table group includes one header table 4A and at least one subsequent table 4B.

[0051] The extraction unit 13B stores each of the classified table groups in the extraction result DB 17.

[0052] When the image processing unit 13 classifies each table 4 included in the image 2 into a table group, the control unit 14 performs control to associate each table 4 included in the table group in accordance with a predetermined association rule 18.

[0053] The association rules 18 are rules that define how to associate each of the tables 4 included in the table group.

[0054] The association rules 18 include at least one of association rules 18 that associate each table 4 in the table group so that they have continuity (hereinafter referred to as ``continuity rules''), and association rules 18 that associate each table 4 in the table group so that they have consistency (hereinafter referred to as ``consistency rules'').

[0055] "Continuity" with respect to Table 4 means that the order of each divided Table 4 is continuous. Also, "consistency" with respect to Table 4 means that when divided Table 4 is combined, the contents of the pre-division Table 4 are included without any omission or addition.

[0056] When a continuity rule is defined as the association rule 18, the control unit 14 associates each of the tables 4 included in the table group so that each table 4 is arranged in the same order as the table 4 before division. Furthermore, when a consistency rule is defined as the association rule 18, the control unit 14 associates each of the tables 4 included in the table group after determining in advance whether the table 4 obtained by associating the tables 4 included in the table group contains the contents of the table 4 before division without excess or deficiency. Note that "associating each table 4" includes, as an example, a mode in which the tables 4 after division are combined so as to satisfy the association rule 18.

[0057] In this way, when subsequent table 4B is associated with table 4A including header row 6, the attribute values ​​for each attribute of table 4 included in header row 6 are arranged in the column direction of table 4. Therefore, a correspondence relationship can be obtained, such as which attribute of table 4 the content of each cell 3 included in subsequent table 4B that does not include header row 6 represents and which attribute of table 4 corresponds to the attribute value. In other words, by associating each table 4 included in the table group in accordance with association rule 18, control unit 14 virtually associates header row 6 with each subsequent table 4B that does not include header row 6, while maintaining at least one of continuity and consistency among the associated tables 4.

[0058] The control unit 14 acquires attribute values ​​for each attribute of the table 4 included in the header row 6 from the table 4 that is associated so as to have at least one of continuity and consistency.

[0059] To obtain the attribute values ​​for each attribute in Table 4, for example, "in-table KV extraction" is used. In-table KV extraction, attribute values ​​for the attributes of Table 4 included in header row 6 are extracted from Table 4 on a page-by-page basis. In in-table KV extraction, "KV" stands for "Key-Value," where "Key" is a string representing an attribute in Image 2, and "Value" is a string representing the attribute value for the attribute.

[0060] For example, if KV extraction is performed on Image 2 of the order table shown in Figure 1, for the attribute "amount," the values ​​"100,000," "2,300," "17,200," "10,000," "2,400," "65,000," and "28,000," which are arranged in the columns of Table 4A, are extracted from Table 4A as the attribute values ​​for amount.

[0061] For such table KV extraction, a header row 6 is required to identify which cell 3 in the table contains the attribute value of which attribute.

[0062] Therefore, it would normally be impossible to extract attribute values ​​for attributes of Table 4 by table KV extraction from subsequent Table 4B, which does not include header row 6. However, as described above, in the case of information processing device 10 according to this embodiment, header row 6 can be virtually associated with subsequent Table 4B. Therefore, information processing device 10 according to this embodiment becomes able to acquire attribute values ​​for attributes of Table 4, which are included in header row 6, from subsequent Table 4B as well, in a state in which the associated Table 4 has at least one of continuity and consistency.

[0063] The control unit 14 stores the attribute values ​​acquired from each of the tables 4 associated with each other so as to have at least one of continuity and consistency in the extraction result DB in association with each attribute of the table 4 included in the header row 6.

[0064] The specific processing contents of the associations in Table 4 according to the association rule 18 in the information processing device 10 will be explained in detail later.

[0065] In accordance with an instruction from the control unit 14, the output unit 15 acquires from the extraction result DB 17 the attribute values ​​for each attribute in Table 4 before division that have been extracted from the image 2 by the control unit 14, and outputs them.

[0066] Outputting the attribute value for each attribute in Table 4 means making the attribute value for each attribute in Table 4 available for confirmation by a user. Therefore, the following are all examples of outputting the attribute value for each attribute in Table 4: transmitting the attribute value for each attribute in Table 4 to an external device via a communication line, displaying the attribute value for each attribute in Table 4 on a display, printing the attribute value for each attribute in Table 4 on a recording medium such as paper by an image forming device, and storing the attribute value for each attribute in Table 4 in a storage device that a user has permission to read.

[0067] 3 is configured using, for example, a computer 20. FIG. 4 is a diagram showing an example of the configuration of the main parts of an electrical system in the information processing device 10 configured using the computer 20.

[0068] 3, a ROM (Read Only Memory) 22 that stores an information processing program executed by the CPU 21, a RAM (Random Access Memory) 23 that is used as a temporary work area for the CPU 21, a nonvolatile memory 24, and an input / output interface (I / O) 25. The CPU 21, the ROM 22, the RAM 23, the nonvolatile memory 24, and the I / O 25 are connected to each other via a bus 26.

[0069] The nonvolatile memory 24 is an example of a storage device that maintains stored information even if the power supplied to the nonvolatile memory 24 is cut off, and is, for example, a semiconductor memory or a hard disk. The nonvolatile memory 24 does not necessarily have to be built into the computer 20, and may be a storage device that is detachable from the computer 20, such as a memory card. The OCR result DB 16 and the extraction result DB 17 are constructed in the nonvolatile memory 24, for example.

[0070] To the I / O 25, for example, a communication unit 27, an input unit 28, and a display unit 29 are connected.

[0071] The communication unit 27 is connected to a communication line and is provided with a communication protocol for communicating with external devices such as storage devices and computers connected to the same connection line.

[0072] The input unit 28 is a device that receives instructions from a user and notifies the CPU 21, and may be, for example, a button, a touch panel, a keyboard, a mouse, etc. The information processing device 10 executes a function instructed by the user via the input unit 28.

[0073] The display unit 29 is a device that displays information processed by the CPU 21 as an image, and may be, for example, a liquid crystal display, an organic EL (Electro Luminescence) display, or a projector that projects an image onto a screen.

[0074] The input unit 28 and the display unit 29 cooperate with the UI unit 12 shown in FIG. 3 to accept various instructions from the user and notify the user of various information relating to the operation and state of the information processing device 10.

[0075] Note that the units connected to the I / O 25 are not limited to the units exemplified in Fig. 4. For example, a scanner unit that optically reads the contents of a document placed on the platen glass and converts the document contents into an image may be connected to the I / O 25. In this case, the CPU 21 receives the document image 2 from the scanner unit via the I / O 25.

[0076] If the scanner unit is not connected to the I / O 25, the information processing device 10 may receive the image 2 from an external device, for example, via the communication unit 27. The information processing device 10 may also receive the image 2 from a storage device that is detachably attached to the computer 20, such as a memory card.

[0077] First Embodiment Next, an example of processing in the information processing device 10 will be described in detail.

[0078] 5 is a flowchart showing an example of the flow of extraction processing executed by the CPU 21 when a document image 2 is received from an external device, for example, via the communication unit 27. It is assumed that a continuity rule is defined in the association rule 18 of the information processing device 10.

[0079] An information processing program that defines the extraction process according to the first embodiment is stored in advance in, for example, the ROM 22 of the information processing device 10. The CPU 21 of the information processing device 10 reads the information processing program stored in the ROM 22 and executes the extraction process.

[0080] In step S10, the CPU 21 performs OCR processing on each of the received images 2, generates OCR results including character information within each image 2 and structural information of Table 4, and stores the results in the OCR result DB 16 constructed in the non-volatile memory 24.

[0081] In step S20, CPU 21 performs table KV extraction, extracts attribute values ​​for each attribute of Table 4A included in header row 6 from Table 4A including header row 6, and stores the extracted attribute values ​​in RAM 23 for each attribute of Table 4A.

[0082] In step S30, the CPU 21 acquires the structure information of each of the Tables 4 included in the image 2 from the OCR result DB 16.

[0083] In step S40, CPU 21 uses the structural information of Table 4 acquired in step S30 to classify each of Table 4 included in image 2 into a group of tables that have common characteristics regarding the structure of Table 4. For ease of explanation, it is assumed here that each of Table 4 is classified into one group of tables.

[0084] In step S50, CPU 21 acquires each attribute included in header row 6 of Table 4A. In the case of Table 4A, which is the order table shown in Fig. 1, "No.", "Product name.", "Quantity.", "Unit price.", and "Amount." are acquired as attributes of Table 4A included in header row 6.

[0085] In step S60, CPU 21 determines whether the attributes of Table 4A acquired in step S50 include an order attribute. An order attribute is an attribute that indicates the order of Table 4. In the case of Table 4A of the order table shown in FIG. 1, "No.", an example of an attribute included in header row 6, is an example of an order attribute because it is information that manages the order of each row in Table 4 using integers arranged in ascending order from 1. In this way, the order of attribute values ​​for the order attribute is determined according to a predetermined rule, such as ascending order or descending order. Therefore, CPU 21 can determine the continuity of Table 4 after division from the order of attribute values ​​for the order attribute.

[0086] If it is determined in the determination process of step S60 that the header row 6 includes an order attribute, the process proceeds to step S70.

[0087] In step S70, CPU 21 sets heading table 4A as the comparison table to be compared with respect to the continuity of table 4. That is, by setting heading table 4A as the comparison table, CPU 21 makes preparations to identify succeeding table 4B that follows heading table 4A from among the divided tables 4 included in the table group.

[0088] In step S80, CPU 21 selects one of the succeeding tables 4B included in the table group. In this case, CPU 21 may randomly select one of the succeeding tables 4B included in the table group, but it is preferable to select the succeeding table 4B included in each image 2 in the order in which the images 2 were received.

[0089] Users are more likely to image a document in page order than to image the document in the wrong page order. Therefore, by selecting the successor table 4B according to the order in which the images 2 were received, the successor table 4B closest to the comparison table at the time the images 2 were received is selected first, which may shorten the time required to associate each of the divided tables 4 so that they have continuity, compared to selecting successor tables 4B randomly. In this embodiment, the successor table 4B selected in step S80 is referred to as the selected successor table 4B.

[0090] In step S90, CPU 21 compares the attribute value of the order attribute in the last row of Table 4 set in the comparison table with the attribute value of the order attribute in the first row of selected successor Table 4B, and determines whether the attribute values ​​of the order attribute are arranged in accordance with a predetermined rule for the order attribute from the comparison table to selected successor Table 4. In other words, CPU 21 determines whether the attribute values ​​of the order attribute are consecutive in the last row of the comparison table and the first row of selected successor Table 4B.

[0091] If the header table 4A shown in FIG. 1 is set as the comparison table and the successor table 4B shown in FIG. 2A is the selected successor table 4B, the attribute value of "No" in the last row of the comparison table is "7," and the attribute value in the same column as "No" in the first row of the selected successor table 4B is "8." If there is a rule that the attribute value of "No" is set to an integer that increases by one from top to bottom of table 4, CPU 21 determines that the comparison table and selected successor table 4B are continuous because the attribute values ​​of the comparison targets are "7" and "8." Note that the comparison table is an example of "one table" according to this embodiment, and the selected successor table 4B is an example of "the other table" according to this embodiment.

[0092] If the attribute values ​​of the order attribute are consecutive in the last row of the comparison table and the first row of the selection successor table 4B, the process proceeds to step S100.

[0093] Since it is determined that the comparison table and the selection successor table 4B are consecutive, in step S100, the CPU 21 associates the comparison table and the selection successor table 4B according to the order of the attribute values ​​of the order attribute so that the selection successor table 4B is consecutive after the comparison table.

[0094] In step S130, the CPU 21 sets the selected successor table 4B as a new comparison table.

[0095] In step S140, CPU 21 determines whether or not there is a successor table 4B for which association has not yet been completed among the successor tables 4B included in the table group. If there is a successor table 4B for which association has not yet been completed, the process proceeds to step S80, and CPU 21 selects a successor table 4B for which association has not yet been completed from among the successor tables 4B included in the table group as a new selected successor table 4B.

[0096] On the other hand, if it is determined in the determination process of step S90 that the comparison table and the selection successor table 4B are not consecutive, the process proceeds to step S110.

[0097] In step S110, CPU 21 determines whether the group of tables includes a successor table 4B for which continuity has not yet been determined with respect to the current comparison table and for which no association has yet been established. If the group of tables includes an unrelated successor table 4B for which continuity has not yet been determined with respect to the current comparison table, the process proceeds to step S80, and CPU 21 selects, from the successor tables 4B included in the group of tables, a successor table 4B for which continuity has not yet been determined with respect to the current comparison table and for which no association has yet been established, as the selected successor table 4B.

[0098] On the other hand, if it is determined in the determination process of step S110 that the continuity determination for the current comparison table has been performed for all subsequent tables 4B in the group of tables that have not yet been associated, the process proceeds to step S120.

[0099] In this case, it is possible that image 2 has been received with at least one of the succeeding tables 4B missing, making it difficult to associate tables 4 according to the attribute value of the order attribute. Therefore, in step S120, CPU 21 associates the selected succeeding table 4B after the comparison table in the order in which image 2 was received. That is, CPU 21 sets succeeding table 4B included in image 2 on the page next to image 2 that contained the comparison table at the time of receiving image 2 as the table 4 following the comparison table.

[0100] Note that when there are multiple table groups, CPU 21 may detect subsequent tables 4B that are contiguous to the comparison table from other table groups, instead of associating subsequent tables 4B after the comparison table in accordance with the order of receipt of images 2. For example, even if tables 4 are split from the same table 4, errors in reading the structural information of table 4 may result in the split tables 4 being classified across multiple table groups. In such a case, if a subsequent table 4B that is contiguous to the comparison table is included in a table group other than the table group for which the subsequent table 4B was selected in step S80, CPU 21 may obtain subsequent tables 4 that are contiguous to the comparison table from the other table groups and associate them after the comparison table.

[0101] After associating the succeeding table 4B with the comparison table in accordance with the order of reception of the images 2, the CPU 21 proceeds to step S130 and sets the succeeding table 4B associated with the comparison table as a new comparison table.

[0102] That is, by repeatedly executing the processes of steps S80 to S140 until all tables 4 included in the table group have been associated, each of the divided tables 4 included in the table group can be associated. Of the associated tables 4, tables 4 that are associated so that the order of the divided tables 4 is continuous are called a "series of tables 4." A series of tables 4 consists of a header table 4A and at least one or more subsequent tables 4B.

[0103] If it is determined in the determination process of step S140 that all tables 4 in the table group have been associated, the process proceeds to step S160.

[0104] By associating Tables 4 included in the table group, the correspondence between each attribute of Table 4 included in the header row 6 of Table 4A and each attribute value in the subsequent Table 4B becomes clear. Therefore, in step S160, CPU 21 adds each attribute value in each subsequent Table 4B to the attribute value for each attribute of Table 4 included in the header row 6 extracted from Table 4A by the in-table KV extraction in step S20, in accordance with the sorting order of the subsequent Table 4B, for each attribute in the header row 6. As a result, the attribute value of each attribute of Table 4 included in the header row 6 is acquired from each associated Table 4.

[0105] If each table 4 is associated with another table 4 so as to have continuity, the attribute values ​​for each attribute of the table 4 included in the header row 6 are obtained according to the order of the tables 4 before division.

[0106] In step S170, the CPU 21 outputs the attribute values ​​for each attribute of Table 4 included in the header row 6, which were acquired in step S160, and ends the extraction process shown in FIG.

[0107] If each table 4 is associated with another table 4 so as to have continuity, the CPU 21 will output the attribute values ​​for each attribute of the table 4 included in the header row 6 in the order of the table 4 before division.

[0108] If it is determined in the determination process of step S60 that the header row 6 does not include an order attribute, the CPU 21 cannot associate each of the tables 4 included in the table group using the attribute value of the order attribute. Therefore, the process proceeds to step S150, where the CPU 21 associates each of the tables 4 in the table group included in each image 2 in accordance with the order in which the images 2 were received, and then proceeds to step S160.

[0109] In this case, the associated Table 4 may not necessarily have continuity, so the order of the attribute values ​​for each attribute of Table 4 obtained in step S160 may not be the same as the order of the pre-split Table 4. However, the attribute values ​​for each attribute of Table 4 included in header row 6 are obtained from at least subsequent Table 4B that does not include header row 6.

[0110] When CPU 21 outputs the attribute values ​​for each attribute of Table 4 included in header row 6 in step S170, it is preferable that CPU 21 also outputs information notifying whether there is continuity. Specifically, if CPU 21 has acquired attribute values ​​from each of Tables 4 in which association between Tables 4 has been performed in accordance with the order of reception of images 2 even once in the processing of step S120 or S150, CPU 21 outputs that the continuity of the attribute values ​​is not guaranteed. Furthermore, if CPU 21 has acquired attribute values ​​from each of Tables 4 in which association between Tables 4 has not been performed in accordance with the order of reception of images 2 even once in the processing of step S120 or S150, CPU 21 outputs that the continuity of the attribute values ​​is guaranteed.

[0111] If the table 4 included in the received image 2 is classified into a plurality of table groups in step S40, the processes of steps S50 to S170 may be performed for each table group.

[0112] In this way, even if one table 4 is split across pages and the subsequent table 4B does not include a header row 6, the information processing device 10 can associate the divided tables 4 so that the attribute values ​​of the subsequent table 4B have continuity based on the arrangement of the attribute values ​​for the order attribute included in the header row 6 of table 4A. Therefore, the information processing device 10 can obtain the attribute values ​​for each attribute included in the header row 6 from each divided table 4 in the arrangement of the table 4 before splitting.

[0113] Second Embodiment The first embodiment has described a processing example in which a continuity rule is defined in the association rules 18 of the information processing device 10. The second embodiment will describe a processing example in which a consistency rule is defined in the association rules 18 of the information processing device 10.

[0114] The information processing device 10 of the second embodiment checks whether the table 4 generated by associating each divided table 4 included in the table group contains any parts that are missing from the table 4 before division or any extra parts that were not present in the table 4 before division, and then obtains attribute values ​​for each attribute contained in the header row 6 from the divided table 4.

[0115] FIG. 6 is a flowchart showing an example of the flow of extraction processing executed by the CPU 21 when a document image 2 is received from an external device via the communication unit 27, for example.

[0116] An information processing program that defines the extraction process according to the second embodiment is stored in advance in, for example, the ROM 22 of the information processing device 10. The CPU 21 of the information processing device 10 reads the information processing program stored in the ROM 22 and executes the extraction process.

[0117] The processes of steps S200 to S230 are the same as the processes of steps S10 to S40 of the extraction process according to the first embodiment shown in Fig. 5, and therefore will not be described again. For ease of explanation, it is assumed that in step S230, each Table 4 is classified into one table group.

[0118] In step S240, CPU 21 acquires each attribute of Table 4 included in heading row 6 of Table 4A. CPU 21 also detects a location (hereinafter referred to as a "total column") in any of Table 4 included in the table group that lists the total of the attribute values ​​for at least one attribute included in heading row 6 of Table 4A. CPU 21 acquires the attribute value listed in the detected total column, i.e., the total of the attribute values ​​for any of the attributes of Table 4 included in heading row 6 of Table 4A.

[0119] For example, in the case of Table 4 shown in Figures 1 and 2, the CPU 21 detects at least one of the "Total Amount" shown in Figure 1 or the "Total Amount" shown in Figure 2(C) as the total column from Table 4, and obtains "566,000" listed in the same row as the total column as the total column. Hereinafter, the total in the total column obtained from Image 2 including any Table 4 included in the table group will be referred to as the "obtained total." In other words, the obtained total is an example of a total that is previously recorded in Image 2 including Table 4.

[0120] To detect the total column from Table 4, it is sufficient to detect characters indicating that some kind of total is written, such as "total," "total amount," and "cumulative total," from Table 4. Also, by predefining information about the direction in which the attribute value of the total column is located as viewed from the total column, CPU 21 can obtain the total of the total column from Table 4 in accordance with the predefinition.

[0121] If the item name of the detected total column contains characters that indicate a relationship with one of the attributes included in header row 6 of Table 4A, it is determined from the relationship that the total column represents the sum of the attribute values ​​for which attribute included in header row 6 of Table 4A. For example, since "Total Amount," an example of a total column shown in Figures 1 and 2(C), contains the character "amount," it is determined that "Total Amount" is a total column that represents the sum of each attribute value for "Amount" in header row 6 of Table 4A shown in Figure 1. Hereinafter, an attribute included in header row 6 of Table 4A and whose sum of attribute values ​​is represented by the total column will be referred to as a "consistent attribute."

[0122] The CPU 21 may determine, from the position where the total in the total column is written, the total of the attribute values ​​for which attribute included in the header row 6 of Table 4A the total column represents. For example, if the position of the column where "566,000" representing the total of "total amount" is written as in Table 4B in Fig. 2(C) is the same as the position of the "amount" column in the header row 6 of Table 4A shown in Fig. 1, the CPU 21 determines that "566,000" is the total column representing the sum of each attribute value for "amount."

[0123] In step S250, CPU 21 selects header table 4A and calculates the sum of the attribute values ​​in header table 4A for the matching attributes. The sum calculated by CPU 21 from each attribute value listed in Table 4 is referred to as the "calculated sum."

[0124] In step S260, CPU 21 determines whether the calculated total for the matching attributes calculated in step S250 matches the acquired total for the matching attributes acquired in step S240. If the calculated total and the acquired total are different values, the process proceeds to step S270.

[0125] If the calculated total and the acquired total are different, it means that at least one or more subsequent tables 4B follow the heading table 4A. Therefore, in step S270, CPU 21 stores the calculated total of heading table 4A calculated in step S250 in RAM 23 as the cumulative calculated total.

[0126] In step S280, CPU 21 selects one successor table 4B that has not yet been selected from the successor tables 4B included in the table group. CPU 21 may randomly select one successor table 4B from the successor tables 4B included in the table group, but it is preferable to select the successor table 4B included in each image 2 in the order in which the images 2 were received. In this embodiment, the successor table 4B selected in step S280 is referred to as the selected successor table 4B.

[0127] Then, CPU 21 calculates the sum of the attribute values ​​in selection successor table 4B for the matching attribute. Note that CPU 21 recognizes the attribute value in selection successor table 4B that is at the same position as the column position of the matching attribute in header row 6 as the attribute value in selection successor table 4B for the matching attribute.

[0128] In step S290, CPU 21 adds the calculated total of selected subsequent table 4B calculated in step S280 to the cumulative calculated total stored in RAM 23, and stores the cumulative calculated total after adding the calculated total of selected subsequent table 4B in RAM 23 as a new cumulative calculated total.

[0129] In step S300, CPU 21 determines whether the cumulative calculated total updated in step S290 and the acquired total for the matching attribute acquired in step S240 are the same value. If the cumulative calculated total and the acquired total are the same value, the process proceeds to step S310.

[0130] Since the cumulative calculated total and the acquired total are the same value, by linking the Tables 4 that were the subject of the calculation of each calculated total that makes up the cumulative calculated total, it is possible to construct Table 4 before division without any omissions or excesses.

[0131] That is, since each of the tables 4 that are the subject of calculation of each of the calculated totals that make up the cumulative calculated total having the same value as the acquired total are consistent, in step S310, CPU 21 sets the association information of table 4 to "normal." Note that the "association information" of table 4 is information that indicates the state in which each table 4 is associated with respect to consistency. "Normal" in the association information indicates that each associated table 4 is in a consistent state.

[0132] In step S320, CPU 21 associates the header table 4A selected in step S250 with the subsequent tables 4B that were the subject of calculation of each of the calculation totals that make up the cumulative calculation total. The subsequent tables 4B are preferably associated with header table 4A in the order in which the subsequent tables 4B were selected in step S280. For example, if the subsequent tables 4B are selected in the order of Figures 2(A), 2(B), and 2(C) for header table 4A shown in Figure 1, CPU 21 associates each table 4 with Figures 1, 2(A), 2(B), and 2(C) from the top. CPU 21 associates association information with each associated table 4 and stores the information in extraction result DB 17.

[0133] From then on, in the processes of steps S370 and S380, the same processes as those of steps S160 and S170 of the extraction process of the first embodiment shown in Figure 5 are performed, and the attribute values ​​for each attribute of Table 4 included in header row 6 extracted from each associated Table 4 are output, and the extraction process shown in Figure 6 is terminated.

[0134] On the other hand, if it is determined in the determination process of step S300 that the cumulative calculated total and the acquired total are different values, the process proceeds to step S340.

[0135] In step S340, CPU 21 determines whether all subsequent tables 4B included in the table group have been selected. If there are any unselected subsequent tables 4B, adding the sum of the attribute values ​​in the unselected subsequent tables 4B to the cumulative calculated total may result in the cumulative calculated total and the acquired total being the same value.

[0136] Therefore, the process proceeds to step S280, where one of the unselected successor tables 4B included in the table group is selected as a new selected successor table 4B. Then, CPU 21 executes the processes from step S290 onwards for the new selected successor table 4B.

[0137] On the other hand, if it is determined in the determination process of step S340 that all succeeding tables 4B have been selected from the table group, the process proceeds to step S350.

[0138] If the cumulative calculated total and the obtained total are not the same value even though all subsequent tables 4B included in the table group have been selected, this means that there is a surplus or shortage of tables 4 included in the table group compared to the number of tables 4 split from the pre-split table 4. This situation occurs when table group extraction is not performed properly, or when the accepted image 2 contains unnecessary tables 4, or conversely, does not contain necessary tables 4.

[0139] That is, since each of the Tables 4 included in the group of tables does not have consistency, in step S350, the CPU 21 determines whether the association state of the Table 4 is "insufficient" or "excessive."

[0140] The association status "insufficient" indicates a state in which pre-division Table 4 cannot be constructed because post-division Table 4 is missing. Specifically, when the cumulative calculation total is less than the acquired total, CPU 21 determines the association status of Table 4 to be "insufficient."

[0141] The association status of "surplus" indicates a state in which the specific pre-split Table 4 cannot be constructed because the group of tables includes a successor Table 4B of another Table 4 that is of a different type from the specific pre-split Table 4. Specifically, the CPU 21 determines the association status of Table 4 to be "surplus" when the cumulative calculated total exceeds the acquired total.

[0142] In step S360, the CPU 21 sets the determination result determined in step S350 in the association information in Table 4, and then proceeds to step S320.

[0143] If it is determined in the determination process of step S260 that the calculated total and the acquired total are the same value, the process proceeds to step S330.

[0144] In this case, Table 4 before division is heading Table 4A itself. In other words, there is no succeeding Table 4B following heading Table 4A. Therefore, heading Table 4A in this case is consistent on its own, so in step S330, CPU 21 sets the association information of Table 4 to "normal" and proceeds to step S370.

[0145] When CPU 21 outputs the attribute values ​​for each attribute in Table 4 included in header row 6 in step S380, it may also output information notifying whether there is consistency. Specifically, when attribute values ​​are acquired from each Table 4 associated with association information set to "normal," CPU 21 outputs that consistency of the attribute values ​​is guaranteed. Furthermore, when attribute values ​​are acquired from each Table 4 associated with association information set to "shortage" or "surplus," CPU 21 outputs that consistency of the attribute values ​​is not guaranteed.

[0146] According to the information processing device 10 of the second embodiment, even if one table 4 is split across pages and the header row 6 is not included in the subsequent table 4B, the attribute values ​​for each attribute included in the header row 6 can be obtained exactly from each of the split tables 4 whose association information is set to "normal" by comparing them with the attribute values ​​recorded in the table 4 before splitting.

[0147] Therefore, even if the heading row 6 of Table 4 does not contain an order attribute, if the heading row 6 of Table 4 contains a consistency attribute, by specifying a consistency rule in the association rule 18 of the information processing device 10, all of the attribute values ​​listed in Table 4 before division can be obtained without any excess or deficiency for each attribute contained in the heading row 6 from each of the divided Tables 4 whose association information is set to "normal".

[0148] Third Embodiment The information processing device 10 according to the first embodiment acquires attribute values ​​for each attribute included in the header row 6 from each divided Table 4 in the order of the attribute values ​​in the pre-divided Table 4. In this case, it is clear that the acquired attribute values ​​have continuity, but it is unclear whether all of the attribute values ​​listed in the pre-divided Table 4 have been acquired for each attribute included in the header row 6 without any excess or deficiency.

[0149] For example, if Table 4 is divided into four pages as shown in Figures 1 and 2, and Tables 4 on pages 1 to 3 are classified into the same table group, and the final page, i.e., Table 4B in Figure 2(C), is classified into a different table group, the attribute values ​​obtained from the table group into which Tables 4 on pages 1 to 3 are classified will have continuity but will not have consistency.

[0150] In the third embodiment, an information processing device 10 that acquires attribute values ​​having continuity and consistency for each attribute included in the header row 6 from each of the divided tables 4 will be described.

[0151] FIG. 7 is a flowchart showing an example of the flow of extraction processing executed by the CPU 21 when a document image 2 is received from an external device via the communication unit 27, for example.

[0152] It is assumed that the association rules 18 of the information processing device 10 include both continuity rules and consistency rules.

[0153] An information processing program that defines the extraction process according to the third embodiment is stored in advance in, for example, the ROM 22 of the information processing device 10. The CPU 21 of the information processing device 10 reads the information processing program stored in the ROM 22 and executes the extraction process.

[0154] In step S400, the CPU 21 executes a first process. The first process is a process that follows the flow of steps S10 to S150 of the extraction process according to the first embodiment shown in Fig. 5. It is assumed that each of the divided Tables 4 is associated with a series of Tables 4 by the first process.

[0155] In step S410, CPU 21 determines whether or not any of images 2 including a series of tables 4 associated by the first process has a total column. If a total column is present, the process proceeds to step S420, where CPU 21 obtains the total written in the total column.

[0156] As explained in step S240 of the extraction process according to the second embodiment shown in Figure 6, the CPU 21 identifies the matching attribute corresponding to the total column using at least one of the information on the item name in the total column and the position where the total in the total column is written.

[0157] In step S430, CPU 21 adds up the calculated totals of attribute values ​​for the matching attributes in each divided Table 4, calculated from each Table 4 that makes up the series of Table 4, for each divided Table 4, to obtain a cumulative calculated total for the attribute values ​​of the matching attributes.

[0158] In step S440, CPU 21 determines whether the acquired total acquired in step S420 and the cumulative calculated total calculated in step S430 are the same value. If the acquired total and the cumulative calculated total are the same value, the process proceeds to step S450.

[0159] In this case, the series of Tables 4 associated in the first process will have consistency. Therefore, in steps S450 and S460, CPU 21 performs the same processes as steps S160 and S170 of the extraction process according to the first embodiment shown in Fig. 5, extracts attribute values ​​for each attribute of Table 4 included in header row 6 from the series of Tables 4 that have continuity and consistency, outputs the extraction results, and ends the extraction process shown in Fig. 7. In this case, CPU 21 may output a notification to the user that the attribute values ​​acquired from the series of Tables 4 have continuity and consistency.

[0160] On the other hand, if it is determined in the determination process of step S410 that none of the images 2 containing each of the tables 4 constituting the series of tables 4 have a total column, it is unclear whether the series of tables 4 are consistent. Therefore, the CPU 21 ends the extraction process shown in Fig. 7 without obtaining attribute values ​​for each attribute of the table 4 included in the header row 6 from the series of tables 4. In this case, it is preferable that the CPU 21 output a notification to the user that the obtaining of attribute values ​​has been stopped because it is unclear whether the series of tables 4 are consistent.

[0161] Furthermore, if it is determined in the determination process of step S440 that the cumulative acquired total and the calculated total are different, the series of Tables 4 have continuity but not consistency. Therefore, CPU 21 ends the extraction process shown in Fig. 7 without acquiring attribute values ​​for each attribute of Table 4 included in header row 6 from the series of Tables 4. In this case, it is preferable that CPU 21 output a notification to the user that acquisition of attribute values ​​has been stopped because the series of Tables 4 are not consistent.

[0162] As described above, the information processing device 10 according to the third embodiment acquires attribute values ​​for each attribute of Table 4 included in the header row 6 when Table 4 formed by associating each of the divided Tables 4 has both continuity and consistency. That is, the attribute values ​​for each attribute acquired from each of the divided Tables 4 by the extraction process according to the third embodiment have continuity and consistency.

[0163] In addition, when the CPU 21 receives a notification that it is unclear whether a series of tables 4 are consistent, or a notification that a series of tables 4 are not consistent, and receives an instruction to obtain attribute values ​​from the user through the input unit 28, the CPU 21 may obtain attribute values ​​for each attribute of tables 4 included in the header row 6 from the series of tables 4, even if the series of tables 4 are not consistent.

[0164] <Fourth embodiment> In the information processing device 10 according to the third embodiment, the presence or absence of consistency is determined for a series of tables 4 having continuity, and if a table 4 formed by associating each of the divided tables 4 has both continuity and consistency, an attribute value for each attribute of the table 4 included in the header row 6 is acquired. However, there are no restrictions on the order of determining the continuity and consistency for the table 4 formed by associating each of the divided tables 4.

[0165] In the fourth embodiment, an information processing device 10 will be described which determines the consistency of table 4 after division, then determines continuity, and acquires attribute values ​​having continuity and consistency for each attribute included in header row 6.

[0166] FIG. 8 is a flowchart showing an example of the flow of extraction processing executed by the CPU 21 when a document image 2 is received from an external device via the communication unit 27, for example.

[0167] It is assumed that the association rules 18 of the information processing device 10 include both continuity rules and consistency rules.

[0168] An information processing program that defines the extraction process according to the fourth embodiment is stored in advance in, for example, the ROM 22 of the information processing device 10. The CPU 21 of the information processing device 10 reads the information processing program stored in the ROM 22 and executes the extraction process.

[0169] In step S500, the CPU 21 executes a second process, which is a process that follows the flow of steps S200 to S360 of the extraction process according to the second embodiment shown in FIG.

[0170] In step S510, CPU 21 refers to the contents of the association information associated with Table 4 associated in step S500 (hereinafter in this embodiment, simply referred to as "associated Table 4") and determines whether "normal" is set in the association information.

[0171] If "normal" is not set in the association information, it is not guaranteed that the associated Table 4 has consistency, and therefore the extraction process shown in Fig. 8 is terminated. In this case, it is preferable that the CPU 21 output a notice to the user that acquisition of the attribute value has been stopped because it is unclear whether the associated Table 4 has consistency.

[0172] If it is determined in the determination process of step S510 that the association information is set to "normal," the process proceeds to step S520. That is, if Table 4 associated by the second process is consistent, the process proceeds to step S520.

[0173] In step S520, CPU 21 determines whether or not there is an order attribute in header row 6 of associated Table 4. If there is no order attribute in header row 6 of associated Table 4, the continuity of associated Table 4 cannot be ensured, and therefore the extraction process shown in Fig. 8 is terminated. In this case, CPU 21 preferably outputs a notification to the user that the associated Table 4 has consistency but it is unclear whether it has continuity, and therefore acquisition of attribute values ​​has been discontinued.

[0174] If it is determined in the determination process of step S520 that the header row 6 of the associated table 4 has an order attribute, the process proceeds to step S530.

[0175] In step S530, CPU 21 determines whether the associated Tables 4 are consecutive. The determination of the consecutiveness of Table 4 is performed using the determination method described in step S90 of the extraction process according to the first embodiment shown in Fig. 5. If the associated Tables 4 are not consecutive, the process proceeds to step S540.

[0176] The judgment process of step S510 ensures that the associated Tables 4 are at least consistent, so if the attribute values ​​for the order attribute of each associated Table 4 are referenced and the order of each associated Table 4 is rearranged so that the attribute values ​​for the order attribute of each associated Table 4 are continuous, the rearranged associated Table 4 will have continuity.

[0177] Therefore, in step S540, the CPU 21 refers to the attribute value for the order attribute of each associated Table 4, rearranges the associated Tables 4 so that they are consecutive, and then proceeds to step S550.

[0178] On the other hand, if the judgment process of step S530 determines that each of the associated Tables 4 is consecutive, there is no need to rearrange the order of each of the associated Tables 4, so the process proceeds to step S550 without executing the process of step S540.

[0179] In steps S550 and S560, the CPU 21 performs the same processes as steps S160 and S170 of the extraction process according to the first embodiment shown in Fig. 5, acquires attribute values ​​for each attribute of Table 4 included in header row 6 from associated Table 4 that has continuity and consistency, outputs the extraction results, and ends the extraction process shown in Fig. 8. In this case, the CPU 21 may output a notification to the user that the attribute values ​​acquired from associated Table 4 have continuity and consistency.

[0180] As described above, the information processing device 10 according to the fourth embodiment acquires attribute values ​​for each attribute of Table 4 included in the header row 6 when Table 4 constructed by associating each of the divided Tables 4 has both continuity and consistency. That is, the attribute values ​​of each attribute acquired from each of the divided Tables 4 by the extraction process according to the fourth embodiment have continuity and consistency.

[0181] In addition, when the CPU 21 receives a notification that it is unclear whether the associated Table 4 has continuity or that the associated Table 4 does not have consistency and receives an instruction to obtain attribute values ​​from the user through the input unit 28, the CPU 21 may obtain attribute values ​​for each attribute of the associated Table 4 included in the header row 6 from the associated Table 4 even if the associated Table 4 does not have at least one of consistency and continuity.

[0182] In the extraction process of the information processing device 10 according to the fourth embodiment described above, if the association information associated with the associated Table 4 is not set to "normal," it is not guaranteed that the associated Table 4 has consistency, and therefore the extraction process is terminated without acquiring attribute values ​​from the associated Table 4. However, even if the association information of the associated Table 4 is not "normal," there are cases in which the associated Table 4 can be corrected to have consistency.

[0183] As explained in the first embodiment, even if Table 4 is split from the same Table 4, an error in reading the structural information of Table 4 may result in the split Table 4 being classified across multiple table groups.

[0184] Therefore, in the case of an associated table 4 whose association information is set to "missing," a table 4 that was split off from the same table 4 before the split may be included in a group of tables other than the group of tables that contains the associated table 4.

[0185] To deal with such a situation, when the association information associated with the associated Table 4 is "insufficient," the CPU 21 may acquire a divided Table 4 that satisfies the consistency of the associated Table 4 from a table group other than the table group that includes the associated Table 4, and associate the divided Table 4 with the associated Table 4. In this case, the associated Table 4 will have consistency, so the CPU 21 performs the processes from step S520 onwards in FIG. 8.

[0186] On the other hand, in the case of an associated table 4 whose association information is set to "redundant," the group of tables that includes the associated table 4 will contain a table 4 that has been split off from another table 4 of a different type.

[0187] To deal with such a situation, if the association information associated with an associated Table 4 is "redundant," the CPU 21 refers to the attribute value for the order attribute included in the header row 6 of the associated Table 4 and deletes non-consecutive Tables 4 from the associated Tables 4. Then, the CPU 21 checks the consistency of the associated Table 4 after the unnecessary Tables 4 have been deleted by comparing the obtained total with the cumulative calculated total of the attribute values ​​for the consistent attributes.

[0188] If the CPU 21 can confirm that the associated Table 4 has consistency, it performs the processes from step S520 onward in FIG.

[0189] By the correction process described above, even if the association information in the associated Table 4 is not "normal," the associated Table 4 may be corrected to have consistency.

[0190] In addition, in step S360 of the extraction process of the information processing device 10 according to the second embodiment shown in FIG. 6, if the association information is set to "shortage" or "surplus", after performing the correction process shown above, the attribute values ​​for each attribute included in the header row 6 may be obtained from each of the associated tables 4.

[0191] <Modification of continuity determination> Even in a situation where there is no order attribute in the header row 6 of a header table 4A included in a table group, the CPU 21 can associate the divided tables 4 so that they have continuity by using external information that exists outside the header table 4A and the subsequent table 4B and that represents the continuity between the header table 4A and the subsequent table 4B.

[0192] 9 is a diagram showing an example of page number 5, which is an example of external information. If page number 5 is assigned to the page of image 2 that includes table 4, CPU 21 can use page number 5 as external information to associate divided table 4, which has been classified into a table group, so that it has continuity.

[0193] 10 is a diagram showing an example of a table number 7, which is an example of external information. As shown in FIG. 10, table numbers 7, such as "Table 1" or "Fig. 1," may be written around each divided table 4 to indicate the order in which the table 4 is written. In particular, in documents such as papers and reports, there is a convention of writing a table number 7 for each table 4 that is divided in the middle of a table 4. Therefore, the CPU 21 can use the table numbers 7 as external information to associate the divided tables 4, which have been classified into a table group, so that they have continuity.

[0194] In addition, if the header row 6 of the header table 4A included in the table group contains an order attribute, the CPU 21 may first associate the divided tables 4 so that they have continuity based on the arrangement of attribute values ​​for the order attribute, and then use external information to reconfirm whether the associated tables 4 have continuity.

[0195] So far, we have explained an example of associating Tables 4 with each other so as to maintain at least one of continuity and consistency when Table 4 is divided by page as shown in Figures 1 and 2. However, the division of Table 4 is not limited to page units, and there are also divisions in which Table 4 is divided within the same page, as shown in Figure 11, for example.

[0196] Naturally, as long as the same Table 4 is originally divided into multiple Tables 4, the information processing device 10 can associate the Tables 4 so as to have at least one of continuity and consistency, and acquire an attribute value for each attribute included in the header row 6, regardless of the division form of the Table 4. Therefore, even if the Table 4 is divided in the form shown in Fig. 11, it goes without saying that it is possible to associate the Tables 4 so as to have at least one of continuity and consistency, and acquire an attribute value for each attribute included in the header row 6.

[0197] While one aspect of the information processing device 10 has been described above using the embodiment, the disclosed form of the information processing device 10 is merely an example, and the form of the information processing device 10 is not limited to the scope described in the embodiment. Various modifications or improvements can be made to the embodiment without departing from the gist of the present disclosure, and forms incorporating such modifications or improvements are also included in the technical scope of the disclosure. For example, the order of the extraction processes shown in Figures 5 to 8 may be changed without departing from the gist of the present disclosure.

[0198] In the above embodiment, the extraction process is implemented by software. However, the same process as the extraction process shown in Figures 5 to 8 may be implemented by hardware. In this case, the process can be performed faster than when the extraction process is implemented by software.

[0199] In the above embodiment, the term "processor" refers to a processor in a broad sense, including a general-purpose processor (e.g., CPU 21) and a dedicated processor (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).

[0200] Furthermore, the operations of the processors in the above embodiments may not only be performed by a single processor, but may also be performed by multiple processors located at physically separate locations working together. Furthermore, the order of the operations of the processors is not limited to the order described in the above embodiments, and may be changed as appropriate.

[0201] In the above embodiment, an example in which the information processing program is stored in the ROM 22 has been described, but the storage destination of the information processing program is not limited to the ROM 22. The information processing program of the present disclosure can also be provided in a form recorded on a storage medium readable by the computer 20. For example, the information processing program may be provided in a form recorded on an optical disc such as a CD-ROM (Compact Disk Read Only Memory) or a DVD-ROM (Digital Versatile Disk Read Only Memory). Furthermore, the information processing program may be provided in a form recorded on a portable semiconductor memory such as a USB (Universal Serial Bus) memory or a memory card.

[0202] ROM 22, non-volatile memory 24, CD-ROM, DVD-ROM, USB, and memory cards are examples of non-transitory storage media.

[0203] Furthermore, the information processing device 10 may download an information processing program from an external device connected to the communication unit 27 via a communication line and store the downloaded information processing program in a non-transitory storage medium. In this case, the CPU 21 of the information processing device 10 reads the information processing program downloaded from the external device from the non-transitory storage medium and executes the extraction process. [Explanation of symbols]

[0204] 2 images 3(3A, 3B) cells 4 tables 4A Heading Table (Table) 4B Subsequent Table (Table) 5 Page Number 6 Heading Row 7 Table Number 8 areas 10. Information processing equipment 11 Image Reception Section 12 User Interface Section (UI Section) 13 Image processing section 13A OCR processing section 13B Extraction part 14 Control Unit 15 Output section 18 Association Rules 20 Computer 21 CPU 22 ROM 23 RAM 24 Non-volatile memory 25 I / O 26 Bus 27 Communication Unit 28 Input Units 29 Display Unit 16 OCR result DB 17 Extraction result DB WD Column Width

Claims

1. a processor; The processor: extracting, from the plurality of tables that have been imaged, a group of tables that have a common feature in terms of the table structure and are made up of a first table with a header line indicating the attributes of the attribute values ​​in the table, and at least one or more second tables without the header line; associating each table included in the group of tables so that the order of the tables is continuous using an attribute value of an attribute included in the header row that indicates continuity between the first table and the second table; acquiring attribute values ​​corresponding to each attribute included in the header row from a series of tables constituted by the first table and the second table arranged in succession; detecting whether or not there is an excess or deficiency in the tables that make up the series of tables using an attribute value of an attribute included in the header row that indicates consistency regarding excess or deficiency in the tables that make up the series of tables; From the series of tables for which no excess or deficiency has been detected, attribute values ​​corresponding to each attribute included in the header row are obtained. Information processing device.

2. the attribute representing consistency is an attribute in which a sum of each attribute value in the series of tables corresponding to any of the attributes included in the header row is previously indicated in an image including any of the tables constituting the series of tables, The processor detects that there is no excess or deficiency in the set of tables if the sum of the attribute values ​​corresponding to the attribute representing the consistency of each of the tables constituting the set of tables matches the sum that is previously displayed in an image including any of the tables constituting the set of tables. The information processing device according to claim 1 .

3. the attribute representing continuity is an attribute representing an arrangement order of attribute values ​​in the table group; The processor detects that one table and the other table are continuous when an attribute value of an attribute representing the continuity in the last row of one table and an attribute value of an attribute representing the continuity in the first row of the other table follow a predetermined regularity as a state in which attribute values ​​are aligned.

3. The information processing device according to claim 1.

4. A processor is provided, The processor: extracting, from the plurality of tables that have been imaged, a group of tables that have a common feature in terms of the table structure and are made up of a first table with a header line indicating the attributes of the attribute values ​​in the table, and at least one or more second tables without the header line; associating each table included in the group of tables so that the order of the tables is continuous, using an attribute value of an attribute included in the header row that indicates continuity between the first table and the second table; acquiring attribute values ​​corresponding to each attribute included in the header row from a series of tables constituted by the first table and the second table arranged in succession; If the header row does not include an attribute representing the continuity, the presence or absence of an excess or deficiency in the tables constituting the table group is detected using an attribute value of an attribute included in the header row that represents consistency regarding excess or deficiency in the tables constituting the table group; From the table group in which no excess or deficiency has been detected, attribute values ​​corresponding to each attribute included in the header row are obtained. Information processing device.

5. if the header row does not include an attribute representing the continuity, the processor associates each of the tables included in the group of tables so that the order of the tables is continuous using external information that is information existing outside the first table and the second table and represents the continuity of the first table and the second table; From a series of tables constituted by the first table and the second table, which are arranged in succession, attribute values ​​corresponding to each attribute included in the header row are acquired. The information processing device according to claim 4 .

6. The processor associates each table included in the group of tables so that the order of the tables is continuous, using, as the external information, a page number of a page including either the first table or the second table, or a table number indicating the order of the first table and the second table. The information processing device according to claim 5 .

7. a processor; The processor: extracting, from the plurality of tables that have been imaged, a group of tables that have a common feature in terms of the table structure and are made up of a first table with a header line indicating the attributes of the attribute values ​​in the table, and at least one or more second tables without the header line; detecting whether or not there is a deficiency or excess in the tables that make up the table group using an attribute value of an attribute included in the header row that indicates consistency regarding deficiencies or excess in the tables that make up the table group; From the table group in which no excess or deficiency has been detected, attribute values ​​corresponding to each attribute included in the header row are obtained. Information processing device.

8. On the computer, extracting, from the plurality of tables that have been imaged, a group of tables that have a common feature in terms of the table structure and are made up of a first table with a header line indicating the attributes of the attribute values ​​in the table, and at least one or more second tables without the header line; associating each table included in the group of tables so that the order of the tables is continuous using an attribute value of an attribute included in the header row that indicates continuity between the first table and the second table; acquiring attribute values ​​corresponding to each attribute included in the header row from a series of tables constituted by the first table and the second table arranged in succession; detecting whether or not there is an excess or deficiency in the tables that make up the series of tables using an attribute value of an attribute included in the header row that indicates consistency regarding excess or deficiency in the tables that make up the series of tables; and executing a process of acquiring attribute values ​​corresponding to each attribute included in the header row from the series of tables in which no excess or deficiency has been detected. Information processing program.

9. A computer, extracting, from the plurality of tables that have been imaged, a group of tables that have a common feature in terms of the table structure and are made up of a first table with a header line indicating the attributes of the attribute values ​​in the table, and at least one or more second tables without the header line; associating each table included in the group of tables so that the order of the tables is continuous using an attribute value of an attribute included in the header row that indicates continuity between the first table and the second table; acquiring attribute values ​​corresponding to each attribute included in the header row from a series of tables constituted by the first table and the second table arranged in succession; If the header row does not include an attribute representing the continuity, the presence or absence of an excess or deficiency in the tables constituting the table group is detected using an attribute value of an attribute included in the header row that represents consistency regarding excess or deficiency in the tables constituting the table group; and executing a process for acquiring attribute values ​​corresponding to each attribute included in the header row from the table group in which no excess or deficiency has been detected. Information processing program.

10. On the computer, extracting, from the plurality of tables that have been imaged, a group of tables that have a common feature in terms of the table structure and are made up of a first table with a header line indicating the attributes of the attribute values ​​in the table, and at least one or more second tables without the header line; detecting whether or not there is a deficiency or excess in the tables that make up the table group using an attribute value of an attribute included in the header row that indicates consistency regarding deficiencies or excess in the tables that make up the table group; and executing a process for acquiring attribute values ​​corresponding to each attribute included in the header row from the table group in which no excess or deficiency has been detected. Information processing program.

Citation Information

Patent Citations

  • Table information reading device, table information reading method and program

    JP2020155054A