Document image processing system, document image processing method, and document image processing program

The document image processing system addresses the challenge of processing documents with different layouts by detecting tables and specifying cell attributes without template data, achieving accurate and efficient information extraction.

JP7699773B2Active Publication Date: 2025-06-30NET SMILE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021105251
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-24
Publication Date
2025-06-30
Estimated Expiration
2041-06-24

AI Technical Summary

Technical Problem

Existing document identification systems require template data for processing documents with different layouts, making it difficult to accurately detect attribute values in documents with unknown layouts, and often necessitate pre-setting character strings for extraction, leading to incomplete or incorrect extractions.

Method used

A document image processing system that detects tables in documents without using template data, identifies cells and text objects within these tables, performs character recognition, and uses node data and classification processes to accurately specify the attributes of cells, enabling accurate extraction of relevant information without pre-defined templates.

Benefits of technology

The system effectively specifies the attributes of cells in tables without template data, ensuring accurate extraction of information regardless of the document layout, and improves the efficiency of processing documents with varying formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699773000001
    Figure 0007699773000001
  • Figure 0007699773000002
    Figure 0007699773000002
  • Figure 0007699773000003
    Figure 0007699773000003
Patent Text Reader

Abstract

To accurately identify an attribute of a cell in a table without using template data.SOLUTION: A table detection unit 22 detects a table in a document image and a cell in the table, and generates cell geometry data for the cell. A text object detection unit 23 detects a text object in the document image and generates text object geometry data for the text object. A cell attribute identification unit 25 identifies the text object in the cell based on the cell geometry data and the text object geometry data, generates, for each cell, node data including the cell geometry data, the text object geometry data of the text object in the cell, and text data of the text object in the cell, and executes predetermined classification processing on a node data set including the node data corresponding to the table to identify an attribute of the cell.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a document image processing system, a document image processing method, and a document image processing program.

Background Art

[0002] In a certain form identification system, a form format table is created in advance by a user, and the form format table includes field information indicating the position, size, character type, etc. of a character recognition target area specified by the user. Then, based on this form format (i.e., field information), character information (text data) in the form image is acquired from the image data of the form image (see, for example, Patent Document 1).

[0003] A certain image recognition device cuts out a partial image from a target image, recognizes characters and numbers in the partial image, and executes an extraction process of extracting characters and numbers that satisfy a predetermined condition from the characters and numbers (see, for example, Patent Document 2). In the extraction process, the image recognition device determines, for example, whether the recognized characters include a predetermined bank name set in advance, and when the characters include the predetermined bank name, extracts the characters and the numbers within a predetermined distance from the characters as a pair of a bank name and an account number.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the above-described document identification system, since template data for specifying the layout of documents such as forms (information on the positions where each attribute is described, etc.) is used, in order to process a plurality of documents with different layouts, template data must be created in advance for each layout, and complicated work is required in advance. Also, for a document with an unknown layout, it is difficult to accurately detect the attribute value in the document image for a certain attribute with the above-described technology.

[0006] Also, in the above-described image recognition apparatus, although template data is not required, it is necessary to preset the character string to be extracted (the above-described bank name), and a character string that is not set is not extracted. Further, in the above-described image recognition apparatus, although the numbers within a predetermined distance from the above-described bank name are extracted as account numbers, even if the distance between two character objects is short, if the two are not related, or even if the distance between the two is long, if the two are related, there is a possibility that the desired character string may not be correctly extracted.

[0007] FIG. 5 is a diagram showing an example of a document image including a table. For example, as shown in FIG. 5, the document image 101 of the medical examination report includes a table 111 showing the results of the medical examination. The table 111 includes a combination of a label indicating an examination item and a value that is the examination result of the examination item. In the table 111, regardless of the presence or absence of a grid line, the cells are two-dimensionally arranged, and labels and values (numerical values or character strings other than numerical values) corresponding to the labels are described in the cells. However, even if the distance between the two (label and value) is short, if the two are not related, or even if the distance between the two is long, if the two are related, there are cases.

[0008] The present invention has been made in view of the above problems, and an object thereof is to obtain a document image processing system, a document image processing method, and a document image processing program that accurately specify the attributes of cells in a table without using template data.

Means for Solving the Problems

[0009] The document image processing system according to the present invention includes a table detection unit that detects a table in a document image, detects at least cells in the table, and generates cell geometry data indicating the positions and sizes of the cells; a text object detection unit that detects text objects in the document image and generates text object geometry data indicating the positions and sizes of the text objects; a character recognition processing unit that performs character recognition processing on the text objects to generate text data corresponding to the text objects; and (a) based on the cell geometry data and the text object geometry data, identifies text objects within the cells, (b) for each cell, generates node data including the cell geometry data, the text object geometry data of the text objects within the cell, and the text data of the text objects within the cell, and generates a node data set including the node data corresponding to the table, and (c) a cell attribute specifying unit that performs a predetermined classification process on the node data set to specify the attributes of the cells for each cell.

[0010] The document image processing method according to the present invention detects a table in a document image The computer is and detects at least cells in the table The computer is and generates cell geometry data indicating the positions and sizes of the cells The computer is and detects text objects in the document image The computer is and generates text object geometry data indicating the positions and sizes of the text objects The computer is and performs character recognition processing on the text objects The computer is to generate text data corresponding to the text objects The computer is and (a) based on the cell geometry data and the text object geometry data, identifies text objects within the cells The computer is and (b) for each cell, generates node data including the cell geometry data, the text object geometry data of the text objects within the cell, and the text data of the text objects within the cell The computer isGenerate a node data set including node data corresponding to a table The computer is Generate and perform a predetermined classification process on (c) the node data set The computer is Execute to specify the attributes of each cell The computer is and comprises the step of identifying.

[0011] The document image processing program according to the present invention causes a computer to function as the above-described table detection unit, the above-described text object detection unit, the above-described character recognition processing unit, and the above-described cell attribute specifying unit.

Advantages of the Invention

[0012] According to the present invention, there are provided a document image processing system, a document image processing method, and a document image processing program that accurately identify the attributes of cells in a table without using template data for specifying the description positions of attributes in a document image.

[0013] The above or other objects, features, and advantages of the present invention will become more apparent from the following detailed description together with the accompanying drawings.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0016] FIG. 1 is a block diagram showing the configuration of a document image processing system according to an embodiment of the present invention. The document image processing system shown in FIG. 1 is composed of a single information processing device (such as a personal computer or a server), but the processing unit described later may be distributed among a plurality of information processing devices capable of data communication with each other. Further, such a plurality of information processing devices may include a GPU (Graphics Processing Unit) that performs parallel processing of specific operations.

[0017] The document image processing system shown in FIG. 1 includes a storage device 1, a communication device 2, an image reading device 3, and an arithmetic processing device 4.

[0018] The storage device 1 is a non-volatile storage device such as a flash memory or a hard disk, and stores various data and programs.

[0019] Here, an image processing program 11 is stored in the storage device 1, and system setting data (such as coefficient setting values of neural networks used in each processing unit described later) is stored as necessary. Note that the image processing program 11 may be stored in a portable computer-readable recording medium such as a CD (Compact Disk). In that case, for example, the image processing program 11 is installed from the recording medium into the storage device 1. Further, the image processing program 11 may be a single program or an aggregate of a plurality of programs.

[0020] The communication device 2 is a device capable of data communication such as a network interface, a peripheral device interface, or a modem, and performs data communication with other devices as necessary. The image reading device 3 optically reads a document image from a document and generates image data (such as raster image data) of the document image. Note that the communication device 2 and the image reading device 3 are provided as necessary.

[0021] The arithmetic processing unit 4 is a computer including a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc., and operates as various processing units by loading a program from the ROM, the storage device 1, etc. into the RAM and executing it with the CPU.

[0022] Here, by executing the image processing program 11, the arithmetic processing unit 4 operates as a document image acquisition unit 21, a table detection unit 22, a text object detection unit 23, a character recognition processing unit 24, a cell attribute specification unit 25, a data output unit 26, and a machine learning processing unit 27.

[0023] The document image acquisition unit 21 acquires a document image as image data such as raster image data. The document image is an image of a document that includes an attribute label (text such as a heading) and an attribute value (text such as a numerical value or other character string) for one or more attributes (description items, etc.) in a table, such as receipts (including invoices), bills, delivery notes, flyers for advertising and announcements, answered questionnaire sheets, health diagnosis reports, etc. For example, the document image acquisition unit 21 reads out a document image as image data stored in the storage device 1, acquires a document image as image data received by the communication device 2 via a communication path such as a network, or acquires a document image as image data generated by the image reading device 3.

[0024] The table detection unit 22 detects a table in the acquired document image without using template data, generates table geometry data indicating the position and size of the table, and at least detects cells in the table and generates cell geometry data indicating the position and size of the cells.

[0025] In this embodiment, the table detection unit 22 detects at least one of rows and columns together with the cells in the table, and generates at least one (here, both) of row geometry data indicating the position and size of the row and column geometry data indicating the position and size of the column.

[0026] Note that the table detection unit 22 detects a table, rows, columns, and cells in the document image, and generates their geometry data according to an existing method using a neural network.

[0027] The text object detection unit 23 detects text objects in the acquired document image without using template data, and generates text object geometry data indicating the position and size of the text objects.

[0028] Specifically, the text object detection unit 23 (a) excludes objects other than characters in the document image (such as photo objects, graphic objects, and ruled line objects) and detects character objects, and (b) groups them into "words" based on the positions of the respective character objects and extracts text objects.

[0029] Note that the text object detection unit 23 extracts character objects in the document image using existing technologies (such as region separation processing and machine-learned deep neural networks).

[0030] The character recognition processing unit 24 performs character recognition processing on the detected text objects (raster images) and generates text data (character code strings) corresponding to the text objects. Note that existing technologies are used for this character recognition processing.

[0031] The cell attribute specifying unit 25 identifies text objects in a cell based on (a) the above-described cell geometry data and text object geometry data, and generates, for each cell, node data including the cell geometry data, the text object geometry data of the text objects in the cell, and the text data of the text objects in the cell, and generates a node data set including the node data corresponding to the table, and (c) executes a predetermined classification process on the node data set to specify the attributes of the cells for each cell.

[0032] Note that the positions of the cells, rows, columns, and text objects in the node data are coordinate values of two-dimensional coordinates, and their sizes are the lengths at the respective coordinates of the two-dimensional coordinates. Also, the positions may be indicated by the positions of predetermined portions (any of the four corners, the center, etc.) of the rectangular regions of the cells, rows, columns, and text objects, and may be represented by the relative positions from predetermined portions (any of the four corners, the center, etc.) of the table. This relative position is derived from the absolute position of the table in the document image (the position in the table geometry data) and the absolute position of the cell, row, column, or text object (the position in the original geometry data).

[0033] Specifically, for each detected cell, when the region of the text object (bounding box) specified from the text object geometry data is included in the region of the cell specified from the cell geometry data of the cell, the cell attribute specifying unit 25 determines that the text object is a text object in the cell.

[0034] FIG. 2 is a diagram for explaining the node data generated in the document image processing system according to the embodiment of the present invention. FIG. 3 is a diagram for explaining the classification process in the document image processing system according to the embodiment of the present invention.

[0035] In this embodiment, as shown in FIG. 2, the node data of a certain cell further includes at least one (both in this case) of the row geometry data of the row to which the cell belongs and the column geometry data of the column to which the cell belongs, in addition to the cell geometry data of the cell, the text object geometry data of the text objects in the cell, and the text data of the text objects in the cell.

[0036] In this embodiment, as shown in FIG. 3, the cell attribute specifying unit 25 inputs the above-described node data set as input data into a graph neural network (GNN) that has been machine-learned, and specifies the output data of the GNN as the attribute of the above-described cell. Note that existing GNNs and their machine learning can be used for this GNN.

[0037] Also, the attribute of the cell includes at least the cell type of the cell, and the cell type is either a label or an attribute value (that is, in this case, the cell is classified into either a label cell or an attribute value cell). Note that the output data of the GNN for each node data is the probability (a numerical value within the range of 0 to 1) for each of the possible values (labels and attribute values here) of the cell type of the cell corresponding to that node data. Based on the value of the probability, the cell type is determined to be one of the possible values (labels and attribute values here) of the cell type, for example, by classification using a threshold value.

[0038] Note that the cell attribute specifying unit 25 may classify the node data into clusters by performing clustering on the node data based on the feature amounts indicated by the node data, and specify the attribute of the cell based on the clusters. For example, the cell attribute specifying unit 25 may specify the attribute of the cell by the above-described clustering instead of the GNN, or may specify the attribute of the cell by the above-described clustering instead of the GNN when the reliability of specifying the attribute of the above-described cell by the GNN is low.

[0039] For example, as this feature quantity, a feature vector corresponding to node data (all or a specific part) generated according to an existing method such as Word2vec (Skip-Gram model) is used. Also, a large number of combinations of node data and cell type values for that node data are collected, and the central value (average of feature vectors) for each value of the cell type (here, label or attribute value) is specified (as the center of each cluster), and from the position indicated by the feature vector of the node data to be classified, the value of the cell type (here, label or attribute value) having the closest central value is selected as the cell type value corresponding to that node data.

[0040] The data output unit 26 adds the cell attribute corresponding to each node data to the node data, and stores the node data set in the storage device 1 in a predetermined data format, or transmits it via the communication device 2.

[0041] With this output data (node data set), for example, it is possible to identify cells of labels and cells of attribute values within a column, or cells of labels and cells of attribute values within a row.

[0042] The machine learning processing unit 27 executes machine learning processing for performing the machine learning of the GNN in the cell attribute specifying unit 25 described above. Note that the above-described machine learning processing unit 27 is not essential and may be provided as necessary. Also, when the machine learning of the cell attribute specifying unit 25 (GNN) is completed, the machine learning processing unit 27 may not be provided.

[0043] Next, the operation of the document image processing system according to the present embodiment will be described. FIG. 4 is a flowchart for explaining the operation of the document image processing system shown in FIG. 1.

[0044] First, the document image acquisition unit 21 acquires a document image (step S1).

[0045] Next, the table detection unit 22 detects a table in the document image without using template data, and also detects cells, columns, and rows within the table, and generates cell geometry data, column geometry data, and row geometry data (step S2).

[0046] Also, the text object detection unit 23 detects text objects in the document image without using template data, and generates text object geometry data (step S2). The character recognition processing unit 24 performs character recognition processing on the detected text objects, and generates text data of the text objects.

[0047] Note that the above-described processing by the table detection unit 22 and the above-described processing by the text object detection unit 23 may be executed in parallel, or when performing these processes in order, either one may be executed first.

[0048] Then, for each detected cell, the cell attribute specifying unit 25 specifies text objects within the cell based on the above-described cell geometry data and text object geometry data, and generates node data including the cell geometry data, the text object geometry data of the text objects within the cell, and the text data of the text objects within the cell (step S3).

[0049] Next, the cell attribute specifying unit 25 generates a node data set with the node data corresponding to all the cells detected within the table for each table, and performs a predetermined classification process on the node data set to classify the attributes of each cell and specify the attributes of each cell (here, the cell type called a label or an attribute value) (step S4).

[0050] Then, the data output unit 26, for example, adds the attributes of each cell to the node data corresponding to the cell, and for each table, stores the node data set in the storage device 1 in a predetermined data format as output data, or transmits it via the communication device 2 (step S5).

[0051] In this way, for each table in the document image, attribute data at the cell unit is generated.

[0052] Further, the data output unit 26 may receive a search request in a predetermined format, and in the node data set generated as described above according to the search request, search for an attribute value corresponding to the label specified by the search request, and output a combination of the label and the attribute value. For example, first, in the generated node data set, cells, rows, and columns including the search target label are specified, and rows or columns including cells with attribute values in the row or column are specified. If there is one cell with an attribute value in the row or column, the attribute value is specified as the attribute value corresponding to the label. If there are multiple cells with attribute values, those attribute values are specified as the attribute values corresponding to the label, and the label of the row or column (column if the cell with the attribute value is specified in the row including the search target label, row if the cell with the attribute value is specified in the column including the search target label) of those cells with attribute values may be associated with and attached to those attribute values respectively.

[0053] For example, in the node data set of table 111 in FIG. 5, when "height" is specified as the search target, as attribute values, three values "161.0", "161.2", and "161.1" in the row including "height" are detected. The label "this time" is attached to "161.0", the label "last time" is attached to "161.2", and the label "the time before last" is attached to "161.1".

[0054] Also, it may be possible to specify two labels as the search target labels in the search request. In that case, an attribute value is detected in the same manner as described above with one of the two labels, and among the detected attribute values, those whose label of the row or column to which the cell of the attribute value belongs matches the other label are determined and detected as the attribute values corresponding to the two labels.

[0055] As described above, according to the above embodiment, the table detection unit 22 detects a table in the document image, at least detects cells in the table, and generates cell geometry data indicating the positions and sizes of the cells. The text object detection unit 23 detects text objects in the document image and generates text object geometry data indicating the positions and sizes of the text objects. The character recognition processing unit 24 performs character recognition processing on the text objects to generate text data corresponding to the text objects. The cell attribute specifying unit 25 (a) specifies text objects in the cell based on the cell geometry data and the text object geometry data, and (b) for each cell, generates node data including the cell geometry data, the text object geometry data of the text objects in the cell, and the text data of the text objects in the cell, and generates a node data set including the node data corresponding to the table, and (c) performs a predetermined classification process on the node data set to specify the attributes of the cells for each cell.

[0056] Thereby, without using template data, the attributes of the cells in the table are accurately specified. Further, since the attribute value corresponding to the label is detected without considering the distance (Euclidean distance) between the label and the attribute value, the combination of the label and the attribute value in the table is accurately specified regardless of the distance (Euclidean distance) between the label and the attribute value.

[0057] Note that various changes and modifications to the above-described embodiment will be apparent to those skilled in the art. Such changes and modifications may be made without departing from the spirit and scope of the subject matter and without weakening the intended advantages. That is, it is intended that such changes and modifications be included in the claims.

[0058] For example, in the above embodiment, after the above-described processing is completed, the image data of the document image may be immediately deleted from the system.

[0059] In the above-described embodiment, the number of nodes in the input data of the GNN is set to a predetermined value that is greater than or equal to the maximum value of the number of node data in the node data set (i.e., the number of cells in the table). When the number of node data is less than the number of nodes in the input data of the GNN, fixed values are used as the missing node data, and the corresponding output data is discarded.

[0060] In the above-described embodiment, the cell attribute is the cell type, and the cell type is either a label or an attribute value. However, the cell type may also take a header (such as the title of the table), among other things, in addition to the label and the attribute value.

[0061] In the above-described embodiment, rows and columns may not be detected, and the node data may not include row geometry data and column geometry data. Even in that case, based on the position of the cell in the cell geometry data, the cell of the attribute value can be searched for and detected along the horizontal and vertical directions from the position of the cell of the label, thereby detecting the attribute value corresponding to the label.

Industrial Applicability

[0062] The present invention is applicable to, for example, recognition processing of document images such as forms.

Explanation of Signs

[0063] 4 Arithmetic processing unit (an example of a computer) 11 Image processing program (an example of a document image processing program) 22 Table detection unit 23 Text object detection unit 24 Character recognition processing unit 25 Cell attribute specifying unit

Claims

1. A table detection unit that detects a table in a document image, detects at least cells in the table, and generates cell geometry data indicating the positions and sizes of the cells; A text object detection unit that detects a text object in the document image and generates text object geometry data indicating the positions and sizes of the text objects; A character recognition processing unit that performs character recognition processing on the text object to generate text data corresponding to the text object; (a) identifying the text object in the cell based on the cell geometry data and the text object geometry data; (b) generating node data for each cell, including the cell geometry data, the text object geometry data of the text object in the cell, and the text data of the text object in the cell, and generating a node data set including the node data corresponding to the table; (c) performing a predetermined classification process on the node data set to identify the attributes of the cell for each cell; a cell attribute identification unit; A document image processing system characterized by comprising the above.

2. The cell attribute identification unit inputs the node data set as input data into a graph neural network that has been trained by machine learning, and identifies the output data of the graph neural network as the attributes of the cell. The document image processing system according to Claim 1.

3. The cell attribute identification unit classifies the node data by clustering the node data based on the feature amounts indicated by the node data, and identifies the attributes of the cell based on the clusters. The document image processing system according to Claim 1.

4. The table detection unit detects at least one of rows and columns together with the cells in the table, and generates at least one of row geometry data indicating the positions and sizes of the rows and column geometry data indicating the positions and sizes of the columns; The node data of a certain cell further includes at least one of the row geometry data of the row to which the cell belongs and the column geometry data of the column to which the cell belongs. The document image processing system according to any one of claims 1 to 3, characterized in that

5. The attributes of the cell include at least the cell type of the cell, The cell type includes a label and an attribute value, The document image processing system according to any one of claims 1 to 4, characterized in that

6. Steps for a computer to detect a table in a document image, detect at least cells in the table, and generate cell geometry data indicating the positions and sizes of the cells by the computer; Steps for a computer to detect text objects in the document image and generate text object geometry data indicating the positions and sizes of the text objects by the computer; Steps for a computer to perform character recognition processing on the text object and generate text data corresponding to the text object by the computer; (a) Based on the cell geometry data and the text object geometry data, the computer identifies the text object within the cell, (b) for each cell, the computer generates node data including the cell geometry data, the text object geometry data of the text object within the cell, and the text data of the text object within the cell, and generates a node data set including the node data corresponding to the table by the computer, (c) the computer performs a predetermined classification process on the node data set to identify the attributes of the cell for each cell; A document image processing method characterized by comprising

7. A computer, A table detection unit that detects a table in a document image, detects at least cells in the table, and generates cell geometry data indicating the positions and sizes of the cells; A text object detection unit that detects text objects in the document image and generates text object geometry data indicating the positions and sizes of the text objects; A character recognition processing unit that performs character recognition processing on the text object and generates text data corresponding to the text object, and (a) identifying the text object within the cell based on the cell geometry data and the text object geometry data; (b) generating, for each cell, node data including the cell geometry data, the text object geometry data of the text object within the cell, and the text data of the text object within the cell, and generating a node data set including the node data corresponding to the table; (c) performing a predetermined classification process on the node data set to identify, for each cell, a cell attribute specifying unit that specifies the attribute of the cell, A document image processing program that functions as.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and program

    JP2013080348A

  • Document identification program, document identification device, document identification system, and document identification method

    JP2016048444A

  • Character estimation system, character estimation method, and character estimation program

    JP2019079347A

  • Image recognition device, image recognition method, image recognition program and image recognition system

    JP2020170264A