Information processing system, information processing method, and information processing program

The information processing system uses machine learning and a three-stage classification process to accurately identify document types by extracting title rectangles and recognizing character strings, overcoming limitations of conventional methods that rely on title and font size.

JP2026057913APending Publication Date: 2026-04-03SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Conventional methods for identifying the type of a document based on the presence of a title and specific font size are ineffective when these conditions are not met, leading to inappropriate identification of document types.

Method used

An information processing system that uses machine learning to extract a title rectangle from document image data, performs character recognition on the nearest string rectangle, and refers to a storage unit associating document types with type information to accurately identify the document type, employing a three-stage classification determination process to ensure accurate identification even when title or font size conditions are not satisfied.

Benefits of technology

The system effectively identifies document types regardless of title location or font size, ensuring accurate categorization through a robust three-stage classification process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026057913000001_ABST
    Figure 2026057913000001_ABST
Patent Text Reader

Abstract

This invention provides an information processing system, an information processing method, and an information processing program that can appropriately identify the type of document. [Solution] The image processing device 1 includes an extraction processing unit 112 that extracts a title rectangle containing the title of a document from image data of the document using machine learning; a recognition processing unit 113 that performs character recognition processing on a first string rectangle closest to the position of the title rectangle from the image data and obtains the recognition result; and a specification processing unit 114 that identifies the type of document based on the type information included in the recognition result by referring to a storage unit 12 that stores in advance the association of the type of document and type information related to that type for each type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for performing image processing on image data of a document to identify the type of the document.

Background Art

[0002] Conventionally, a technique for performing image processing on image data of a document (such as a form) to identify the type of the document (such as an invoice, estimate, delivery note, order form, receipt, etc.) has been known. For example, there is known a technique for storing form type information in which a form type and a title of a form classified into the form type are associated with each other, and identifying the form type of the form by identifying the title of the form from character strings extracted from a read image of the form (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the conventional technique, for example, the type of a form is identified based on specific conditions such as the title of the form being located above and being a character string with a font size of a certain level or more. Therefore, in the case of a form that does not satisfy the specific conditions, there arises a problem that the type cannot be appropriately identified.

[0005] An object of the present disclosure is to provide an information processing system, an information processing method, and an information processing program capable of appropriately identifying the type of a document.

Means for Solving the Problems

[0006] An information processing system according to one aspect of this disclosure comprises an extraction processing unit, a recognition processing unit, and a specification processing unit. The extraction processing unit extracts a title rectangle containing the title of a document from image data of the document using machine learning. The recognition processing unit performs character recognition processing on a first string rectangle closest to the position of the title rectangle from the image data and obtains a recognition result. The specification processing unit refers to a storage unit that stores in advance the association of document types and type information related to those types for each type, and identifies the type of the document based on the type information included in the recognition result.

[0007] An information processing method relating to another aspect of this disclosure is a method in which one or more processors perform the following actions: extract a title rectangle containing the title of a document from image data of the document using machine learning; perform character recognition processing on a first string rectangle closest to the position of the title rectangle from the image data to obtain a recognition result; and identify the type of the document based on the type information included in the recognition result by referring to a storage unit that stores in advance, in association with the type of document and type information related to that type for each type.

[0008] An information processing program relating to another aspect of this disclosure is a program that causes one or more processors to perform the following actions: extract a title rectangle containing the title of a document from image data of the document using machine learning; perform character recognition processing on a first string rectangle closest to the position of the title rectangle from the image data to obtain a recognition result; and identify the type of the document based on the type information included in the recognition result by referring to a storage unit that stores in advance the association of the type of document and type information related to that type for each type. [Effects of the Invention]

[0009] This disclosure provides an information processing system, an information processing method, and an information processing program that can appropriately identify the type of document. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 is a functional block diagram showing the configuration of an image processing system according to an embodiment of this disclosure. [Figure 2] Figure 2 shows an example of an invoice included in a document according to the embodiment of this disclosure. [Figure 3] Figure 3 shows an example of a receipt included in the document according to the embodiment of this disclosure. [Figure 4] Figure 4 is a flowchart illustrating an example of a procedure for processing a document image performed by the image processing system according to the embodiment of this disclosure. [Figure 5] Figure 5 is a flowchart illustrating an example of the procedure for document segmentation extraction performed by the image processing system according to the embodiment of this disclosure. [Figure 6] Figure 6 shows a specific example of a character rectangle extracted from document image data by AI in an image processing system according to the embodiment of this disclosure. [Figure 7] Figure 7 shows an example of a string list in which strings extracted in an image processing system according to the embodiment of this disclosure are registered. [Figure 8] Figure 8 shows a method for calculating the unit character area in an image processing system according to an embodiment of this disclosure. [Figure 9] Figure 9 is a flowchart illustrating an example of the procedure for the first classification determination process performed by the image processing system according to the embodiment of this disclosure. [Figure 10] Figure 10 shows a specific example of keyword information referenced in the document segmentation extraction process according to the embodiment of this disclosure. [Figure 11] Figure 11 is a flowchart illustrating an example of the procedure for the second classification determination process performed by the image processing system according to the embodiment of this disclosure. [Figure 12] Figure 12 shows a specific example of a character rectangle extracted from document image data by AI in an image processing system according to an embodiment of this disclosure. [Figure 13]FIG. 13 is a flowchart for explaining an example of the procedure of the third section determination process executed in the image processing system according to the embodiment of the present disclosure. [Figure 14] FIG. 14 is a diagram for explaining the expansion process of the unit character area executed in the image processing system according to the embodiment of the present disclosure. [Figure 15] FIG. 15 is a flowchart for explaining another example of the procedure of the document section extraction process executed in the image processing system according to the embodiment of the present disclosure.

Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are an example of embodying the present disclosure and do not have the character of limiting the technical scope of the present disclosure.

[0012] [Image Processing System 10] FIG. 1 is a block diagram showing the configuration of an image processing system 10 according to an embodiment of the present disclosure. The image processing system 10 includes an image processing apparatus 1 and an operation terminal 2. The image processing apparatus 1 and the operation terminal 2 are connected to each other via a network N1 (for example, the Internet, a LAN, etc.). The image processing system 10 may include a plurality of operation terminals 2.

[0013] In the image processing system 10, the image processing apparatus 1 acquires image data of documents such as forms transmitted from the operation terminal 2, and extracts a desired character string (character string to be managed) from the image data. For example, the operation terminal 2 scans paper-based forms such as invoices, estimates, delivery notes, orders, receipts, receipts, and other documents to generate image data (such as PDF data) and transmits it to the image processing apparatus 1. Further, the operation terminal 2 creates a document file of the form based on the user's operation by, for example, a document creation application, and transmits the document file as image data (such as searchable PDF data (image + text data)) to the image processing apparatus 1. When the image processing apparatus 1 receives the image data transmitted from the operation terminal 2, it executes various processes described later on the image data to extract the character string to be managed included in the form. For example, the image processing apparatus 1 extracts the classification (type) of each form, the date of each form, the amount (total amount, etc.), company information (issuer, destination, registration number, etc.), and the like. Further, the image processing apparatus 1 registers the extracted character string in a predetermined database. For example, every time the image processing apparatus 1 acquires image data of an invoice, it extracts a character string related to the content of the invoice (for example, issue date, invoice amount, issuer, etc.) from the image data and registers it in the database for managing invoices. Further, every time the image processing apparatus 1 acquires image data of a receipt, it extracts a character string related to the content of the receipt (for example, issue date, total amount, issuer, etc.) from the image data and registers it in the database for managing receipts. Thereby, each form can be stored and managed as electronic data. Further, the image processing apparatus 1 outputs the extracted character string to the operation terminal 2 or the like and presents the character recognition result to the user.

[0014] The image processing system 10 is an example of the information processing system of the present disclosure. Note that the information processing system of the present disclosure may be configured by the image processing apparatus 1 alone.

[0015] [Image Processing Apparatus 1] As shown in Figure 1, the image processing device 1 includes a control unit 11, a storage unit 12, an operation display unit 13, a communication unit 14, and the like. The image processing device 1 may consist of one or more cloud servers, or one or more physical servers.

[0016] The communication unit 14 is a communication interface for connecting the image processing device 1 to the network N1 by wire or wireless connection and for performing data communication with the operating terminal 2 via the network N1 in accordance with a predetermined communication protocol. The network N1 consists of, for example, the internet, a LAN, etc.

[0017] The operation display unit 13 is a user interface comprising a display unit such as a liquid crystal display or an organic EL display that displays various types of information, and an operation unit such as a mouse, keyboard, or touch panel that accepts input.

[0018] The storage unit 12 is a non-volatile storage unit such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory that stores various types of information. The storage unit 12 stores a control program that causes the control unit 11 to execute the following processes: form determination processing, operation mode setting processing, document classification extraction processing, date extraction processing, amount extraction processing, company information extraction processing, etc. For example, the control program is non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, read by a reading device (not shown) such as a CD drive or DVD drive provided by the image processing device 1, and stored in the storage unit 12. The control program may also be distributed from a cloud server and stored in the storage unit 12.

[0019] Furthermore, the memory unit 12 stores image data (such as scanned data) of documents such as forms acquired from the operating terminal 2.

[0020] Figure 2 shows an invoice as an example of a document. As shown in Figure 2, the invoice includes strings of information such as the document type ("Invoice"), issue date, issuer's contact information (address, telephone number, fax number, person in charge), invoice amount, product name, quantity, standard price, discount amount, subtotal, consumption tax, and total amount. The user scans the invoice using the operation terminal 2 and uploads the image data P1 to the image processing device 1. The control unit 11 acquires the image data P1 of the invoice and stores it in the storage unit 12.

[0021] Figure 3 shows a receipt as another example of a form. As shown in Figure 3, the receipt contains strings of characters such as the issuer's contact information (address, telephone number), issue date, product name, sales price, subtotal, consumption tax, total amount, and change. The user scans the receipt using the operating terminal 2 and uploads the image data P1 to the image processing device 1. The control unit 11 acquires the image data P1 of the receipt and stores it in the storage unit 12.

[0022] In another embodiment, the control unit 11 may acquire the document file of the form created on the operation terminal 2 and store the document file in the storage unit 12.

[0023] The control unit 11 includes control devices such as a CPU, ROM, and RAM. The CPU is a processor that performs various arithmetic operations. The ROM stores control programs such as a BIOS and OS in advance to allow the CPU to perform various operations. The RAM stores various information and is used as a temporary storage memory (work area) for the various operations performed by the CPU. The control unit 11 controls the image processing device 1 by executing various control programs stored in advance in the ROM or storage unit 12 using the CPU.

[0024] Incidentally, in order to manage documents, it is necessary to accurately identify (distinguish) the category (type) of the document. Conventionally, a technique for identifying the category of a document is known, for example, to extract a string of characters where the title of the document is located at the top and the font size is above a certain level, and to identify the category of the document based on the character recognition result of that string. However, this method has the problem that it is not possible to appropriately identify the category of a document if it does not meet the specific conditions. In contrast to this, the image processing device 1 according to this disclosure has a configuration that makes it possible to appropriately identify the category of a document. The specific processing in the control unit 11 of the image processing device 1 will be described below. In this embodiment, "category" is synonymous with "type," and refers to types such as invoices, quotations, delivery notes, purchase orders, and receipts.

[0025] As shown in Figure 1, the control unit 11 includes various processing units such as an acquisition processing unit 111, an extraction processing unit 112, a recognition processing unit 113, a specific processing unit 114, and an output processing unit 115. The control unit 11 functions as these various processing units by executing various processes according to the control program. Some or all of the processing units included in the control unit 11 may be composed of electronic circuits. The control program may be a program that causes multiple processors to function as these various processing units.

[0026] [Document Image Processing (Overall Flow)] Here, when the control unit 11 acquires image data P1, it sequentially executes the following processes: form determination process, operation mode setting process, document classification extraction process, date extraction process, amount extraction process, and company information extraction process, extracting the string to be managed from the image data P1 and outputting the extraction result. Figure 4 shows an example of the procedure for processing a form image, including each of the above processes.

[0027] This disclosure can be understood as a method for performing one or more steps included in the document image processing described herein. Furthermore, one or more steps included in the document image processing described herein may be omitted as appropriate. In addition, the execution order of each step in the document image processing may differ to the extent that similar effects are produced. Furthermore, although this description uses the case in which the control unit 11 of the image processing apparatus 1 performs each step in the document image processing as an example, in other embodiments, one or more processors may distribute and execute each step in the document image processing. In addition, when the control unit 11 acquires image data P1 from each of the multiple operation terminals 2, it is possible to execute the document image processing in parallel for each image data P1.

[0028] In step S1, the control unit 11 (acquisition processing unit 111) determines whether or not it has acquired image data P1 of the document to be processed from the operation terminal 2. If the control unit 11 has acquired image data P1 from the operation terminal 2 (S1:Yes), it proceeds to step S2. The control unit 11 waits until it has acquired image data P1 from the operation terminal 2 (S1:No).

[0029] In step S2, the control unit 11 executes a document determination process. Specifically, the control unit 11 determines the format of the document based on the image data P1. For example, the control unit 11 determines whether the document corresponding to the image data P1 is a receipt-format document (hereinafter also referred to as a "receipt") or a general document format other than a receipt (hereinafter also referred to as a "general document"). A specific example of the document determination process will be described later.

[0030] If the document determination process determines that the document is a receipt (S3: Yes), in step S4, the control unit 11 temporarily sets the operating mode of the character recognition process to receipt mode. After step S4, the control unit 11 proceeds to step S5. Conversely, if the document determination process determines that the document is a general document (S3: No), in step S31, the control unit 11 temporarily sets the operating mode of the character recognition process (OCR process) to general document mode. After step S31, the control unit 11 proceeds to step S32.

[0031] In the aforementioned character recognition (OCR) process, the control unit 11 extracts characters one by one and performs a word formation process to group the characters into words (strings). When the control unit 11 extracts each string in the character recognition process, it associates information such as a number, string, and position coordinates with each extracted string and registers it in a string list (not shown). The position coordinates are, for example, represented by the starting point coordinates (top left coordinate) and ending point coordinates (bottom right coordinate) of the string rectangle.

[0032] In step S5, the control unit 11 performs character recognition processing (OCR processing for receipt mode) corresponding to the receipt mode. The OCR processing for receipt mode is a process that can recognize special characters on a receipt (such as double-width characters, double-width characters, characters with narrow line spacing, characters with narrow spacing, characters with wide spacing, etc.). After step S5, the control unit 11 moves the process to step S6.

[0033] In step S32, the control unit 11 performs character recognition processing (OCR processing for general form mode) corresponding to the general form mode. The OCR processing for general form mode performs pre-processing such as seal impression removal, background removal, character inversion, and italic correction, and then performs word formation processing. After step S32, the control unit 11 moves the processing to step S6.

[0034] In step S6, the control unit 11 executes an operation mode setting process. Specifically, the control unit 11 confirms the operation mode that was provisionally set in steps S4 and S31 to either receipt mode or general form mode. The control unit 11 then performs the following document classification extraction process, date extraction process, amount extraction process, and company information extraction process based on character recognition processing according to the confirmed operation mode. A specific example of the operation mode setting process will be described later.

[0035] Next, in step S7, the control unit 11 executes a document classification extraction process. Specifically, the control unit 11 extracts the classification of the document by performing a string rectangle extraction process using AI (machine learning) and a character recognition process (OCR processing) on ​​the string rectangle. A specific example of the document classification extraction process will be described later.

[0036] Next, in step S8, the control unit 11 performs a date extraction process. Specifically, the control unit 11 extracts dates such as the issue date included in the document by character recognition processing according to the operation mode. A specific example of the date extraction process will be described later.

[0037] Next, in step S9, the control unit 11 performs an amount extraction process. Specifically, the control unit 11 extracts the total amount included in the report by character recognition processing according to the operation mode. A specific example of the amount extraction process will be described later.

[0038] Next, in step S10, the control unit 11 executes a company information extraction process. Specifically, the control unit 11 extracts company information (issuer, recipient, registration number, etc.) contained in the form by character recognition processing according to the operation mode. A specific example of the company information extraction process will be described later.

[0039] Next, in step S11, the control unit 11 (output processing unit 115) outputs the extraction results of the document classification extraction process, the date extraction process, the amount extraction process, and the company information extraction process. For example, the control unit 11 outputs "Invoice" as the document classification, along with the invoice issuance date, invoice amount, and issuer (company name).

[0040] The following describes the specific configurations of each of the following processes included in the document image processing: the document determination process (S2), the operation mode setting process (S6), the document classification extraction process (S7), the date extraction process (S8), the amount extraction process (S9), and the company information extraction process (S10). The configuration relating to this disclosure is particularly characterized by the "document classification extraction process" among the above processes. Therefore, the following will mainly describe the "document classification extraction process," and the other processes will be described in the [Reference Form] section below.

[0041] [Document Classification Extraction Process] Figure 5 shows an example of the procedure for the document classification extraction process. In the document classification extraction process, the control unit 11 extracts the classification (type) of the document (form) based on the image data P1. For example, the control unit 11 identifies one of the following as the form classification of the image data P1: "invoice," "quotation," "delivery note," "order form," or "receipt." If none of these apply, it identifies "other document."

[0042] <Step S41> In step S41, the control unit 11 (extraction processing unit 112) extracts string rectangles from image data P1 and extracts the top three string rectangles in descending order of unit character area. Specifically, first, the extraction processing unit 112 extracts string rectangles from image data P1 using AI (machine learning). In this machine learning, training data (teacher data) is generated by attaching labels such as titles, strings, and tables to each string rectangle in the image data, and multiple such training data are used for training. The extraction processing unit 112 uses the trained model (AI) obtained by the machine learning to perform string rectangle extraction processing on the target image data P1. As a result, the extraction processing unit 112 extracts all string rectangles contained in the image data P1. Figure 6 shows the string rectangles extracted by the AI. Figure 7 shows the string rectangle list D1 for each string rectangle. The strings shown in Figure 7 are the recognition results of the OCR process. In another embodiment, the strings shown in Figure 7 may be the recognition results recognized by the AI.

[0043] Next, the extraction processing unit 112 extracts the top three string rectangles from the extracted string rectangles in descending order of their unit character area. The unit character area is the value obtained by dividing the area of ​​the string rectangle by the number of characters contained in the string rectangle. For example, for the string rectangle K1 shown in Figure 8, if the number of characters is 3, the number of pixels in the horizontal direction W1 is "565", and the number of pixels in the vertical direction H1 is "93", the total area of ​​string rectangle K1 is "52545". Dividing this by the number of characters "3", the unit character area of ​​string rectangle K1 is calculated to be "17515". The extraction processing unit 112 calculates the unit character area for each string rectangle and extracts the top three string rectangles with the largest areas. Note that in the string rectangle extraction process, strings with 3 or more characters or strings containing only numbers may be excluded from the extraction target.

[0044] <Step S42> Next, in step S42, the control unit 11 executes the first classification determination process (see Figure 9), which will be described later, to extract the classification of the document.

[0045] <Step S43> Next, in step S43, the control unit 11 determines whether the classification extracted by the first classification determination process has been finalized. If the classification extracted by the first classification determination process has been finalized (S43: Yes), the control unit 11 proceeds to step S49. If the classification extracted by the first classification determination process has not been finalized (S43: No), the control unit 11 proceeds to step S44.

[0046] <Step S44> Next, in step S44, the control unit 11 executes the second classification determination process (see Figure 11), which will be described later, to extract the classification of the document.

[0047] <Step S45> Next, in step S45, the control unit 11 determines whether the classification extracted by the second classification determination process has been finalized. If the classification extracted by the second classification determination process has been finalized (S45: Yes), the control unit 11 proceeds to step S49. If the classification extracted by the second classification determination process has not been finalized (S45: No), the control unit 11 proceeds to step S46.

[0048] <Step S46> Next, in step S46, the control unit 11 executes the third classification determination process (see Figure 13), which will be described later, to extract the classification of the document.

[0049] <Step S47> Next, in step S47, the control unit 11 determines whether the classification extracted by the third classification determination process has been finalized. If the classification extracted by the third classification determination process has been finalized (S47: Yes), the control unit 11 proceeds to step S49. If the classification extracted by the third classification determination process has not been finalized (S47: No), the control unit 11 proceeds to step S48.

[0050] <Step S48> Next, in step S48, if the classification is not determined in the third classification determination process, the control unit 11 determines the classification of the document to "other document".

[0051] <Step S49> Next, in step S49, the control unit 11 extracts the confirmed classifications, terminates the document classification extraction process, and moves the process to step S8 (see Figure 4).

[0052] [First classification determination process] Figure 9 shows an example of the procedure for the first classification determination process. The first classification determination process determines the classification of a document using the result of AI-based detection of the title rectangle.

[0053] <Step S421> In step S421, the extraction processing unit 112 determines whether or not there is a string rectangle labeled "Title" among the string rectangles extracted by AI from the image data P1 of the form to be judged. That is, the extraction processing unit 112 extracts string rectangles that appear to be titles from among all the string rectangles extracted by AI. If the extraction processing unit 112 determines that there is a title rectangle (S421: Yes), it proceeds to step S422; if it determines that there is no title rectangle (S421: No), it proceeds to step S44 (second classification determination process).

[0054] <Step S422> In step S422, the extraction processing unit 112 searches the image data P1 for the nearest neighbor string rectangle (first string rectangle) that is closest to the position of the title rectangle.

[0055] <Step S423> Next, in step S423, the recognition processing unit 113 performs character recognition processing (OCR processing) on ​​the nearest character rectangle to obtain the recognition result. The identification processing unit 114 then determines whether the recognition result matches a pre-set classification-related keyword. Figure 10 shows keyword information D2, which is pre-registered as a link between the classification of a form (document) and keywords related to that classification (classification-related keywords). The identification processing unit 114 determines which of the classification-related keywords in keyword information D2 the recognition result corresponds to.

[0056] For example, if the recognition processing unit 113 performs OCR processing on the string in the nearest string rectangle K1 and recognizes "invoice", the identification processing unit 114 refers to the keyword information D2 and searches for a category-related keyword that matches "invoice". In this case, the identification processing unit 114 determines that there is a category-related keyword that matches "invoice".

[0057] If the identification processing unit 114 determines that there is a classification-related keyword that matches the recognition result of the OCR processing (S423: Yes), it proceeds to step S424. On the other hand, if the identification processing unit 114 determines that there is no classification-related keyword that matches the recognition result (S423: No), it proceeds to step S44 (second classification determination processing).

[0058] <Step S424> Next, in step S424, the identification processing unit 114 determines whether the unit character area of ​​the nearest string rectangle is among the top three. That is, the identification processing unit 114 determines whether the nearest string rectangle is included in the top three string rectangles with large unit character areas extracted in step S41 (see Figure 5). If the identification processing unit 114 determines that the unit character area of ​​the nearest string rectangle is among the top three (S424: Yes), the process moves to step S425. On the other hand, if the identification processing unit 114 determines that the unit character area of ​​the nearest string rectangle is not among the top three (S424: No), the process moves to step S44 (second classification determination process).

[0059] <Step S425> Next, in step S425, the identification processing unit 114 determines the classification of the document. For example, if the recognition result of the nearest string rectangle is "Invoice", and "Invoice" is included in the classification-related keywords, the identification processing unit 114 determines the classification of the document to be "Invoice" if the unit character area of ​​the string rectangle is among the top three. Once the classification is determined, the control unit 11 moves the process to step S49 (see Figure 5) and extracts the determined classification.

[0060] Thus, in the first classification determination process, the identification processing unit 114 refers to the keyword information D2 (see Figure 10) and identifies the type of document based on the classification included in the recognition result. Specifically, the identification processing unit 114 identifies the classification identified based on the title rectangle as the document classification when the unit character area in the nearest string rectangle satisfies a predetermined condition (unit character area is within the top 3). If the control unit 11 cannot determine the document classification by the first classification determination process described above, it executes the following second classification determination process.

[0061] [Second Classification Determination Process] Figure 11 shows an example of the procedure for the second classification determination process. The second classification determination process focuses on the top three character rectangles with the largest unit character area and uses the recognition results of the OCR process to determine the classification of the document.

[0062] <Step S441> In step S441, the identification processing unit 114 determines whether the classification-related keyword exists in the string rectangle extracted by the extraction processing unit 112 using AI. If the identification processing unit 114 determines that the classification-related keyword exists in the string rectangle (S441: Yes), the process proceeds to step S442. On the other hand, if the identification processing unit 114 determines that the classification-related keyword does not exist in the string rectangle (S441: No), the process proceeds to step S46 (third classification determination process).

[0063] Specifically, the identification processing unit 114 determines which of the classification-related keywords in keyword information D2 (see Figure 10) corresponds to the recognition result of the OCR processing for each character rectangle. If the identification processing unit 114 determines that there is a classification-related keyword that matches the recognition result (S441: Yes), it proceeds to step S442. If it determines that there is no classification-related keyword that matches the recognition result (S441: No), it proceeds to step S46 (third classification determination processing). The identification processing unit 114 performs OCR processing on all character rectangles to determine the presence or absence of the classification-related keyword.

[0064] <Step S442> In step S442, the identification processing unit 114 determines whether the unit character area of ​​the string rectangle containing the classification-related keyword is among the top three. That is, the identification processing unit 114 determines whether the string rectangle is included in the top three string rectangles with large unit character areas extracted in step S41 (see Figure 5). If the identification processing unit 114 determines that the unit character area of ​​the string rectangle is among the top three (S442: Yes), the process moves to step S443. On the other hand, if the identification processing unit 114 determines that the unit character area of ​​the string rectangle is not among the top three (S442: No), the process moves to step S46 (third classification determination process). The identification processing unit 114 determines whether the unit character area of ​​all string rectangles containing the classification-related keyword is among the top three.

[0065] <Step S443> In step S443, the identification processing unit 114 adds the unit character area for each candidate category. Specifically, if the category-related keyword exists in the string rectangle, the category associated with the category-related keyword (see Figure 10) becomes the candidate category for extraction (candidate category). The identification processing unit 114 identifies candidate categories for each of the top three string rectangles in terms of unit character area and calculates the sum of the unit character areas for each candidate category. For example, in the example shown in Figure 12, if string rectangle K1 and string rectangle K2 contain the category-related keyword, are among the top three in terms of unit character area, and their respective candidate categories are "invoice", the identification processing unit 114 adds the unit character area of ​​string rectangle K1 and the unit character area of ​​string rectangle K2 to calculate the total area (cumulative character area) corresponding to the category "invoice". For example, if the string rectangle K3 contains the aforementioned category-related keyword, and its unit character area is among the top three, and the candidate category is "quotation," the identification processing unit 114 calculates the unit character area of ​​the string rectangle K3 as the total area corresponding to the category "quotation." In this way, the identification processing unit 114 calculates the total area (cumulative character area) by accumulating the unit character area for each candidate category for string rectangles whose unit character area is among the top three.

[0066] Furthermore, the specific processing unit 114 may exclude string rectangle K4 from the accumulation target if, for example, the number of characters is greater than or equal to a predetermined number (e.g., 15 characters), such as in string rectangle K4.

[0067] <Step S444> In step S444, the identification processing unit 114 determines whether there are multiple candidate categories. If there are multiple candidate categories (S444:Yes), the identification processing unit 114 proceeds to step S445, and if there is only one candidate category (S444:No), the processing proceeds to step S446. For example, if the candidate categories are "invoice" and "quotation," the identification processing unit 114 proceeds to step S445, and if the candidate category is only "invoice," the processing proceeds to step S446.

[0068] <Step S445> In step S445, the identification processing unit 114 determines the candidate category with the largest total unit character area (cumulative character area) as the document category. For example, if the cumulative character area of ​​"Invoice" is greater than the cumulative character area of ​​"Quotation", the identification processing unit 114 determines "Invoice" as the document category. Also, for example, if the cumulative character area of ​​"Quotation" is greater than the cumulative character area of ​​"Invoice", the identification processing unit 114 determines "Quotation" as the document category.

[0069] In another embodiment, if there are multiple candidate categories, the identification processing unit 114 may select the string rectangle whose position (vertical position, Y coordinate) corresponding to the candidate category is located at the top of the entire page, and confirm the candidate category of that string rectangle as the category of the form.

[0070] <Step S446> In step S446, the identification processing unit 114 determines one candidate category as the category for the document. For example, if the only candidate category is "Invoice", the identification processing unit 114 determines "Invoice" as the category for the document.

[0071] When the control unit 11 determines the classification, it moves the process to step S49 (see Figure 5) and extracts the determined classification. In this way, in the second classification determination process, if the extraction processing unit 112 cannot extract the title rectangle, the recognition processing unit 113 obtains multiple recognition results from the OCR processing of each string rectangle extracted by the AI ​​that satisfy predetermined conditions including the classification-related keyword, and the identification processing unit 114 identifies the classification of the document based on the unit character area in the string rectangle corresponding to each of the multiple recognition results. If the control unit 11 cannot determine the classification of the document by the second classification determination process described above, it executes the following third classification determination process.

[0072] [Third Category Determination Process] Figure 13 shows an example of the procedure for the third classification determination process. The third classification determination process focuses on all character rectangles and uses the recognition results of the OCR process to determine the classification of the document.

[0073] <Step S461> In step S461, the identification processing unit 114 determines whether the classification-related keyword exists in the string rectangle extracted by the extraction processing unit 112 using AI. If the identification processing unit 114 determines that the classification-related keyword exists in the string rectangle (S461: Yes), the process proceeds to step S462. On the other hand, if the identification processing unit 114 determines that the classification-related keyword does not exist in the string rectangle (S461: No), the process proceeds to step S48. The identification processing unit 114 performs OCR processing on all string rectangles to determine the presence or absence of the classification-related keyword.

[0074] <Step S462> In step S462, the identification processing unit 114 adds the unit character area for each candidate category. The identification processing unit 114 identifies candidate categories for each of the string rectangles containing the category-related keywords and calculates the total unit character area (cumulative character area) for each candidate category. For example, if the image data P1 contains the category-related keywords "invoice", "quotation", and "delivery note", the identification processing unit 114 calculates the cumulative character area for all string rectangles containing the category-related keyword "invoice", the cumulative character area for all string rectangles containing the category-related keyword "quotation", and the cumulative character area for all string rectangles containing the category-related keyword "delivery note".

[0075] <Step S463> In step S463, the identification processing unit 114 determines the candidate category with the largest total unit character area (cumulative character area) as the document category. For example, if the cumulative character area of ​​"Invoice" is the largest, the identification processing unit 114 determines "Invoice" as the document category; if the cumulative character area of ​​"Quotation" is the largest, it determines "Quotation" as the document category; and if the cumulative character area of ​​"Delivery Note" is the largest, it determines "Delivery Note" as the document category.

[0076] <Step S48> In step S48, the specific processing unit 114 determines "other documents" as the form category. That is, if the category cannot be determined in the first category determination process, the second category determination process, or the third category determination process, the specific processing unit 114 determines "other documents" as the form category.

[0077] Thus, in the third classification determination process, if the unit character area in the nearest string rectangle does not satisfy the predetermined conditions, the identification processing unit 114 obtains all recognition results from the OCR processing for each string rectangle extracted by the AI ​​that include the classification-related keywords, and identifies the classification of the document based on the unit character area in the string rectangle corresponding to each of the recognition results. Once the classification of the document is determined, the control unit 11 moves the process to step S49 (see Figure 5) and extracts the determined classification. For example, the control unit 11 extracts one of the following as the classification of the document corresponding to the image data P1: "Invoice", "Quotation", "Delivery Note", "Purchase Order", "Receipt", or "Other Document". In this way, the control unit 11 executes the document classification extraction process to extract the classification (type) of the document.

[0078] As described above, in the document classification extraction process according to this embodiment, the image processing system 10 extracts a title rectangle containing the title of the document from the image data P1 of the document using AI (machine learning), performs character recognition processing (OCR processing) on ​​the nearest neighbor string rectangle (first string rectangle) closest to the position of the title rectangle from the image data P1 to obtain the recognition result, and refers to a storage unit (keyword information D2) that stores in advance an association between the document classification (type) and type information (classification-related keywords) for each classification, and identifies the classification of the document based on the type information included in the recognition result. Furthermore, if the unit character area in the nearest neighbor string rectangle satisfies a predetermined condition (the area size is within the top 3), the image processing system 10 identifies the classification identified based on the title rectangle as the type of the document (the "first classification determination process").

[0079] Furthermore, if the title rectangle cannot be extracted by the first classification determination process, the image processing system 10 obtains multiple recognition results (the top three string rectangles) that satisfy predetermined conditions including classification-related keywords from the recognition results of the OCR processing for each string rectangle extracted by the AI, and identifies the type of document based on the unit character area of ​​these results. For example, the image processing system 10 calculates the unit character area for each of the top three string rectangles, calculates the sum of the unit character areas for each candidate classification, and identifies the candidate classification with the largest total area among the total areas for each candidate classification as the classification of the document (the "second classification determination process").

[0080] Furthermore, if the unit character area in a string rectangle containing classification-related keywords does not satisfy the predetermined conditions, the image processing system 10 acquires all recognition results containing the classification-related keywords from the OCR processing recognition results for each string rectangle extracted by the AI, and identifies the type of document based on the unit character area in the string rectangle corresponding to each of the recognition results. For example, the image processing system 10 calculates the unit character area for each of the recognition results, calculates the total unit character area for each candidate classification, and identifies the candidate classification with the largest total area among the total areas for each candidate classification as the type of document. Also, if the classification-related keywords are not included in the recognition results, the image processing system 10 identifies the type of document as a predetermined type ("other document") ("third classification determination process").

[0081] Thus, the image processing system 10 is configured to identify the classification of a document through a three-stage classification determination process (first classification determination process, second classification determination process, and third classification determination process).

[0082] According to the above configuration, for example, even if the form does not have a title located at the top, or if the font size of the text strings does not meet the predetermined size, the form category (type) can be identified. Furthermore, even if the title of a form cannot be recognized, the category can be identified by using the OCR results of the text rectangle other than the title through the second or third category determination process. Therefore, the form category can be appropriately identified.

[0083] [Other embodiments] The image processing system 10 relating to this disclosure is not limited to the embodiments described above, but may also be the following embodiments.

[0084] In the image data P1 shown in Figure 14, for example, if the character rectangle K1 ("···") is the first in unit character area, the character rectangle K2 ("···") is the second in unit character area, the character rectangle K3 ("Invoice") is the third in unit character area, and the character rectangle K4 ("Delivery Note") is the fourth in unit character area, then in the second classification determination process, the control unit 11 will target the top three character rectangles K1 to K3 for determination. In this case, the fourth-ranked "Delivery Note," which has the potential to be classified, will be excluded from the determination.

[0085] Therefore, the control unit 11 executes a process to expand the unit character area (expansion process) if the recognition result of the OCR process exactly matches any of the strings "invoice", "delivery note", "receipt", "order form", or "quotation". Specifically, the control unit 11 corrects (expands) the unit character area by multiplying it by a coefficient (for example, "1.05"). For example, as shown in Figure 14, the control unit 11 expands the unit character area of ​​the third-place "invoice" and the fourth-place "delivery note". Then, the control unit 11 sets the unit character area of ​​the third-place "invoice" before expansion ("10063") as the base size, and determines string rectangles whose expanded unit character area is equal to or greater than the base size ("10063") as targets for determination. As a result, the fourth-place "delivery note" is added to the targets for determination.

[0086] Thus, the control unit 11 (specific processing unit 114) may, when the character corresponding to the recognition result matches a character of a category registered in keyword information D2 (see Figure 10), expand the size (unit character area) of the string rectangle corresponding to the recognition result, and determine whether the unit character area satisfies a predetermined condition (greater than or equal to the standard size) based on the expanded string rectangle.

[0087] In the example above, "Invoice" and "Delivery Note" are candidate categories, so the control unit 11 confirms "Delivery Note," which is located at the top of the entire page, as the document category. With this configuration, when a category-related keyword is three characters long, such as "○○ document," those characters can be treated as the most important keyword.

[0088] In another embodiment of this disclosure, the image processing system 10 may use the title recognition result by AI to identify the document category. Specifically, when the AI ​​extracts a title rectangle and extracts the title (category) of the title rectangle, the control unit 11 compares the category extracted by the AI ​​with the recognition result obtained by OCR processing of the nearest character string rectangle to the title rectangle. If the categories match, the control unit 11 confirms that category as the document category. Figure 15 shows an example of the procedure for this configuration.

[0089] <Step S51> In step S51, the control unit 11 determines whether or not there is a string rectangle labeled "Title" among the string rectangles extracted by AI from the image data P1. If the control unit 11 determines that there is a title rectangle (S51: Yes), it proceeds to step S52; if it determines that there is no title rectangle (S51: No), it proceeds to step S57.

[0090] <Step S52> In step S52, the control unit 11 searches the image data P1 for the nearest neighbor string rectangle closest to the position of the title rectangle and extracts the recognition result of the OCR processing of that nearest neighbor string rectangle.

[0091] <Step S53> In step S53, the control unit 11 determines whether the recognition result exists in a pre-set classification-related keyword. If the control unit 11 determines that there is a classification-related keyword that matches the recognition result (S53: Yes), the process moves to step S54. On the other hand, if the identification processing unit 114 determines that there is no classification-related keyword that matches the recognition result (S53: No), the process moves to step S57.

[0092] <Step S54> In step S54, the control unit 11 determines whether the division of the title rectangle extracted by the AI ​​matches the division recognized by the OCR process. If the control unit 11 determines that the division extracted by the AI ​​matches the division recognized by the OCR process (S54: Yes), it proceeds to step S55. On the other hand, if the identification processing unit 114 determines that the division extracted by the AI ​​does not match the division recognized by the OCR process (S54: No), it proceeds to step S56.

[0093] <Step S55> In step S55, the control unit 11 confirms the matching classification as the classification for the form. Once the classification is confirmed, the control unit 11 moves the process to step S49 (see Figure 5) and extracts the confirmed classification.

[0094] <Step S56> In step S56, the control unit 11 determines whether the accuracy (evaluation value, score, etc.) of the categories recognized by OCR processing is greater than the accuracy of the categories extracted by AI. If the accuracy of the categories recognized by OCR processing is greater than the accuracy of the categories extracted by AI (S56: Yes), the control unit 11 proceeds to step S57.

[0095] <Step S57> In step S57, the control unit 11 determines whether or not the classification-related keyword exists in the string rectangle extracted by the AI. If the identification processing unit 114 determines that the classification-related keyword exists in the string rectangle (S57: Yes), it proceeds to step S58. On the other hand, if the identification processing unit 114 determines that the classification-related keyword does not exist in the string rectangle (S57: No), it proceeds to step S59.

[0096] The control unit 11 performs OCR processing on all character rectangles to determine whether or not the classification-related keywords are present.

[0097] <Step S58> In step S58, the control unit 11 adds up the unit character area for each candidate category. The control unit 11 identifies the candidate category for each of the string rectangles containing the category-related keywords and calculates the total unit character area (cumulative character area) for each candidate category.

[0098] <Step S59> In step S59, the control unit 11 determines "other documents" as the category of the form.

[0099] <Step S60> In step S60, the control unit 11 determines whether there are multiple candidate categories. If there are multiple candidate categories (S60:Yes), the identification processing unit 114 proceeds to step S61, and if there is only one candidate category (S60:No), it proceeds to step S62.

[0100] <Step S61> In step S61, the identification processing unit 114 determines the candidate character rectangle whose rectangular coordinate (Y coordinate) is the highest level as the form category. In another embodiment, the control unit 11 may determine the candidate category with the largest total area among the cumulative character areas for each candidate category as the form category.

[0101] <Step S62> In step S62, the control unit 11 confirms one candidate category as the category for the form.

[0102] Once the control unit 11 determines the classification (steps S59, S61, S62), it moves the process to step S49 (see Figure 5) and extracts the determined classification.

[0103] As described above, the control unit 11 may use the AI ​​title recognition result and the OCR processing recognition result to identify the document classification.

[0104] [Reference form] The outlines of the document determination process (step S2 in Figure 4), the operation mode setting process (step S6 in Figure 4), the date extraction process (step S8 in Figure 4), the amount extraction process (step S9 in Figure 4), and the company information extraction process (step S10 in Figure 4) will be described below. In this disclosure, the document determination process, the operation mode setting process, the date extraction process, the amount extraction process, and the company information extraction process are not limited to the following configurations, and well-known technologies may be applied.

[0105] [Document Recognition Process] In the document determination process, the control unit 11 determines the format of the document based on the image data P1. For example, the control unit 11 determines whether the document (document) in the image data P1 is a receipt or a general document.

[0106] Specifically, the control unit 11 obtains a text rectangle surrounding the text contained in the document (form) based on the image data P1 of the document, and determines whether the document is a receipt or a general form based on the arrangement density of the text rectangle in an area corresponding to the size of the document. Specifically, the control unit 11 determines whether the document is a receipt or a general form based on the arrangement density (height ratio) of the text rectangle in the vertical direction of the area.

[0107] Furthermore, the control unit 11 may calculate the arrangement density as the sum of the vertical heights of each of the multiple string rectangles relative to the vertical reference height of the area, and if the arrangement density is equal to or greater than a threshold, it may determine the document to be a receipt, and if the arrangement density is less than a threshold, it may determine the document to be a general form.

[0108] Furthermore, the control unit 11 may calculate the total height obtained by summing the vertical heights of each of the multiple character rectangles relative to the vertical reference height of the area as the arrangement density, and if the arrangement density is less than a threshold, the document may be determined to be a general form.

[0109] Furthermore, the control unit 11 may determine the document to be a general form if the arrangement density is above a threshold, the aspect ratio of the area matches a predetermined value, and the estimated number of characters in the horizontal direction of the area is above a predetermined number of characters. If the arrangement density is above a threshold and the aspect ratio of the area does not match the predetermined value, or if the arrangement density is above a threshold and the estimated number of characters is less than the predetermined number of characters, the control unit 11 may determine the document to be a receipt.

[0110] Furthermore, the control unit 11 may estimate the vertical and horizontal size (image size) of the region based on the maximum and minimum horizontal position coordinates and the maximum and minimum vertical position coordinates of each of the multiple string rectangles. For example, the control unit 11 estimates the image size by using the minimum horizontal value as the left vertical line of the crop, the maximum horizontal value as the right vertical line of the crop, the minimum vertical value as the top line of the crop, and the maximum vertical value as the bottom line of the crop for each of the multiple string rectangles. Alternatively, the control unit 11 may determine the region based on the difference between the size of the image data P1 (input image size) and the estimated image size estimated from the arrangement positions of the multiple string rectangles.

[0111] Furthermore, the control unit 11 may calculate a reference height by subtracting the height of a non-string rectangle (such as a one-dimensional code, two-dimensional code, or illustration) placed between the first and second string rectangles from the length from the first string rectangle Ps located at the uppermost position of the area to the second string rectangle Pe located at the lowermost position of the area.

[0112] Furthermore, the control unit 11 may calculate the arrangement density corresponding to a plurality of character rectangles included in a divided region obtained by dividing the region horizontally (left and right) by vertical lines.

[0113] According to the above configuration, it is possible to appropriately determine whether the input image data P1 is a receipt or a general form.

[0114] [Operation mode setting process] In the operation mode setting process, the control unit 11 determines whether the document (form) is a receipt or a general form based on the image data P1 of the document, and extracts a category from among the multiple items contained in the document that represents the type of document. The control unit 11 also sets the operation mode for extracting target characters contained in the multiple items to receipt mode (first operation mode) or general form mode (second operation mode) based on the determination result and the extracted category.

[0115] Furthermore, if the control unit 11 determines that the document is a general document during the document determination process and extracts a quotation, order form, delivery note, invoice, or receipt as the category, it sets the operation mode to general document mode. Also, if the control unit 11 determines that the document is a receipt during the document determination process and extracts a quotation, order form, delivery note, or invoice as the category, it sets the operation mode to general document mode.

[0116] Furthermore, if the control unit 11 determines that the document is a receipt in the document determination process and extracts "receipt" or "other document" (other document) as the category, it sets the operation mode to receipt mode.

[0117] Furthermore, if the control unit 11 determines that the document is a general form, it extracts the classification from the image data P1 on which the general form mode OCR processing (second character recognition processing) corresponding to the general form mode has been performed. If the control unit 11 determines that the document is a receipt, it extracts the classification from the image data P1 on which the receipt mode OCR processing (first character recognition processing) corresponding to the receipt mode has been performed.

[0118] According to the above configuration, even if the type of document is incorrectly identified, for example, the correct target characters can be extracted. For example, even if the control unit 11 incorrectly identifies a general form document as a receipt and sets it to receipt mode, if the extracted category is a quotation, order form, delivery note, or invoice, it will change the receipt mode to general form mode. In other words, even if the control unit 11 incorrectly identifies a general form document as a receipt, it can use the category extraction result to change (recover) it to a general form. As a result, the control unit 11 can extract the target characters as a general form in subsequent processing.

[0119] When the control unit 11 sets an operating mode, it extracts strings of the target (managed object), such as dates, amounts, company information, etc., according to that operating mode.

[0120] In other words, the control unit 11 automatically determines whether the input image data P1 is a general form or a receipt. If it is determined to be a general form, it performs OCR processing using an OCR engine for general forms (OCR for general form mode). Using the OCR results, it performs a classification extraction process to determine which category of general form it belongs to. If the classification is determined to be a quotation, order form, delivery note, invoice, or receipt, it performs a process to extract items for general forms (total amount, date, issuer, recipient, etc.). If it is determined to be any other document, it does not perform the item extraction process. In response, if the control unit 11 determines that the document is a receipt, it performs OCR processing using an OCR engine for receipts (OCR for receipt mode), and uses the OCR results to determine which category of general documents it belongs to in the category extraction process. If the category is determined to be a quotation, order form, delivery note, or invoice, it updates the document identification result to a general document (recovery process) and extracts items for general documents (total amount, date, issuer, recipient, etc.) according to the updated document identification result and the determined category. On the other hand, if the control unit 11 determines that the category is a receipt or other document, it updates the category of other documents to a receipt (recovery process) and extracts items for receipts (total amount, date, issuer, etc.) according to the document identification result and the updated category.

[0121] Furthermore, regardless of the document identification result, if the control unit 11 determines the category to be "other document," it updates the document identification result to "receipt" as necessary (recovery process), updates the category of "other document" to "receipt" (recovery process), and extracts items for receipts (total amount, date, issuer, etc.).

[0122] Furthermore, if the control unit 11 determines the category to be "other document," it updates the category of "other document" to "receipt" (recovery processing). The control unit 11 also processes to extract items for general forms (total amount, date, issuer, recipient, etc.) for the categories of quotation, order, delivery note, invoice, and receipt. The control unit 11 also processes to extract items for receipts (total amount, date, issuer, etc.) for the categories of quotation, order, delivery note, invoice, and receipt. Finally, based on the document discrimination result, the category extraction result, and the results of each extraction process, the control unit 11 updates the document discrimination result (recovery processing), and selects and outputs output items according to the updated document discrimination result and category.

[0123] [Date extraction process] In the date extraction process, the control unit 11 obtains the results of character recognition (OCR) processing performed on the character strings contained in the image of the document (form), calculates an evaluation value for each of the multiple date string candidates included in the results of the character recognition processing according to the format of the document (receipt, general form), and extracts a date string based on the calculated evaluation value for each of the multiple date string candidates.

[0124] Furthermore, the control unit 11 sets the evaluation value for the date string candidate according to the distance from the center position P0 of the image to the position of the date string candidate. For example, if the document is a receipt, the control unit 11 sets a higher evaluation value for the date string candidate the shorter the distance. Also, for example, if the document is a general form, the control unit 11 sets a higher evaluation value for the date string candidate the longer the distance.

[0125] Furthermore, the control unit 11 may add a predetermined value to the evaluation value of the date string candidate if the image contains related characters related to the date within a predetermined range from the position of the date string candidate. Also, the control unit 11 may subtract a predetermined value from the evaluation value of the date string candidate if the image contains unrelated characters (NG words) that are not related to the date within a predetermined range from the position of the date string candidate.

[0126] According to the above configuration, the dates of documents (receipts, general ledger forms) with different formats can be appropriately extracted.

[0127] [Amount extraction process] In the amount extraction process, the control unit 11 extracts a character string of a specific amount (total amount) corresponding to a target item name ("total", "amount", etc.) to be extracted from the image data P1 of the document (ledger form). Specifically, based on the image data P1, the control unit 11 determines whether the ledger form is a receipt or a general ledger form, extracts the item name related to the target item name from the image data P1, and within the search range (the same line, one line below, etc.) corresponding to the format (receipt or general ledger form) of the ledger form based on the item name, extracts the amount corresponding to the item name, and outputs the specific amount based on the extracted amount.

[0128] Also, when the ledger form is a receipt, the control unit 11 extracts the amount within the search range of the same line as the first line where the item name is arranged or the search range of the second line one line below the first line.

[0129] On the other hand, when the ledger form is a general ledger form, the control unit 11 extracts the amount within the search range on the right side of the first position where the item name is arranged or the search range below the first position.

[0130] Further, when the control unit 11 extracts the item name related to the total from the image data P1 and extracts a character string including the numbers included in the search range and specific characters ("¥", "yen", etc.) representing the amount, the control unit 11 may output the number as the specific amount.

[0131] Also, when the control unit 11 does not extract the item name related to the total from the image data P1, extracts the item name related to the amount, and extracts a plurality of character strings including the numbers and the specific characters included in the search range, the control unit 11 outputs a calculated value calculated based on the plurality of numbers as the specific amount.

[0132] Furthermore, if the control unit 11 does not extract the item names related to the total and amount from the image data P1, and extracts multiple strings of numbers from the image data P1, it may output the number identified from among the multiple numbers based on the numbers related to consumption tax and change as the specified amount.

[0133] Furthermore, the control unit 11 extracts keywords related to total amounts, such as "Total" and "Current Total," from the recognized string information of the image data P1 input as, for example, a receipt image. It then extracts numerical strings that exist to the right of the keyword on the same line or to the right of the line below the keyword as candidate amounts corresponding to that keyword. When extracting keywords, the control unit 11 searches for keywords using the recognition results of both a non-learning OCR, which is an OCR for receipt mode, and a general-purpose AI-learning OCR. On the other hand, when extracting candidate amounts, the control unit 11 preferentially uses the recognition results of the general-purpose AI-learning OCR to search for numerical strings with monetary symbols (¥, yen, etc.) and extract candidate amounts. If the control unit 11 does not find keywords related to total amounts, it searches for keywords related to amounts other than total amounts and extracts candidate amounts. When the control unit 11 extracts candidate amounts, it assigns a higher score to the candidate amounts extracted using the total amount keyword than to the candidate amounts extracted using keywords other than total amounts. Furthermore, the control unit 11 adds scores based on indicators such as whether they are on the same line, whether they have an amount symbol, or whether they are the largest numbers. Finally, the control unit 11 outputs the amount candidate with the highest score as the amount to be extracted (total amount).

[0134] Furthermore, if no numerical string with an amount symbol exists, the control unit 11 searches for numerical strings located as far apart as possible on the same line to extract candidate amounts.

[0135] Furthermore, if the control unit 11 does not find a keyword, it extracts monetary candidates without a keyword. When extracting monetary candidates, the control unit 11 prioritizes using the recognition results of the AI-learning OCR to search for strings of numbers with monetary symbols and extracts the amounts. If the control unit 11 does not find strings of numbers with monetary symbols, it also extracts strings of numbers without monetary symbols. For example, if the control unit 11 does not find a keyword, it may extract all strings containing numbers as monetary candidates, and if a string of numbers with monetary symbols is found, it may cancel all strings of numbers without monetary symbols. Also, if the control unit 11 extracts only strings without monetary symbols, it may cancel strings of numbers that contain characters other than numbers and commas. Furthermore, the control unit 11 may cancel strings of numbers extracted using keywords such as "cash" or "deposit". Also, if the control unit 11 finds a string of numbers that is the same as an amount calculated using the keyword and keywords such as "consumption tax" or "change", it may add a score to make it a monetary candidate.

[0136] With the above configuration, the total amount to be extracted can be accurately extracted for each general report and receipt. Furthermore, even if the receipt layout is different, the total amount to be extracted can be accurately extracted.

[0137] [Company Information Extraction Process] In the company information extraction process described above, the control unit 11 extracts company information such as the issuer, recipient, and registration number contained in the document. Specifically, if the document is a receipt, the control unit 11 extracts the company information using OCR processing for receipt mode, and if the document is a general document, it extracts the company information using OCR processing for general document mode.

[0138] For example, the control unit 11 extracts strings related to company information such as "store name," "corporate name," "person's name," and "registration number" through OCR processing according to the type of document, and determines the company information to be extracted based on an evaluation value corresponding to the position of the string in the image data P1. After step S10, the control unit 11 moves the process to step S11.

[0139] [Extraction result output processing] After performing the document classification extraction process, date extraction process, amount extraction process, and company information extraction process described above, the control unit 11 outputs the extraction results in step S10 (see Figure 4). Here, the control unit 11 (output processing unit 115) outputs the date (issue date, etc.), amount (total amount, etc.), and company information (issuer, etc.) for general documents (quotation, purchase order, invoice, delivery note, receipt), and the date (issue date, etc.), amount (total amount, etc.), and company information (store name, etc.) for receipts.

[0140] In the embodiments described above, the image processing device 1 alone corresponds to the information processing system according to this disclosure. However, the information processing system according to this disclosure may consist of the image processing device 1 and the operation terminal 2. For example, when the components of the image processing device 1 and the operation terminal 2 cooperate to perform the aforementioned processing, a system including multiple components that perform the processing corresponds to the information processing system according to this disclosure. Furthermore, when the operation terminal 2 performs the processing, the operation terminal 2 alone may constitute the information processing system according to this disclosure.

[0141] The control unit 11 of the image processing device 1 controls the entire image processing device 1. The control unit 11 realizes various functions by reading and executing various programs stored in the storage unit 12 (for example, storage or ROM). The control unit 11 may be realized by one or more control devices / arithmetic units (CPU (Central Processing Unit), SoC (System on a Chip)). The control unit 11 may also be composed of one or more control circuits (electronic circuits).

[0142] [Note on disclosure] The following is an overview of the disclosures extracted from the above-described embodiments. Note that each configuration and processing function described in the following notes can be selected and combined as desired.

[0143] <Note 1> An extraction processing unit that uses machine learning to extract a title rectangle containing the title of the document from the image data of the document, A recognition processing unit that performs character recognition processing on the first string rectangle closest to the position of the title rectangle from the image data and obtains the recognition result, A specific processing unit identifies the type of document based on the type information included in the recognition result, by referring to a storage unit that stores in advance the association of document types and type information related to those types for each type, An information processing system equipped with the following features.

[0144] <Note 2> The identification processing unit identifies the type identified based on the title rectangle as the type of document when the unit character area in the first string rectangle satisfies predetermined conditions. The information processing system described in Appendix 1.

[0145] <Note 3> If the extraction processing unit is unable to extract the title rectangle, The recognition processing unit acquires a plurality of first recognition results that satisfy predetermined conditions, including the type information, from among the recognition results of the character recognition process for each string rectangle extracted by the machine learning, The identification processing unit identifies the type of document based on the unit character area in the string rectangle corresponding to each of the plurality of first recognition results. The information processing system described in Appendix 1 or 2.

[0146] <Note 4> The identification processing unit calculates the unit character area for each of the plurality of first recognition results, calculates the total unit character area for each candidate type, and identifies the candidate with the largest total area among the total areas for each candidate type as the type of document. The information processing system described in Appendix 3.

[0147] <Note 5> If the unit character area in the first string rectangle does not satisfy the predetermined conditions, The specified processing unit acquires all of the first recognition results, including the type information, from the recognition results of the character recognition process for each string rectangle extracted by the machine learning, The identification processing unit identifies the type of document based on the unit character area in the string rectangle corresponding to each of the first recognition results. The information processing system described in Appendix 3 or 4.

[0148] <Note 6> The identification processing unit calculates the unit character area for each of the first recognition results, calculates the total unit character area for each candidate type, and identifies the candidate with the largest total area among the total areas for each candidate type as the type of document. The information processing system described in Appendix 5.

[0149] <Note 7> The identification processing unit identifies the type of the document as a predetermined type if the type information is not included in the recognition result. An information processing system as described in any of the appendices 1 to 6.

[0150] <Note 8> The identification processing unit expands the size of the first string rectangle corresponding to the recognition result when the character corresponding to the recognition result matches a character of a type stored in the storage unit, and determines whether the unit character area satisfies the predetermined conditions based on the expanded first string rectangle. The information processing system described in Appendix 2.

[0151] <Note 9> Extracting a title rectangle containing the document's title from the document's image data using machine learning, The process involves performing character recognition on the first character rectangle closest to the position of the title rectangle from the image data and obtaining the recognition result. The process involves referring to a storage unit that stores, in advance, a relationship between the type of document and type information related to that type for each type, and identifying the type of document based on the type information included in the recognition result. An information processing method performed by one or more processors.

[0152] <Note 10> Extracting a title rectangle containing the document's title from the document's image data using machine learning, The process involves performing character recognition on the first character rectangle closest to the position of the title rectangle from the image data and obtaining the recognition result. The process involves referring to a storage unit that stores, in advance, a relationship between the type of document and type information related to that type for each type, and identifying the type of document based on the type information included in the recognition result. An information processing program for causing one or more processors to execute, or a non-temporary computer-readable recording medium on which such information processing program is recorded. [Explanation of symbols]

[0153] 1: Image processing device 2: Operating terminal 10: Image Processing System 11: Control Unit 12: Storage section 13: Operation display section 14: Communications Department 111: Acquisition Processing Unit 112: Determination Processing Unit 113: Extraction Processing Unit 114: Configuration Processing Unit 115: Calculation Processing Unit 116: Output Processing Unit

Claims

1. An extraction processing unit that uses machine learning to extract a title rectangle containing the title of the document from the image data of the document, A recognition processing unit that performs character recognition processing on the first string rectangle closest to the position of the title rectangle from the image data and obtains the recognition result, A specific processing unit identifies the type of document based on the type information included in the recognition result, by referring to a storage unit that stores in advance the association of document types and type information related to those types for each type, An information processing system equipped with the following features.

2. The identification processing unit identifies the type identified based on the title rectangle as the type of document when the unit character area in the first string rectangle satisfies predetermined conditions. The information processing system according to claim 1.

3. If the extraction processing unit is unable to extract the title rectangle, The recognition processing unit acquires a plurality of first recognition results that satisfy predetermined conditions, including the type information, from among the recognition results of the character recognition process for each string rectangle extracted by the machine learning, The identification processing unit identifies the type of document based on the unit character area in the string rectangle corresponding to each of the plurality of first recognition results. The information processing system according to claim 1.

4. The identification processing unit calculates the unit character area for each of the plurality of first recognition results, calculates the total unit character area for each candidate type, and identifies the candidate with the largest total area among the total areas for each candidate type as the type of document. The information processing system according to claim 3.

5. If the unit character area in the first string rectangle does not satisfy the predetermined conditions, The identification processing unit acquires all of the first recognition results, including the type information, from the recognition results of the character recognition process for each string rectangle extracted by the machine learning, and identifies the type of document based on the unit character area in the string rectangle corresponding to each of the first recognition results. The information processing system according to claim 3.

6. The identification processing unit calculates the unit character area for each of the first recognition results, calculates the total unit character area for each candidate type, and identifies the candidate with the largest total area among the total areas for each candidate type as the type of document. The information processing system according to claim 5.

7. The identification processing unit identifies the type of the document as a predetermined type if the type information is not included in the recognition result. The information processing system according to claim 1.

8. The identification processing unit expands the size of the first string rectangle corresponding to the recognition result when the character corresponding to the recognition result matches a character of a type stored in the storage unit, and determines whether the unit character area satisfies the predetermined conditions based on the expanded first string rectangle. The information processing system according to claim 2.

9. Extracting a title rectangle containing the document's title from the document's image data using machine learning, The process involves performing character recognition on the first character rectangle closest to the position of the title rectangle from the image data and obtaining the recognition result. The process involves referring to a storage unit that stores, in advance, a relationship between the type of document and type information related to that type for each type, and identifying the type of document based on the type information included in the recognition result. An information processing method performed by one or more processors.

10. Extracting a title rectangle containing the document's title from the document's image data using machine learning, The process involves performing character recognition on the first character rectangle closest to the position of the title rectangle from the image data and obtaining the recognition result. The process involves referring to a storage unit that stores, in advance, a relationship between the type of document and type information related to that type for each type, and identifying the type of document based on the type information included in the recognition result. An information processing program that causes one or more processors to execute.

Citation Information

Patent Citations

  • Information processing device and program

    JP2021056722A