Information processing system, information processing method, and information processing program
The system enhances character recognition accuracy by determining document formats based on character string rectangle placement density, effectively extracting relevant information from diverse document layouts.
Patent Information
- Application Number
- JP2024011971
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-12
AI Technical Summary
Conventional character recognition technologies struggle with documents having a relatively free layout, such as receipts, leading to reduced accuracy in distinguishing document formats and recognizing characters.
An information processing system that acquires character string rectangles from document images, determines the document format based on the placement density of these rectangles, and extracts target characters using an appropriate operating mode.
Improves character recognition accuracy by appropriately distinguishing between documents of different formats, enhancing the extraction of specific items like dates, amounts, and company information.
Smart Images

Figure 2025117234000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing program that perform image processing such as character recognition on image data. [Background technology]
[0002] Conventionally, there are known techniques for extracting character strings from image data of documents such as forms. For example, there is known a technique for determining the document type of the image data and performing character recognition processing on the image data using a character recognition engine associated with the determined document type (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2020-181369 Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventional technology determines the type of document based on characteristics such as whether the document has been treated with measures to improve recognition accuracy, and whether it has an identifier that can identify the document, such as a document ID.Therefore, in the case of documents with a relatively free layout, such as receipts, the type of document cannot be properly determined, resulting in a problem of reduced character recognition accuracy.
[0005] An object of the present disclosure is to provide an information processing system, an information processing method, and an information processing program that are capable of appropriately distinguishing between documents of different formats and improving character recognition accuracy. [Means for solving the problem]
[0006] An information processing system according to one aspect of the present disclosure includes an acquisition processing unit that acquires character string rectangles surrounding character strings contained in a document based on image data of the document, a determination processing unit that determines the format of the document based on the placement density of the character string rectangles in an area corresponding to the size of the document, and an extraction processing unit that extracts target characters corresponding to specified items contained in the document from the image data using an operating mode according to the determination result of the determination processing unit.
[0007] An information processing method according to another aspect of the present disclosure is an information processing method executed by one or more processors to obtain character string rectangles surrounding character strings contained in a document based on image data of the document, determine the format of the document based on the placement density of the character string rectangles in an area corresponding to the size of the document, and extract target characters corresponding to specified items contained in the document from the image data using an operating mode according to the determination result.
[0008] An information processing program according to another aspect of the present disclosure is an information processing program for causing one or more processors to execute the following steps: obtain character string rectangles surrounding character strings contained in a document based on image data of the document; determine the format of the document based on the placement density of the character string rectangles in an area corresponding to the size of the document; and extract target characters corresponding to specified items contained in the document from the image data using an operating mode according to the determination result. [Effects of the Invention]
[0009] According to the present disclosure, it is possible to provide an information processing system, an information processing method, and an information processing program that are capable of appropriately distinguishing between documents of different formats and improving character recognition accuracy. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a functional block diagram showing the configuration of an image processing system according to an embodiment of the present disclosure. [Figure 2]FIG. 2 is a diagram illustrating an example of an estimate included in a form according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a diagram illustrating an example of a receipt included in a form according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a flowchart illustrating an example of a procedure for form image processing executed by the image processing system according to the embodiment of the present disclosure. [Figure 5A] FIG. 5A is a diagram showing an example of a character string list in which character strings extracted in character recognition processing of the image processing system according to an embodiment of the present disclosure are registered. [Figure 5B] FIG. 5B is a diagram illustrating an example of character recognition processing on image data of a general form according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram showing an example of character recognition processing on image data of a receipt according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a flowchart illustrating an example of the procedure of the form determination process executed in the image processing system according to the embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating a specific example of the form determination process according to the embodiment of the present disclosure. [Figure 9] FIG. 9 is a diagram illustrating a specific example of the form determination process according to the embodiment of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating a specific example of the form determination process according to the embodiment of the present disclosure. [Figure 11] FIG. 11 is a diagram illustrating a specific example of the form determination process according to the embodiment of the present disclosure. [Figure 12] FIG. 12 is a diagram illustrating a specific example of the form determination process according to the embodiment of the present disclosure. [Figure 13] FIG. 13 is a flowchart illustrating an example of the procedure of an operation mode setting process executed in the image processing system according to an embodiment of the present disclosure. [Figure 14] FIG. 14 is a diagram illustrating a specific example of the operation mode setting process according to an embodiment of the present disclosure. [Figure 15]FIG. 15 is a flowchart for explaining an example of the procedure of an operation mode setting process executed in the image processing system according to an embodiment of the present disclosure. [Figure 16] FIG. 16 is a flowchart for explaining an example of the procedure of an operation mode setting process executed in the image processing system according to an embodiment of the present disclosure. [Figure 17] FIG. 17 is a flowchart illustrating an example of the procedure of a date extraction process executed in the image processing system according to an embodiment of the present disclosure. [Figure 18] FIG. 18 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 19] FIG. 19 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 20] FIG. 20 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 21] FIG. 21 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 22] FIG. 22 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 23] FIG. 23 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 24] FIG. 24 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 25] FIG. 25 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 26A] FIG. 26A is a diagram showing a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 26B] FIG. 26B is a diagram showing a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 27] FIG. 27 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 28] FIG. 28 is a diagram illustrating a specific example of the date extraction process according to an embodiment of the present disclosure. [Figure 29]FIG. 29 is a flowchart for explaining an example of the procedure of the amount extraction process executed in the image processing system according to the embodiment of the present disclosure. [Figure 30] FIG. 30 is a diagram showing a specific example of the amount extraction process according to an embodiment of the present disclosure. [Figure 31] FIG. 31 is a flowchart illustrating an example of the procedure of the amount extraction process executed in the image processing system according to the embodiment of the present disclosure. [Figure 32] FIG. 32 is a flowchart illustrating an example of the procedure of the amount extraction process executed in the image processing system according to the embodiment of the present disclosure. [Figure 33] FIG. 33 is a flowchart illustrating an example of the procedure of the amount extraction process executed in the image processing system according to the embodiment of the present disclosure. [Figure 34] FIG. 34 is a diagram showing a specific example of the amount extraction process according to an embodiment of the present disclosure. [Figure 35] FIG. 35 is a diagram showing an extraction result of form image processing according to an embodiment of the present disclosure. [Figure 36] FIG. 36 is a diagram showing an extraction result of form image processing according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples that embody the present disclosure and do not limit the technical scope of the present disclosure.
[0012] [Image Processing System 10] 1 is a block diagram showing a configuration of an image processing system 10 according to an embodiment of the present disclosure. The image processing system 10 includes an image processing device 1 and an operation terminal 2. The image processing device 1 and the operation terminal 2 are connected to each other via a network N1 (e.g., the Internet, a LAN, etc.). The image processing system 10 may include multiple operation terminals 2.
[0013] In the image processing system 10, the image processing device 1 acquires image data of documents such as forms transmitted from the operation terminal 2 and extracts desired character strings (character strings to be managed) from the image data. For example, the operation terminal 2 scans paper forms such as estimates, purchase orders, invoices, delivery notes, receipts, and receipts, and transmits the generated image data (e.g., PDF data) to the image processing device 1. The operation terminal 2 also creates a document file of the form based on a user's operation using, for example, a word processing application, and transmits the document file to the image processing device 1 as image data (e.g., searchable PDF data (image + text data)). Upon receiving the image data transmitted from the operation terminal 2, the image processing device 1 performs various processes (described later) on the image data to extract character strings to be managed from the form. For example, the image processing device 1 extracts the date, amount (e.g., total amount), company information (e.g., issuer, destination, registration number), etc., of each form. The image processing device 1 also registers the extracted character strings in a predetermined database. For example, each time image data of an estimate is acquired, the image processing device 1 extracts character strings related to the contents of the estimate (e.g., issue date, estimated amount, issuer, etc.) from the image data and registers them in a database that manages estimates. Also, each time image data of a receipt is acquired, the image processing device 1 extracts character strings related to the contents of the receipt (e.g., issue date, total amount, issuer, etc.) from the image data and registers them in a database that manages receipts. This allows each document to be saved and managed as electronic data. The image processing device 1 also outputs the extracted character strings to the operation terminal 2 or the like and presents the character recognition results to the user.
[0014] The image processing system 10 is an example of an information processing system according to the present disclosure. Note that the information processing system according to the present disclosure may be configured with the image processing device 1 alone.
[0015] [Image processing device 1] 1, the image processing device 1 includes a control unit 11, a storage unit 12, an operation display unit 13, and a communication unit 14. The image processing device 1 may be one or more cloud servers, or one or more physical servers.
[0016] The communication unit 14 is a communication interface for connecting the image processing device 1 to a network N1 by wire or wirelessly and for executing data communication with the operation terminal 2 via the network N1 in accordance with a predetermined communication protocol. The network N1 is configured, for example, by the Internet or a LAN.
[0017] The operation display unit 13 is a user interface that includes a display unit such as a liquid crystal display or an organic EL display that displays various information, and an operation unit such as a mouse, keyboard, or touch panel that accepts operations.
[0018] The storage unit 12 is a non-volatile storage unit such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory that stores various types of information. The storage unit 12 stores control programs that cause the control unit 11 to execute the document determination process (FIG. 7), the operation mode setting process (FIGS. 13, 15, and 16), the date extraction process (FIG. 17), the amount extraction process (FIGS. 29, 31 to 33), and the company information extraction process, which will be described later. For example, the control programs are non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, and are read by a reading device (not shown) such as a CD drive or a DVD drive provided in the image processing device 1 and stored in the storage unit 12. The control programs may also be distributed from a cloud server and stored in the storage unit 12.
[0019] Furthermore, the storage unit 12 stores image data (scanned data, etc.) of documents such as forms acquired from the operation terminal 2.
[0020] FIG. 2 shows an estimate as an example of a document. As shown in FIG. 2, the estimate includes text such as the document classification ("estimate"), issue date, contact information for the estimator (address, telephone number, fax number, person in charge), estimated amount, product name, quantity, standard price, discount amount, subtotal, consumption tax, and total amount. The user scans the estimate using the operation terminal 2 and uploads image data P1 to the image processing device 1. The control unit 11 acquires the image data P1 of the estimate and stores it in the memory unit 12.
[0021] FIG. 3 shows a receipt as another example of a document. As shown in FIG. 3, a receipt contains text such as the issuer's contact information (address, phone number), issue date, product name, sales price, subtotal, consumption tax, total amount, and change. The user scans the receipt using the operation terminal 2 and uploads image data P1 to the image processing device 1. When the control unit 11 acquires the receipt image data P1, it stores it in the memory unit 12.
[0022] In another embodiment, the control unit 11 may acquire a document file of a form created on the operation terminal 2 and store the document file in the storage unit 12.
[0023] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various types of arithmetic processing. The ROM stores in advance control programs such as a BIOS and an OS that cause the CPU to execute various types of processing. The RAM stores various types of information and is used as a temporary storage memory (work area) for the various types of processing executed by the CPU. The control unit 11 controls the image processing device 1 by having the CPU execute various control programs that are pre-stored in the ROM or the storage unit 12.
[0024] Incidentally, there are forms with different formats (formats) such as layout, character font, etc. For example, a general form such as an estimate shown in Figure 2 and a receipt shown in Figure 3 have different spacing between characters, font, etc.
[0025] The receipt is a document printed out from a store cash register, and is a document on which the date of purchase of the product, the store name (issuer), the product (item), the unit price (selling price), transaction details, etc. are printed. The receipt is a document that does not include the recipient's name (the purchaser's name), and documents that include the recipient's name (such as receipts) are excluded from the receipt. The receipt also includes receipts printed out from ticket machines for public transportation (trains, buses, etc.). The receipt does not include coupons, discount coupons, postcards, etc. The general document is a document other than the receipt, such as an estimate, order form, invoice, delivery note, receipt, etc.
[0026] If the same character recognition process (OCR process) is performed on the receipt and the general form, the characters can be properly recognized (character strings extracted) on one form, but the characters cannot be properly recognized on the other form. For example, if the character recognition process corresponding to the general form is performed on the receipt, the problem occurs that adjacent characters with close character spacing or adjacent characters with wide character spacing cannot be properly recognized, resulting in a decrease in character recognition accuracy.
[0027] The image processing device 1 according to the present disclosure has a configuration that can appropriately distinguish between forms (documents) of different formats and improve character recognition accuracy. Specific processing in the control unit 11 of the image processing device 1 will be described below.
[0028] As shown in Fig. 1, the control unit 11 includes various processing units such as an acquisition processing unit 111, a determination processing unit 112, an extraction processing unit 113, a setting processing unit 114, a calculation processing unit 115, and an output processing unit 116. The control unit 11 functions as the various processing units by executing various processes in accordance with the control program. Some or all of the processing units included in the control unit 11 may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the various processing units.
[0029] [Form image processing (overall flow)] Here, when the control unit 11 acquires the image data P1, it sequentially executes the form determination process (Fig. 7), the operation mode setting process (Figs. 13, 15, 16), the date extraction process (Fig. 17), the amount extraction process (Fig. 29, 31 to 33), and the company information extraction process to extract the character string to be managed from the image data P1 and output the extraction results. Fig. 4 shows an example of the procedure for form image processing including each of the above processes.
[0030] The present disclosure can be considered as a form image processing method that executes one or more steps included in the form image processing. One or more steps included in the form image processing described herein may be omitted as appropriate. The steps in the form image processing may be executed in a different order as long as the same effects are achieved. While the following description uses an example in which the control unit 11 of the image processing device 1 executes each step in the form image processing, in other embodiments, one or more processors may execute each step in the form image processing in a distributed manner. Furthermore, when the control unit 11 acquires image data P1 from each of multiple operation terminals 2, it can execute the form image processing for each piece of image data P1 in parallel.
[0031] In step S1, the control unit 11 (acquisition processing unit 111) determines whether or not image data P1 of the form to be processed has been acquired from the operation terminal 2. When the control unit 11 acquires the image data P1 from the operation terminal 2 (S1: Yes), the control unit 11 shifts the processing to step S2. The control unit 11 waits until the image data P1 is acquired from the operation terminal 2 (S1: No).
[0032] In step S2, the control unit 11 (determination processing unit 112) executes a form determination process. Specifically, the control unit 11 determines the form format based on the image data P1. For example, the control unit 11 determines whether the form of the image corresponding to the image data P1 is a receipt-format form (hereinafter also referred to as a "receipt") or a general form format other than a receipt format (hereinafter also referred to as a "general form"). A specific example of the form determination process will be described later.
[0033] If the result of the form determination process is that the form is a receipt (S3: Yes), in step S4, the control unit 11 (setting processing unit 114) provisionally sets the operating mode of the character recognition process to receipt mode. After step S4, the control unit 11 transitions the process to step S5. On the other hand, if the result of the form determination process is that the form is a general form (S3: No), in step S31, the control unit 11 provisionally sets the operating mode of the character recognition process (OCR processing) to general form mode. After step S31, the control unit 11 transitions the process to step S32.
[0034] In the character recognition process (OCR process), the control unit 11 (extraction processing unit 113) extracts characters on a character-by-character basis and executes a wordization process that groups the characters on a character-by-character basis into words (character strings). When the control unit 11 extracts each character string in the character recognition process, it associates information such as a number, character string, and position coordinates with each extracted character string and registers it in the character string list D1. Note that the position coordinates are represented by, for example, the start coordinates (top left coordinates) and end coordinates (bottom right coordinates) of the rectangle of the character string. FIG. 5A shows a specific example of the character string list D1 of the estimate shown in FIG. 2.
[0035] In step S5, the control unit 11 (extraction processing unit 113) executes character recognition processing corresponding to receipt mode (receipt mode OCR processing). The receipt mode OCR processing is processing that can recognize, for example, special characters on receipts (double-width characters, double-height characters, characters with narrow line spacing, characters with narrow character spacing, characters with wide character spacing, etc.). Figure 6 shows the results of the receipt mode OCR processing. After step S5, the control unit 11 proceeds to step S6.
[0036] In step S32, the control unit 11 (extraction processing unit 113) executes character recognition processing corresponding to the general form mode (general form mode OCR processing). The general form mode OCR processing executes preprocessing such as seal imprint removal, background removal, character inversion, and italic correction, and then executes word processing. FIG. 5B shows the result of the general form mode OCR processing. After step S32, the control unit 11 shifts the processing to step S6.
[0037] In step S6, the control unit 11 executes an operation mode setting process. Specifically, the control unit 11 confirms the operation mode provisionally set in steps S4 and S31 as the receipt mode or general form mode. The control unit 11 executes the following date extraction process (FIG. 17), amount extraction process (FIG. 29, FIGS. 31 to 33), and company information extraction process through character recognition processing according to the confirmed operation mode. Specific examples of the operation mode setting process will be described later.
[0038] Next, in step S7, the control unit 11 (extraction processing unit 113) executes a date extraction process. Specifically, the control unit 11 extracts dates, such as the issue date, included in the document by character recognition processing according to the operation mode. A specific example of the date extraction process will be described later.
[0039] Next, in step S8, the control unit 11 (extraction processing unit 113) executes an amount extraction process. Specifically, the control unit 11 extracts the total amount (an example of the specific amount of the present disclosure) included in the document by character recognition processing according to the operation mode. A specific example of the amount extraction process will be described later.
[0040] Next, in step S9, the control unit 11 (extraction processing unit 113) executes a company information extraction process. Specifically, the control unit 11 extracts company information (issuer, destination, registration number, etc.) included in the form by character recognition processing according to the operation mode. A specific example of the company information extraction process will be described later.
[0041] Next, in step S10, the control unit 11 (output processing unit 116) outputs the extraction results of the date extraction process, the amount extraction process, and the company information extraction process. For example, the control unit 11 outputs the issuer (store name), total amount, and date of the receipt, and the issuer (company name), total amount, and date of the estimate.
[0042] Below, we will explain the specific configurations of the document determination process (S2), the operation mode setting process (S6), the date extraction process (S7), the amount extraction process (S8), and the company information extraction process (S9) included in the document image processing (overall flow).
[0043] [Report Judgment Processing] 7 shows an example of the procedure for the form determination process. In the form determination process, the control unit 11 (determination processing unit 112) determines the form format based on the image data P1. For example, the control unit 11 determines whether the document (form) of the image data P1 is a receipt or a general form.
[0044] In step S101, the control unit 11 estimates the image (angle of view) size of the form based on the image data P1 of the form. Specifically, the control unit 11 estimates the image size (receipt size, A-size paper size, etc.) based on the layout of multiple character string rectangles extracted from the image data P1. For example, the control unit 11 estimates the image size based on the coordinates of the character string rectangle located in the upper left and the coordinates of the character string rectangle located in the lower right of the multiple character string rectangles.
[0045] Next, in step S102, the control unit 11 estimates the maximum number of characters in the horizontal direction of the form (estimated number of characters) based on the estimated image size (estimated image size). Specifically, the control unit 11 calculates the average value (average height) of the heights of multiple character string rectangles, and determines the estimated number of characters by dividing the width of the estimated image size by the average height.
[0046] Next, in step S103, the control unit 11 calculates vertical and horizontal offset values based on the image data P1. Specifically, the control unit 11 calculates the difference (offset value) between the image size of the image data P1 and the estimated image size in both the horizontal and vertical directions. For example, for the image data P1 of the receipt shown in FIG. 8, the control unit 11 calculates the difference (x1, y1) between the image size of the image data P1 (H0 × W0) and the estimated image size (H1 × W1) as the offset value.
[0047] Next, in step S104, the control unit 11 determines whether the set conditions are satisfied. If the control unit 11 determines that the set conditions are satisfied (S104: Yes), the control unit 11 shifts the process to step S105. On the other hand, if the control unit 11 determines that the set conditions are not satisfied (S104: No), the control unit 11 shifts the process to step S106.
[0048] For example, if the horizontal margin ratio (W1 / W0) or the vertical margin ratio (H1 / H0) is less than a preset ratio (threshold) and the estimated number of characters is less than a preset number of characters (threshold), the control unit 11 determines that the setting condition is met and shifts the process to step S105. On the other hand, if the margin ratio is equal to or greater than the preset ratio or the estimated number of characters is equal to or greater than the preset number of characters, the control unit 11 determines that the setting condition is not met and shifts the process to step S106. For example, the setting threshold is set to 0.7, and the setting number of characters is set to 30 characters. Note that the setting threshold and the setting number of characters are not limited to these. The setting threshold may be set to different values for the horizontal direction and the vertical direction.
[0049] In step S105, the control unit 11 executes offset processing. Specifically, the control unit 11 executes processing (image crop) to move (offset) the form portion (estimated image portion) to the upper left and cut it out so that the difference (offset value) is eliminated (x1=0, y1=0), as shown in Fig. 9.
[0050] In step S106, control unit 11 calculates the reference height of the form. Specifically, control unit 11 calculates the distance (maximum vertical height) between the vertical coordinate (Y coordinate) of the character string rectangle located in the top position among the character string rectangles extracted from image data P1 and the vertical coordinate (Y coordinate) of the character string rectangle located in the bottom position as reference height Hm. For example, as shown in Fig. 10, control unit 11 calculates the distance between the Y coordinate Ps of the character string rectangle Ks located in the top position and the Y coordinate Pe of the character string rectangle Ke located in the bottom position as reference height Hm.
[0051] In this way, when the estimated image size relative to the input image size is small and the estimated number of characters in the horizontal direction is small, the control unit 11 calculates the reference height Hm in the image after the offset process is performed (see FIG. 9).
[0052] Next, in step S107, the control unit 11 executes a process of correcting the calculated reference height Hm. Specifically, the control unit 11 corrects the reference height Hm by excluding images other than character strings, such as one-dimensional codes, two-dimensional codes, illustrations, and tables. For example, as shown in FIG. 11, if the image of a two-dimensional code is included in the range of the calculated reference height Hm, the control unit 11 subtracts the height Hq of the rectangle of the two-dimensional code from the reference height Hm. The control unit 11 sets the value obtained by subtracting the height Hq of the rectangle as the corrected reference height Hm.
[0053] Next, in step S108, the control unit 11 calculates the total height of the target character string rectangles for each divided area. Specifically, as shown in Fig. 10, the control unit 11 divides the form image (estimated image) into an area (first divided area) that is (W1) / 5 from both the left and right ends, and an area (second divided area) that is (W1) / 3 from both the left and right ends. Then, the control unit 11 calculates the sum of the heights of the multiple character string rectangles included in each divided area (total height).
[0054] For example, the control unit 11 extracts character string rectangles whose upper left x1 coordinate is located within the first divided area on the left, and is located within 5 times the height of the character string rectangle Ks from the left edge, and whose rectangle ratio (width / height) of the horizontally long rectangle is 20 or less, and calculates the total height SL1 of each of the heights of the extracted character string rectangles.
[0055] For example, the control unit 11 extracts character string rectangles whose upper left x1 coordinate is located within the second divided area on the left, and is located within 8 times the height of the character string rectangle Ks from the left edge, and whose rectangle ratio (width / height) of the horizontally long rectangle is 20 or less, and calculates the total height SL2 of each of the heights of the extracted character string rectangles.
[0056] For example, the control unit 11 extracts character string rectangles whose bottom right x2 coordinate is located within the first divided area on the right and whose horizontally long rectangle has a rectangle ratio (width / height) of 20 or less, and calculates the total height SR1 of each of the extracted character string rectangles.
[0057] For example, the control unit 11 extracts character string rectangles whose bottom right x2 coordinate is located within the second divided area on the right and whose horizontally long rectangle has a rectangle ratio (width / height) of 20 or less, and calculates the total height SR2 of each of the extracted character string rectangles.
[0058] For example, control unit 11 extracts horizontally long character string rectangles whose horizontal center coordinates are located at least (W1) / 4 but not more than (W1) × 3 / 4 from the left end and whose upper left x1 coordinate is located at least 5 times the height of character string rectangle Ks from the left end, and calculates the total height SM of the extracted character string rectangles. Furthermore, for example, as shown in Fig. 12, if a character string rectangle is included in the coordinate range of a table, control unit 11 excludes that character string rectangle from the character string rectangles for which the total height SM is calculated.
[0059] Next, in step S109, the control unit 11 calculates the height ratio (distribution density) for each divided area. Specifically, the control unit 11 calculates the height ratio by dividing each calculated total height by the height in the vertical direction (reference height Hm).
[0060] Next, in step S110, the control unit 11 determines whether the height ratio is greater than a preset height ratio (threshold value). If the height ratio is equal to or greater than the preset height ratio (S110: Yes), the control unit 11 shifts the process to step S111. On the other hand, if the height ratio is less than the preset height ratio (S110: No), the control unit 11 shifts the process to step S114.
[0061] Next, in step S110, the control unit 11 determines whether the height ratio is greater than a preset height (threshold value). If the height ratio is equal to or greater than the preset height (S110: Yes), the control unit 11 shifts the process to step S111. On the other hand, if the height ratio is less than the preset height (S110: No), the control unit 11 shifts the process to step S114.
[0062] For example, the control unit 11 shifts the processing to step S111 when the ratio of total height SL1 to reference height Hm is 0.24 or more and the ratio of total height SR1 to reference height Hm is 0.24 or more (first condition). Furthermore, for example, the control unit 11 shifts the processing to step S111 when the ratio of total height SL2 to reference height Hm is 0.30 or more and the ratio of total height SR2 to reference height Hm is 0.30 or more (second condition). Furthermore, for example, the control unit 11 shifts the processing to step S111 when the ratio of total height SM to reference height Hm is 0.60 or more (third condition). Also, for example, the control unit 11 transitions the processing to step S111 when the total ratio of the total height SL1 / reference height Hm and the total height SL2 / reference height Hm is 0.55 or more, and the total ratio of the total height SR1 / reference height Hm and the total height SR2 / reference height Hm is 0.55 or more (fourth condition).
[0063] If the first to fourth conditions are not satisfied, the control unit 11 shifts the process to step S114.
[0064] In step S111, the control unit 11 determines the form corresponding to the image data P1 as a candidate for a receipt.
[0065] Next, in step S112, control unit 11 determines whether the image size of image data P1 corresponds to a predetermined aspect ratio. For example, if the image size corresponds to the aspect ratio of a predetermined standard size such as A4 size, A5 size, B size, or letter size, and the estimated number of characters is 30 characters or more (for portrait size), control unit 11 shifts the process to step S114. On the other hand, if the image size does not correspond to the aspect ratio of the standard size, or if the estimated number of characters is less than 30 characters, control unit 11 shifts the process to step S113.
[0066] In step S113, the control unit 11 determines that the form corresponding to the image data P1 is a receipt.
[0067] In step S114, the control unit 11 determines that the form corresponding to the image data P1 is a general form.
[0068] In this way, the control unit 11 executes the form determination process. After executing the form determination process, the control unit 11 executes the processes from step S3 onwards shown in FIG.
[0069] As described above, the control unit 11 acquires character string rectangles surrounding character strings included in a document (form) based on image data P1 of the document, and determines whether the document is a receipt or a general form based on the layout density of the character string rectangles in an area corresponding to the size of the document. Specifically, the control unit 11 determines whether the document is a receipt or a general form based on the layout density (height ratio) of the character string rectangles in the vertical direction of the area.
[0070] The control unit 11 may also calculate the arrangement density as the total height obtained by adding up the vertical heights of the multiple character string rectangles relative to the vertical reference height Hm of the area, and determine the document to be the receipt if the arrangement density is equal to or greater than a threshold value, and determine the document to be the general form if the arrangement density is less than the threshold value.
[0071] In addition, the control unit 11 may calculate the placement density as the total height obtained by adding up the vertical heights of each of the multiple character string rectangles relative to the vertical reference height Hm of the area, and determine that the document is a general form if the placement density is less than a threshold value.
[0072] In addition, the control unit 11 may determine that the document is a general form if the layout density is equal to or greater than a threshold, the aspect ratio of the area matches a predetermined value, and the estimated number of characters horizontally in the area is equal to or greater than a predetermined number of characters, and may determine that the document is a receipt if the layout density is equal to or greater than a threshold and the aspect ratio of the area does not match the predetermined value, or if the layout density is equal to or greater than a threshold and the estimated number of characters is less than the predetermined number of characters.
[0073] The control unit 11 may also estimate the vertical and horizontal sizes (image size) of the region based on the maximum and minimum values of the horizontal position coordinates and the maximum and minimum values of the vertical position coordinates of each of the plurality of character string rectangles. For example, the control unit 11 estimates the image size by setting the horizontal minimum value as the left vertical line of the crop, the horizontal maximum value as the right vertical line of the crop, the vertical minimum value as the upper line of the crop, and the vertical maximum value as the lower line of the crop, for the position coordinates of each of the plurality of character string rectangles. The control unit 11 may also determine the region based on the difference between the size of the image data P1 (input image size) and the estimated image size estimated from the arrangement positions of the plurality of character string rectangles.
[0074] In addition, the control unit 11 may calculate the reference height Hm (see Figure 11) as the length from the first character string rectangle Ps placed at the top position of the area to the second character string rectangle Pe placed at the bottom position of the area, excluding the height of a non-character string rectangle (one-dimensional code, two-dimensional code, illustration, etc.) placed between the first character string rectangle and the second character string rectangle.
[0075] Furthermore, the control unit 11 may calculate the arrangement densities corresponding to the plurality of character string rectangles included in divided areas (see FIG. 10) obtained by dividing the area in the horizontal (left and right) direction by vertical lines.
[0076] According to the above configuration, it is possible to appropriately determine whether the document (form) of the input image data P1 is a receipt or a general form.
[0077] [Operation mode setting process] 13 shows an example of the procedure for the operation mode setting process. In the operation mode setting process, the control unit 11 (extraction processing unit 113, setting processing unit 114) sets an operation mode for extracting character strings to be extracted (date, amount, company information, etc.) from the document (ledger) of the image data P1. After the OCR process of steps S5 and S32 shown in FIG. 4, the control unit 11 executes the following operation mode setting process.
[0078] In step S201, the control unit 11 (extraction processing unit 113) extracts a category of the form based on the image data P1 of the form. For example, the control unit 11 extracts a character string consisting of relatively large characters (for example, characters with a font size of 16 points or more) in the area above the center position in the vertical direction of the form, and including "quotation," "order form," "invoice," "delivery note," and "receipt."
[0079] Next, in step S202, the control unit 11 (setting processing unit 114) determines whether the operation mode provisionally set in steps S4 and S31 shown in Fig. 4 is the general form mode. If the provisionally set operation mode is the general form mode (S202: Yes), the control unit 11 shifts the process to step S206. On the other hand, if the provisionally set operation mode is the receipt mode (S202: No), the control unit 11 shifts the process to step S203.
[0080] In step S206, the control unit 11 (setting processing unit 114) determines whether the extracted category is a document other than an "estimate," "order form," "invoice," "delivery note," or "receipt" ("other document"). If the category is a document other than an "estimate," "order form," "invoice," "delivery note," or "receipt" ("other document"), the control unit 11 proceeds to step S12 and outputs the document as a document other than an "other document" (S206: Yes). In this case, the control unit 11 ends the process without outputting the extracted character string (date, amount, company information, etc.). On the other hand, if the category is not a document other than an "other document" (S206: No), the control unit 11 proceeds to step S207 and sets the operating mode to the general form mode. After step S207, the control unit 11 proceeds to step S7 (see FIG. 6).
[0081] In step S203, the control unit 11 (setting processing unit 114) determines whether the extracted category is "receipt" or "other document." If the category is receipt or other document (S203: Yes), the control unit 11 proceeds to step S204 and sets the category to "receipt." Then, in step S205, the control unit 11 sets the operating mode to receipt mode. After step S205, the control unit 11 proceeds to step S7 (see FIG. 6).
[0082] On the other hand, if the extracted category is not "receipt" or "other document" (S203: No), the control unit 11 proceeds to step S207 and sets the operation mode to general form mode. In this way, even if the control unit 11 has provisionally set the operation mode to receipt mode based on the layout of the character string rectangles, etc. (step S4 in FIG. 4), if it has determined that the category is "quotation," "order form," "invoice," or "delivery note," it changes the receipt mode to general form mode.
[0083] Steps S201 to S207 of the above-described operation mode setting process correspond to step S6 shown in FIG.
[0084] FIG. 14 shows the relationship between extraction categories and operation modes. As shown in FIG. 14, when the target form (image data P1) is provisionally set to general form mode and the extraction category is "quote," "order form," "invoice," "delivery note," or "receipt," the operation mode is set to "general form mode." When the general form mode is provisionally set and the extraction category is "other documents," the character string to be extracted is not extracted. For example, if documents accompanying the form, such as quotation details, billing details, or catalogs, are scanned together, they are excluded as other documents.
[0085] Furthermore, if the target form (image data P1) is provisionally set to receipt mode and the extraction classification is "receipt" or "other document," the operating mode is set to "receipt mode." For example, some receipts may not contain any mention of receipt or may not contain the keywords required to determine whether they are receipts, and may be determined to be "other document." Therefore, an input image that is determined (provisionally set) to be in receipt mode and determined to be an "other document" is processed as a receipt. Note that receipts are treated as a type of receipt and are classified as "receipt." In contrast, if the receipt mode is provisionally set and the extraction classification is "quotation," "order form," "invoice," or "delivery note," the operating mode is changed to "general form mode."
[0086] As described above, the control unit 11 determines whether the document (form) is a receipt or a general form based on the image data P1 of the document, and extracts a category that represents the type of the document from among the multiple items included in the document. Furthermore, based on the determination result and the extracted category, the control unit 11 sets the operation mode for extracting target characters included in the multiple items to a receipt mode (first operation mode) or a general form mode (second operation mode).
[0087] Furthermore, when the control unit 11 determines that the document is a general document in the document determination process (see FIG. 7) and extracts an estimate, order form, delivery note, invoice, or receipt as the category, it sets the operation mode to the general document mode. Furthermore, when the control unit 11 determines that the document is a receipt in the document determination process (see FIG. 7) and extracts an estimate, order form, delivery note, or invoice as the category, it sets the operation mode to the general document mode.
[0088] In addition, when the control unit 11 determines that the document is a receipt in the document determination process (see FIG. 7) and extracts a receipt or other document (other document) as the classification, it sets the operation mode to receipt mode.
[0089] In addition, when the control unit 11 determines that the document is a general form, it extracts the classification from image data P1 on which OCR processing for general form mode (second character recognition processing) corresponding to the general form mode has been performed, and when the control unit 11 determines that the document is a receipt, it extracts the classification from image data P1 on which OCR processing for receipt mode (first character recognition processing) corresponding to the receipt mode has been performed.
[0090] With the above configuration, even if the document type is erroneously determined, the correct target characters can be extracted. For example, even if the control unit 11 erroneously determines that a general form document is a receipt and sets the receipt mode, if the extracted category is an estimate, purchase order, delivery note, or invoice, the control unit 11 changes the receipt mode to the general form mode. In other words, even if the control unit 11 erroneously determines that a general form document is a receipt, the control unit 11 can change (recover) the document to a general form using the category extraction result. This allows the control unit 11 to extract target characters as a general form in subsequent processing.
[0091] Another embodiment of the operation mode setting process will now be described. Figures 15 and 16 show an example of the procedure of the operation mode setting process according to this embodiment.
[0092] In the configuration shown in FIG. 15, if the category is "other documents" in step S206 (S206: Yes), the control unit 11 proceeds to step S204 and sets the category to "receipt." Also, in the configuration shown in FIG. 15, the control unit 11 determines in step S10 whether the extraction of each item to be extracted (date, amount, company information, etc.) is appropriate. If the extraction of each item to be extracted is appropriate (S13: Yes), the control unit 11 proceeds to step S11 and outputs the extraction results. On the other hand, if the extraction of each item to be extracted is inappropriate (S13: No), the control unit 11 proceeds to step S12 and outputs the result as "other documents." In step S10, for example, in the case of a receipt, the control unit 11 determines whether the destination was extracted, whether the date was extracted in receipt mode, whether the date was extracted in general form mode, etc. For example, if the control unit 11 cannot extract the destination but can extract the issuer and amount, it determines that the extraction is appropriate. Also, for example, when only the date is extracted, the control unit 11 determines that the extraction is inappropriate. The other steps in the process shown in FIG.
[0093] In this way, when the control unit 11 determines that the document is the general form or the receipt and extracts other documents as the classification, it sets the operation mode to the receipt mode.
[0094] 16, the control unit 11 performs a category recovery process (if Yes in S206, set to receipt in S204), and then extracts each item in both the receipt mode and the general form mode (S7 to S9). Then, in step S10 (S11 to S13), the control unit 11 determines the final document classification result and category from the document classification result, category extraction result, and each item extraction result, and selects output of items according to the document classification result and category.
[0095] In this way, the control unit 11 executes a first process to extract target characters from the image data P1 in the general form mode, and a second process to extract target characters from the image data P1 in the receipt mode, and determines the output result corresponding to the category based on the results of the first process and the second process.
[0096] When the operation mode is set as described above, the control unit 11 extracts character strings of the extraction target (management target), such as date, amount, company information, etc., in accordance with the operation mode.
[0097] In other words, the control unit 11 automatically determines whether the input image data P1 is a general form or a receipt, and if it is determined to be a general form, performs OCR processing using an OCR engine for general forms (OCR for general form mode), and uses the OCR results to determine which general form category it falls into in a category extraction process.If it determines the category to be an estimate, purchase order, delivery note, invoice, or receipt, it performs a process to extract items for general forms (total amount, date, issuer, destination, etc.), but if it determines it to be some other document, it does not perform the item extraction process. In response to this, if the control unit 11 determines that the document is a receipt, it performs OCR processing using an OCR engine for receipts (receipt mode OCR), and uses the OCR results to determine which general document category it falls into in a category extraction process.If the category is determined to be an estimate, purchase order, delivery note, or invoice, it updates the document classification result to a general document (recovery process) and performs processing to extract general document items (total amount, date, issuer, destination, etc.) according to the updated document classification result and the determined category.On the other hand, if the control unit 11 determines that the category is a receipt or other document, it updates the other document category to a receipt (recovery process) and performs processing to extract receipt items (total amount, date, issuer, destination, etc.) according to the document classification result and the updated category.
[0098] In addition, if the control unit 11 determines that the document category is other document regardless of the document classification result, it updates the document classification result to receipt (recovery processing) as necessary, updates the classification of other document to receipt (recovery processing), and performs processing to extract items for receipts (total amount, date, issuer, etc.).
[0099] Furthermore, when the control unit 11 determines that the category is other documents, it updates the category of other documents to receipts (recovery processing). Furthermore, the control unit 11 performs processing to extract items for general documents (total amount, date, issuer, destination, etc.) from the categories of estimates, purchase orders, delivery notes, invoices, and receipts. Furthermore, the control unit 11 performs processing to extract items for receipts (total amount, date, issuer, etc.) from the categories of estimates, purchase orders, delivery notes, invoices, and receipts. Furthermore, the control unit 11 updates the document discrimination result (recovery processing) based on the document discrimination result, category extraction result, and the results of each extraction processing, and selects and outputs output items according to the updated document discrimination result and category.
[0100] Below, specific examples of the date extraction process (FIG. 17), the amount extraction process (FIG. 29, FIGS. 31 to 33), and the company information extraction process will be explained.
[0101] [Date extraction process] The control unit 11 is configured to appropriately extract dates such as the issue date for each of general forms and receipts. In the date extraction process, the control unit 11 (extraction processing unit 113, calculation processing unit 115) extracts dates according to the document type (general form, receipt). Figure 17 shows an example of the procedure for the date extraction process.
[0102] In step S301, the control unit 11 extracts target character string candidates. Specifically, the control unit 11 extracts, from a plurality of character strings (character string list D1) acquired by OCR processing, character strings that include characters related to dates and are composed of a predetermined number of characters, as target character string candidates. For example, the control unit 11 extracts, from the character string list D1 (see FIG. 5A), character strings that begin with "20," "19," "Reiwa," or "Heisei," as well as character strings that include " / ," ".", ",", or "-." FIG. 18 shows a target character string candidate list D2 in which target character string candidates are registered.
[0103] Next, in step S302, if the target character string candidate contains an incorrectly recognized character, the control unit 11 executes a correction process to correct the character. For example, if the characters for "year," "month," and "day" are recognized as incorrect characters, the control unit 11 corrects the character error. Also, for example, if the characters for the date numeric value are recognized as incorrect characters, the control unit 11 corrects the numeric value character error. The control unit 11 registers the corrected character string in the target character string candidate list D2.
[0104] Next, in step S302, the control unit 11 extracts date string candidates from the multiple string candidates registered in the target string candidate list D2. Specifically, the control unit 11 extracts date string candidates by determining whether the target string candidate contains a "year," a "month," or a "day." The control unit 11 also identifies the characters for "year," "month," and "day," and identifies numeric strings indicating the year, month, and day from the identified positions. The control unit 11 extracts date string candidates by excluding target strings for which numeric strings cannot be identified. FIG. 19 shows the date string candidate list D3 in which date string candidates are registered.
[0105] Furthermore, the control unit 11 identifies the character string of numbers indicating the year, month, and day from the positions of the characters "year," "month," and "day," and converts (corrects) the character string of numbers into a numerical value.
[0106] Next, in step S303, the control unit 11 (calculation processing unit 115) calculates an evaluation value (reliability) for each of the extracted date string candidates. Specifically, the control unit 11 calculates the evaluation value for each date string candidate based on at least one of the following conditions: (1) the positional relationship between the date string candidate and related words, (2) the positional relationship between the date string candidate and NG words, and (3) the position of the date string candidate within the image. Fig. 20 shows an example of the evaluation value (total evaluation value) for each date string candidate.
[0107] Next, in step S304, the control unit 11 (extraction processing unit 113) identifies a date to be extracted. Specifically, the control unit 11 identifies the date character string candidate with the highest total evaluation value (reliability) from among the multiple date character string candidates. In another embodiment, the control unit 11 may identify a predetermined number of date character string candidates from among the multiple date character string candidates in descending order of reliability.
[0108] Here, a specific example of a method for calculating the evaluation values corresponding to the conditions (1) to (3) and a method for identifying the date character string candidates will be described below.
[0109] (1) Positional relationship between date string candidates and related words Specifically, when a related word related to a date is included within a predetermined range from the position of the date string candidate in the input image, the calculation processing unit 115 adds a predetermined evaluation value to the date string candidate. FIG. 21 shows an example of keyword information D4. In the keyword information D4, related words and NG words are associated and registered for each type (category) of form. The related words include terms related to forms and terms essential to forms. For example, related words for invoices include terms such as issue date and billing date, and related words for estimates include terms such as issue date and estimate date. The NG words include terms with little relevance to forms and terms that are not essential to documents. For example, NG words for invoices include terms such as payment deadline, closing date, and delivery date, NG words for estimates include terms such as payment deadline and order date, and NG words for receipts include terms such as campaign, points, coupon, application deadline, rank, lottery, redemption, application, and deadline. The related words and the NG words may be registered in advance by a user's registration operation, or may be registered in advance by the control unit 11 according to the classification of the form.
[0110] The calculation processing unit 115 assigns a predetermined evaluation value to a date string candidate when a related word related to a date is included within a predetermined range from the position of the date string candidate in the input image (image data P1). For example, the related word "issue date" (position (x2, y2)) of "invoice" is located near the position (x3, y3) of the character string "March 31, 2023" of the date string candidate. In this case, the calculation processing unit 115 assigns a positive evaluation value to "March 31, 2023." The predetermined range is set to a predetermined length in each of the vertical and horizontal directions based on the center position (center coordinates) of the rectangle of the date string candidate. Note that if a character string within the predetermined range does not completely match a related word in the keyword information D4, the calculation processing unit 115 may assign an evaluation value according to the number of characters included in the character string that match characters included in the related word. For example, the calculation processing unit 115 assigns a higher positive evaluation value the more matching characters there are.
[0111] (2) Positional relationship between date string candidates and NG words If the input image (image data P1) contains a NG word (non-related character) unrelated to dates within a predetermined range from the position of the date string candidate, the calculation processing unit 115 subtracts a predetermined evaluation value from the date string candidate (assigning a negative evaluation value). For example, the NG word "payment deadline" (position (x17, y17)) for "invoice" is located near the position (x18, y18) of the character string "April 4, 2023" of the date string candidate. In this case, the calculation processing unit 115 assigns a negative evaluation value to "April 4, 2023." Note that if a character string within the predetermined range does not completely match a related word in the keyword information D4, the calculation processing unit 115 may assign an evaluation value according to the number of characters included in the character string that match characters included in the related word. For example, the calculation processing unit 115 assigns a higher negative evaluation value the more matching characters there are.
[0112] (3) Position of the date string candidate in the image The calculation processing unit 115 assigns an evaluation value to the date string candidate according to the distance from the center position (P0 in FIG. 22) of the input image (image data P1) to the position of the date string candidate. For example, for the date string candidate ("March 31, 2023") that includes a related word related to the date within the predetermined range, the calculation processing unit 115 assigns an evaluation value according to the distance from the center position P0 (see FIG. 22) to the center position of the rectangle of "March 31, 2023." For example, if the form is a general form, the calculation processing unit 115 sets a higher positive evaluation value the farther the distance. On the other hand, if the form is a receipt, for example, the date tends to be printed closer to the center of the form, so the calculation processing unit 115 sets a higher positive evaluation value the closer the distance. Note that the calculation processing unit 115 may set an evaluation value according to the level of the distance. Furthermore, the calculation processing unit 115 may calculate the ratio of the distance to half the length of the diagonal of the angle of view, and set an evaluation value according to this ratio.
[0113] In the above configuration, the reference point for the distance is set to the center position of the input image. For example, when the document category is "invoice," the control unit 11 determines that character strings to be extracted, such as "issue date" and "invoice date," are usually placed in the corners of the document (top right, top left, etc.) and are rarely placed in the center of the document. Therefore, the control unit 11 sets the reference point to a position where they are unlikely to be placed. However, the position of the reference point is not limited to the center of the document, and may be set to the top right corner, top left corner, etc., depending on the document category.
[0114] The calculation processing unit 115 calculates a total evaluation value for each date string candidate by adding up the evaluation values calculated based on at least one of the following conditions: (1) the positional relationship between the date string candidate and related words, (2) the positional relationship between the date string candidate and NG words, and (3) the position of the date string candidate within the image. FIG. 23 shows an example of the total evaluation value for each date string candidate. The calculation processing unit 115 may also convert the total evaluation value into a reliability indicating the likelihood of the extraction target. For example, the calculation processing unit 115 sets the reference value (initial value) of the reliability to 50% and calculates a reliability of 0 to 100% according to the total evaluation value.
[0115] The extraction processing unit 113 identifies a date to be extracted. Specifically, the extraction processing unit 113 identifies a predetermined date string candidate based on the evaluation values of each of the multiple date string candidates calculated by the calculation processing unit 115. For example, the extraction processing unit 113 identifies the date string candidate with the highest evaluation value (total evaluation value) from among the multiple date string candidates. Furthermore, the extraction processing unit 113 identifies the date string candidate with the highest reliability from among the multiple date string candidates. In the above example, the extraction processing unit 113 identifies the date string candidate "March 31, 2023" with a reliability of 90% (see FIG. 24). In another embodiment, the extraction processing unit 113 may identify a predetermined number (e.g., the top five) of date string candidates from among the multiple date string candidates in descending order of reliability.
[0116] An example of a receipt is shown in Figure 25. Also, Figure 26A shows an example of calculation of an evaluation value corresponding to the condition (3) in an estimate, and Figure 26B shows an example of calculation of an evaluation value corresponding to the condition (3) in a receipt.
[0117] For estimates, as shown in Fig. 26A, the calculation processing unit 115 calculates a higher evaluation value the farther the distance from the center position P0 of the form. In contrast, for receipts, as shown in Fig. 26B, the calculation processing unit 115 calculates a higher evaluation value the closer the distance from the center position P0 of the form.
[0118] A receipt may contain information such as advertisements. In this case, the vertical length of the document tends to increase depending on the amount of information. A longer receipt length may prevent the extraction of an appropriate date based on the center position P0. For example, in the example shown in FIG. 27, if "September 1, 2023" is closer to the center position P0 than the actual issue date, "August 22, 2023," the evaluation value of "September 1, 2023" may be the highest. Taking into account the influence of receipt length, the control unit 11 assigns additional values to each date string candidate in ascending order of y-coordinate. For example, an evaluation value of "60" is added to "August 22, 2023," the top date string candidate, an evaluation value of "30" is added to "September 1, 2023," the second date string candidate, and an evaluation value of "September 30, 2023," the third date string candidate.
[0119] As a result, as shown in FIG. 28, "August 22, 2023" has the highest evaluation value ("90"), allowing the receipt date to be extracted appropriately. The control unit 11 may also pre-set a predetermined number of additional values (for example, the top five) in ascending order of the y coordinate of the character string rectangle of the date string candidate. In this way, the control unit 11 sets an evaluation value for the date string candidate according to the position coordinate of the date string candidate in the image.
[0120] When the control unit 11 identifies the date character string candidates in the above manner (S305), it proceeds to step S8 (FIG. 4). The control unit 11 outputs the date to be extracted each time an input image is acquired. The control unit 11 may also display multiple date character string candidates in descending order of the evaluation value.
[0121] As described above, the control unit 11 obtains the results of the character recognition process (OCR process) performed on the character strings contained in the image of the document (form), calculates an evaluation value for each of the multiple date string candidates contained in the result of the character recognition process according to the format of the document (receipt, general form), and extracts a date string based on the evaluation value calculated for each of the multiple date string candidates.
[0122] Furthermore, the control unit 11 sets the evaluation value for the date string candidate according to the distance from the center position P0 of the image to the position of the date string candidate. For example, when the document is the receipt, the control unit 11 sets a higher evaluation value for the date string candidate as the distance is shorter. Also, for example, when the document is the general form, the control unit 11 sets a higher evaluation value for the date string candidate as the distance is longer.
[0123] Furthermore, when the image contains date-related characters within a predetermined range from the position of the date string candidate, the control unit 11 may add a predetermined value to the evaluation value of the date string candidate. Furthermore, when the image contains non-date-related characters (NG words) within the predetermined range from the position of the date string candidate, the control unit 11 may subtract a predetermined value from the evaluation value of the date string candidate.
[0124] According to the above configuration, dates can be appropriately extracted from documents of different formats (receipts, general forms).
[0125] [Amount extraction process] The control unit 11 is configured to appropriately extract a specific amount (specific amount of the present disclosure) such as a total amount for each of a general form and a receipt. In the amount extraction process, the control unit 11 (extraction processing unit 113) extracts an amount according to the document type and format (general form, receipt). Figure 29 shows an example of the procedure for the amount extraction process corresponding to a general form.
[0126] In step S401, the control unit 11 (extraction processing unit 113) extracts a string of numbers that seems to be an amount (amount candidate) from a plurality of character strings obtained by OCR processing. Specifically, the control unit 11 extracts a string of numbers as an amount candidate according to a predetermined extraction rule in the central region Ar1 (see FIG. 30) excluding the upper and lower parts of the form. For example, the control unit 11 excludes a character string containing symbols other than numbers ("-", " / ", "〒", etc.), and extracts a character string composed only of numbers not containing such symbols as an amount candidate. Note that the control unit 11 may extract a character string composed of numbers and characters such as "¥", ",", "yen", "JPY", "gold", etc. that represent an amount as an amount candidate. In the example shown in FIG. 30, the control unit 11 excludes the character string E1 of the postal code, the character string E2 of the telephone number, and the character string E3 of the FAX number, and extracts the character string A1 of "¥5,489,000", the character string A2 of "¥5,000,000", the character string A3 of "¥10,000", the character string A4 of "¥4,990,000", the character string A5 of "¥4,990,000", the character string A6 of "¥499,000", and the character string A7 of "¥5,489,000" as amount candidates, respectively. Also, for example, the control unit 11 searches for a character string corresponding to "standard price", "discount amount", "consumption tax" from the result of the OCR processing, and searches for a character string of numbers (amount) existing in the right direction or the downward direction from the position information of the character string, and extracts the number at the closest position as an amount candidate character string. Thus, the control unit 11 extracts an amount candidate from the image data P1.
[0127] Next, in step S402, the control unit 11 sets an item name (keyword) necessary for calculating the target item value. For example, when the estimated amount (estimated amount including tax, estimated amount excluding tax) in the estimate is the target amount (target item value) to be extracted, the control unit 11 sets "consumption tax" as the keyword.
[0128] Next, in step S403, the control unit 11 extracts a candidate amount corresponding to the keyword ("consumption tax"). For example, the control unit 11 searches for a character string containing "tax" among the multiple candidate amount amounts in the image data P1, and determines whether there is a candidate amount to the right or below the searched character string. If there is a candidate amount to the right or below the character string, the control unit 11 extracts it as the candidate amount corresponding to "consumption tax." In this case, the control unit 11 extracts "¥499,000" present to the right of the character string "tax" as the candidate amount.
[0129] Next, in step S404, the control unit 11 sets "discount amount" as an item name (keyword) required to calculate the target item value.
[0130] Next, in step S405, the control unit 11 extracts a price candidate corresponding to the keyword ("discount amount"). For example, the control unit 11 searches for a character string containing "discount" among the multiple price candidates in the image data P1, and determines whether there is a price candidate to the right or below the searched character string. If there is a price candidate to the right or below the character string, the control unit 11 extracts it as the price candidate corresponding to "discount amount." In this case, the control unit 11 extracts "¥10,000" located below the character string "discount" as the price candidate.
[0131] Next, in step S406, the control unit 11 sets "standard price" as an item name (keyword) required to calculate the target item value.
[0132] Next, in step S407, the control unit 11 extracts a price candidate corresponding to the keyword ("standard price"). For example, the control unit 11 searches for a character string containing "standard" or "reference" among the multiple price candidates in the image data P1, and determines whether there is a price candidate to the right of or below the searched character string. If there is a price candidate to the right of or below the character string, the control unit 11 extracts it as the price candidate corresponding to "standard price." In this case, the control unit 11 extracts "¥5,000,000" located below the character string "standard" as the price candidate.
[0133] Next, in step S408, the control unit 11 identifies the largest and second largest prices from among the multiple price candidates corresponding to the keyword. For example, if the control unit 11 finds a price candidate for "standard price," it excludes that price candidate, and if the control unit 11 finds a price candidate for "discount amount," it excludes the largest price as the price before discount if the difference between the largest and second largest prices corresponds to the discount amount.
[0134] In the estimate shown in FIG. 2, first, the control unit 11 extracts the largest amount ("¥5,489,000"), the second largest amount ("¥5,000,000"), and the third largest amount ("¥4,990,000") from the price candidate extracted in step S401. Next, the control unit 11 determines whether or not there is a price candidate ("¥5,000,000") that corresponds to the "standard price" among these three amounts. In the example of FIG. 2, the second largest amount ("¥5,000,000") matches the standard price candidate ("¥5,000,000"), so the control unit 11 determines that the second amount ("¥5,000,000") may be the standard price. Next, the control unit 11 subtracts the second largest amount ("¥5,000,000") from the largest amount ("¥5,489,000"), the third largest amount ("¥4,990,000") from the first largest amount ("¥5,489,000"), and the third largest amount ("¥4,990,000") from the second largest amount ("¥5,000,000"). If any of the differences obtained by each subtraction match the discount amount, the control unit 11 determines that the amount to be subtracted is the amount before the discount. If none of the differences obtained by each subtraction match the discount amount, the control unit 11 excludes only the candidate amount that is determined to be the standard price. If there is no candidate amount that is determined to be the standard price, the control unit 11 leaves the amount as is. In the example shown in FIG. 2, first, the difference (¥489,000) obtained by subtracting the second largest amount from the first largest amount does not match the discount amount (¥10,000). Next, the difference (¥499,000) obtained by subtracting the third largest amount from the first largest amount also does not match the discount amount (¥10,000). Finally, the difference (¥10,000) obtained by subtracting the third largest amount from the second largest amount matches the discount amount (¥10,000). Therefore, the control unit 11 determines that the second largest amount is the amount before the discount. The control unit 11 excludes the second largest amount (¥5,000,000), which may be the standard price and has been determined to be the amount before the discount.The control unit 11 extracts the largest amount ("5,489,000 yen") and the second largest amount ("4,990,000 yen"), excluding the excluded amount ("5,000,000 yen").The control unit 11 then identifies "5,489,000 yen" as the largest amount and "4,990,000 yen" as the second largest amount.
[0135] Next, in step S409, the control unit 11 determines the estimated amount (the estimated amount including tax and the estimated amount excluding tax). Here, if the amount obtained by subtracting the second largest amount ("4,990,000 yen") from the largest amount ("5,489,000 yen") matches the amount corresponding to "consumption tax" ("499,000 yen"), the control unit 11 determines the largest amount ("5,489,000 yen") as the estimated amount including tax, and the second largest amount ("4,990,000 yen") as the estimated amount excluding tax.
[0136] If "consumption tax" is not included in the keywords, the control unit 11 calculates the ratio between the largest amount and the second largest amount, and determines whether the calculated ratio corresponds to the consumption tax rate. In this case, if the control unit 11 has been able to extract the issue date from the OCR results in advance, the control unit 11 determines the largest amount as the estimated amount including tax if the tax rate corresponding to that date matches the calculated tax rate according to the issue date, and determines the second largest amount as the estimated amount excluding tax.
[0137] In addition, if the control unit 11 is able to extract a string such as "10%" along with the keyword "consumption tax," and the tax rate (percentage) of the string matches the calculated tax rate, it determines the largest amount as the estimated amount including tax, and determines the second largest amount as the estimated amount excluding tax.
[0138] Here, the control unit 11 may determine the amount depending on the classification of the document. For example, if the document is an estimate or an order form, and the keyword "consumption tax" is not present, and the ratio between the largest amount and the second largest amount does not correspond to the consumption tax rate, the control unit 11 determines whether the strings "excluding tax" and "including tax" are present around the position of the largest amount, and if "excluding tax" or neither is present, it determines that the amount including tax is not displayed, and determines the largest amount as the amount including tax. Also, if the string "including tax" is present, the control unit 11 determines that the amount including tax is not displayed, and determines the largest amount as the amount including tax.
[0139] For example, if the document is an invoice or delivery note, and the keyword "consumption tax" is not present, and the ratio between the largest and second largest amounts does not correspond to the consumption tax rate, the control unit 11 determines whether the character strings "excluding tax" and "including tax" are present around the position of the largest amount, and if "including tax" is present or neither, it determines that the amount excluding tax is not displayed, and determines the largest amount as the amount including tax. Also, if the character string "excluding tax" is present, the control unit 11 determines that the amount including tax is not displayed, and determines the largest amount as the amount excluding tax.
[0140] If "consumption tax" is set as the target item name of the extraction target, the control unit 11 determines the amount of consumption tax.
[0141] The control unit 11 executes the amount determination process for the general form in the above manner. After determining the estimated amount (estimated amount including tax, estimated amount excluding tax), which is the target item value, the control unit 11 outputs the result of the amount determination process in step S10 shown in FIG.
[0142] FIG. 31 shows an example of the procedure for the amount extraction process corresponding to a receipt.
[0143] In step S501, the control unit 11 executes a search process for rule 1. Specifically, the control unit 11 searches for a keyword "total" and a character string of an amount corresponding to the keyword. A specific example of rule 1 will be described later.
[0144] Next, in step S502, the control unit 11 executes a search process for rule 2. Specifically, the control unit 11 searches for a character string of the total amount. A specific example of rule 2 will be described later.
[0145] Next, in step S503, the control unit 11 executes a search process for rule 3. Specifically, the control unit 11 searches for a keyword "consumption tax" and a character string of an amount corresponding to the keyword. A specific example of rule 3 will be described later.
[0146] Next, in step S504, the control unit 11 executes a search process for rule 4. Specifically, the control unit 11 searches for a keyword "change" and a character string of an amount corresponding to the keyword. A specific example of rule 4 will be described later.
[0147] Next, in step S505, the control unit 11 executes the search process of rule 5. Specifically, the control unit 11 searches for numeric keywords to be excluded and numeric character strings corresponding to the keywords, and deletes the character strings from the character strings searched for by the processes of rules 1 to 4. A specific example of rule 5 will be described later.
[0148] The control unit 11 identifies and outputs the total amount based on the character string searched (extracted) by the processing of the rules 1 to 5.
[0149] Specifically, in step S506, if the total amount corresponding to the keyword "total" is clear, the control unit 11 outputs the total amount and ends the process.
[0150] If the total amount is not clear in step S506, in step S507, the control unit 11 calculates the total amount from the maximum and second values of the total amount candidates and the change, outputs the calculated total amount, and ends the process.
[0151] If the control unit 11 is unable to extract the change character string, in step S508, it calculates the total amount from the largest and second largest total amount candidates and the consumption tax, outputs the calculated total amount, and ends the process. In steps S507 to S508, the control unit 11 determines whether the amount includes tax or excludes tax and outputs the total amount. After steps S507 to S508, the control unit 11 causes the process to proceed to step S9 shown in FIG. 4.
[0152] 32 shows an example of the procedure for the search process of rule 1. In step S511, the control unit 11 (extraction processing unit 113) searches for the keyword "total." Specifically, the control unit 11 searches for a character string containing the character "total" from multiple character strings acquired by OCR processing. If the control unit 11 extracts the keyword "total" (S512: Yes), it shifts the process to step S515.
[0153] On the other hand, if the control unit 11 has not extracted the keyword "total" (S512: No), it shifts the process to step S513. In step S513, the control unit 11 searches for a keyword related to "amount". Specifically, the control unit 11 searches for a character string containing characters related to "amount" (for example, "amount", "actual total", etc.) from multiple character strings acquired by the OCR process. If the control unit 11 has extracted a keyword related to "amount" (S514: Yes), it shifts the process to step S515.
[0154] When searching for keywords, the control unit 11 executes the conventional (non-AI learning) OCR process for the receipt mode and the AI learning OCR process, and searches for keywords based on the results of both OCR processes (OCR engines).
[0155] In step S515, the control unit 11 executes a process of searching for an amount corresponding to the keyword (amount search process). FIG. 33 shows a specific example of the amount search process. In the amount search process, the control unit 11 executes AI learning-based OCR processing to search for an amount.
[0156] In AI learning, for example, the control unit 11 pre-learns, using a learning model, a character string including numbers such as an amount, a telephone number, a postal code, a company number, and a document number as a character string of numbers. Then, when the control unit 11 acquires the image data P1, it executes extraction of a character string of numbers by inference processing of AI and OCR processing of the entire image. Next, the control unit 11 performs amount extraction using the result of the OCR processing (with delimiter information in string units) and the result of extraction of a character string of numbers (coordinate positions of the upper left and lower right of the string rectangle) as inputs. Since an AI infers and outputs a rectangular area that seems to be a character string of numbers from the features of the image, if the learning accuracy is improved, even if a number is misrecognized as an alphabet or other symbols in the OCR processing, the output of the AI is trusted and corrected to a numeric character, and it becomes possible to return to the correct number. Also, if it becomes possible to more accurately infer only a character string that seems to be an amount, a part of the amount extraction process can be simplified, and both the accuracy and the processing performance are improved.
[0157] In step S521 of FIG. 33, the control unit 11 (extraction processing unit 113) searches for a character string of numbers. Specifically, the control unit 11 searches for a character string including a specific character (amount mark) representing an amount such as "¥" or "yen" and numbers. When the control unit 11 extracts the character string (S522: Yes), it transfers the process to step S523. On the other hand, when the control unit 11 does not extract the character string (S522: No), it transfers the process to step S529.
[0158] In step S523, the control unit 11 determines whether the character string is located on the same line as the keyword. If the control unit 11 determines that the character string is located on the same line as the keyword (S523: Yes), in step S524, the control unit 11 determines the character string as a candidate for the amount character string. After processing step S524, the control unit 11 proceeds to step S525.
[0159] On the other hand, if the control unit 11 determines that the character string is not located on the same line as the keyword (S523: No), it determines in step S526 whether the character string is located on the line immediately below the keyword. If the control unit 11 determines that the character string is located on the line immediately below the keyword (S526: Yes), it extracts the character string as a candidate amount character string in step S527. If the control unit 11 determines that the character string is not located on the line immediately below the keyword (S526: No), it proceeds to step S522.
[0160] If there are other character strings extracted in step S521 (S528: No), the control unit 11 returns the process to step S522 and executes the above-described process for the next character string. When the control unit 11 has completed the above-described process for all character strings, the control unit 11 shifts the process to step S525.
[0161] If the control unit 11 does not extract the character string containing the specific character (amount mark) and a number in the search process of step S521 (S522: No), it searches for a character string of numbers that does not contain the specific character (amount mark) in step S529. If the control unit 11 extracts the character string (S530: Yes), it shifts the process to step S531. On the other hand, if the control unit 11 does not extract the character string (S530: No), it shifts the process to step S522.
[0162] In step S531, the control unit 11 determines whether the character string is located on the same line as the keyword. If the control unit 11 determines that the character string is located on the same line as the keyword (S531: Yes), in step S532, the control unit 11 extracts the character string as a candidate for an amount character string. After processing step S532, the control unit 11 proceeds to step S533.
[0163] On the other hand, if the control unit 11 determines that the character string is not located on the same line as the keyword (S531: No), it determines in step S534 whether the character string is located on the line immediately below the keyword. If the control unit 11 determines that the character string is located on the line immediately below the keyword (S534: Yes), it extracts the character string as a price character string candidate in step S532. If the control unit 11 determines that the character string is not located on the line immediately below the keyword (S534: No), it proceeds to step S522.
[0164] If there are other character strings extracted in step S529 (S533: No), the control unit 11 returns the process to step S522 and executes the above-described process for the next character string. When the control unit 11 has completed the above-described process for all character strings, the control unit 11 shifts the process to step S525.
[0165] In step S525, the control unit 11 determines the character string for the amount. Specifically, the control unit 11 determines that a character string for the amount satisfies the following first to third conditions as a character string for the amount. The first condition is that "at least two of the first three characters of the character string must be numbers." Note that in the amount numeric conversion process described below, characters that are misrecognized by OCR and are to be corrected into numeric characters are also counted as numbers, and currency unit symbols (such as ¥, JPY, US$, and ?) are also counted as 1. For example, "¥1,000" is counted as two characters ("¥1,"), "580 yen" is counted as three characters ("580"), and "JPY1,234" is counted as two characters ("JPY" is one character).
[0166] The second condition is that the string must not contain symbols used in telephone numbers, fax numbers, postal codes, dates, or times. These symbols include left and right parentheses (, ) , hyphens (-), postal codes (〒), slashes ( / ), and colons (:). For this reason, for example, the following are not eligible for the price string: telephone number "XXX(XXX)XXXX," fax number "XXX-XXX-XXXX," postal code "〒000-0000" (there is also a condition that the hyphen is in the third character (if there is no postal code) or the fourth character (if there is a postal code)), date "XXXX / XX / XX," and time "00:00:00."
[0167] The third condition is that "there is a certain number of numeric characters compared to the total number of characters in the string." For a string such as "¥100,000,000" that represents an amount, the number of numeric characters is defined based on the total number of characters, taking into account the "¥" mark and comma ",". For example, the condition for an amount of 1 to 3 digits is "the difference between the total number of characters and the number of numeric characters is 3 characters or less, and the number of numeric characters is 1 to 3 characters." The condition for an amount of 4 to 6 digits is "the difference between the total number of characters and the number of numeric characters is 4 characters or less, and the number of numeric characters is 4 to 6 characters." The condition for an amount of 7 to 9 digits is "the difference between the total number of characters and the number of numeric characters is 5 characters or less, and the number of numeric characters is 7 to 9 characters." The condition for an amount of 10 digits is "the number of numeric characters is 10 characters (here, up to 10 digits are defined, but more than 10 digits can be defined in the same way)." For example, in the case of "¥100,000,000", the total number of characters is 12, the number of numeric characters is 9, the difference is 3, and the condition of a 7- to 9-digit amount is met. The conditions for the other digits are defined in the same way. The control unit 11 determines that a character string that meets all of the above conditions is a monetary amount character string.
[0168] Next, in step S535, the control unit 11 converts the amount string into a numerical value. Specifically, the control unit 11 converts the numeric characters into a numerical value and ignores the characters other than the numeric characters and the minus sign (such as "¥", ",", "yen", etc.). For example, the control unit 11 ignores the comma "," in "1,234" and converts it to "1234" as a numerical value. Specifically, the control unit 11 converts and adds the numbers in order from the leftmost number. At that time, the number converted after multiplying the value already substituted into the variable by 10 as follows is added. X = 0 X = X × 10 + 1 X = X × 10 + 2 X = X × 10 + 3 X = X × 10 + 4 Through the above processing, X becomes 1234.
[0169] In addition, since there may be cases of decimal notation, in the case of a string ending with two characters of "「.」 + number" (the second decimal place), it is output up to the second decimal place as decimal notation. For example, "68.99" is numerically converted from "$68.99" as follows (the number converted from the digit after the decimal point is multiplied by 0.1 and 0.01 and then added). X = 0 X = X × 10 + 6 [[ID=2二十二]]X = X × 10 + 8 二十五]]X = X + 9 × 0.1 X = X + 9 × 0.01 Through the above processing, X becomes 68.99.
[0170] In addition, since the OCR may mis-convert a number into a non-numeric character, for a target string for amount numerical conversion, if there are some specific non-numeric characters, the control unit 11 corrects them to numeric characters if possible and then performs numerical conversion. For example, "l" in "¥l,000" is corrected to "1", and then it is numerically converted to "1000". For example, characters similar to "0" such as "O (oh)", "○ (circle symbol)", etc., characters similar to "1" such as "l (ell)", "I (eye)", etc., characters similar to "2" such as "Z", etc., in this way, only the necessary numbers are defined in advance as the conversion target according to the tendency of OCR mis-conversion for "0" to "9".
[0171] In addition, for misrecognized OCR characters that may have multiple correction candidates, the control unit 11 checks whether any of the second through fifth candidate characters in the OCR result, which has up to five candidates, contain any numeric characters. If any numeric characters are found, the control unit 11 corrects them to the first numeric character found and then performs numeric conversion. If none of the first five candidates contain any numeric characters, the control unit 11 corrects them to a predefined numeric character. For example, since the letter "S" can be "6," "8," or "9," the control unit 11 checks whether any of the numbers are found in the second through fifth candidates in the OCR result. If any are found, the control unit 11 corrects them to the numeric character. If none are found, the control unit 11 assumes the predefined "8." For example, if the first candidate is "S," the second candidate is "s," the third candidate is "6," the fourth candidate is "8," and the fifth candidate is "9," the control unit 11 corrects the "S" to "6." If the first candidate is "S," the second candidate is "s," and the third candidate is "$," the control unit 11 corrects the "S" to the predefined correction digit "8."
[0172] After executing the amount search process, the control unit 11 determines whether or not any amounts have been extracted by the amount search process in step S516 shown in Fig. 32. If any amounts have been extracted by the amount search process (S516: Yes), the control unit 11 adds a predetermined score to the character string of the corresponding amount in step S517 and determines whether the amount includes tax. Note that the control unit 11 may set a high added score for the character string of the amount corresponding to the keyword "total" and a low added score for the character string of the amount corresponding to the keyword "amount."
[0173] In this way, the control unit 11 executes the processing of the rule 1. After executing the processing of the rule 1, the control unit 11 executes the processing of the rule 2.
[0174] In the rule 2, the control unit 11 searches for all character strings that seem to be amounts, regardless of the keywords "total" and "amount." The process of the rule 2 is the same as the process shown in Fig. 33, except that the processes for determining the positional relationship with the keywords (S523, S526, S531, S534) are omitted.
[0175] The control unit 11 executes the processing of rule 3 following rule 2. In rule 3, the control unit 11 searches for a keyword related to "consumption tax" in step S513 of Fig. 32, and searches for a character string of an amount corresponding to the keyword by the amount search processing of Fig. 33.
[0176] Furthermore, the control unit 11 executes the processing of rule 4 following rule 3. In accordance with rule 4, the control unit 11 searches for a keyword related to "change" in step S513 of Fig. 32, and searches for a character string of an amount corresponding to the keyword by the amount search processing of Fig. 33.
[0177] Furthermore, the control unit 11 executes the processing of rule 5 following rule 4. In rule 5, the control unit 11 searches for keywords related to the excluded numerical value in step S513 of Fig. 32, and searches for a numeric character string corresponding to the keyword by the amount search processing of Fig. 33. The excluded numerical value includes characters such as the product number, type number, register number, person in charge number, receipt number, and NO.
[0178] The control unit 11 identifies and outputs the total amount based on the character string searched (extracted) by the processing of the rules 1 to 5. For example, when the control unit 11 extracts one candidate amount string corresponding to a keyword related to "total," it determines the amount string as the total amount and ends the processing.
[0179] In addition, when the control unit 11 extracts candidate change amount strings, it calculates the total amount based on the maximum and second largest values among the multiple candidate change amount strings and the change, determines the calculated total amount, and ends the process.
[0180] In addition, when the control unit 11 extracts a candidate amount string for consumption tax, it calculates the total amount based on the largest and second largest values among the multiple candidate amount strings and the consumption tax, determines the calculated total amount, and ends the processing.
[0181] As described above, the control unit 11 extracts keywords related to "total" using two types of character recognition processing (non-learning OCR, which is OCR for receipt mode, and AI learning OCR), and extracts keywords related to "amount" if the keyword "total" is not present. Furthermore, if the control unit 11 extracts the keyword, it preferentially uses the AI learning OCR to extract the amount.
[0182] The control unit 11 also preferentially extracts character strings of numbers that include the specific character (amount mark). The control unit 11 also preferentially extracts character strings that are located on the same line as the keyword or one line below. The control unit 11 also preferentially extracts character strings of numbers that do not include the specific character (amount mark) when the character string is located on the same line as the keyword but far from the keyword. The control unit 11 also extracts multiple candidate amount strings for keywords related to "amount."
[0183] Furthermore, if the control unit 11 is unable to extract a keyword, it extracts multiple character strings of numbers. In this case, if the specific character is present, the control unit 11 may extract only the character string of numbers. Furthermore, if the specific character is not present, the control unit 11 may extract a character string consisting of only numbers.
[0184] For example, in the example shown in Figure 34(a), the amount "¥3,024" located on the same line as the keyword "total" is extracted as the total amount. In the example shown in Figure 34(b), the amount "¥1,512" located on the line one line below the keyword "current total" is extracted as the total amount. In the example shown in Figure 34(c), the amount "¥3,024" located on the same line as the keyword "amount" but farther away is extracted as the total amount. In the example shown in Figure 34(d), since there are no keywords related to "total" or "amount," the numeric string "880 (tax included)" is extracted as the total amount.
[0185] When the control unit 11 determines the total amount as described above (S506 to S508), it shifts the process to step S9 (FIG. 4). The control unit 11 outputs the total amount of the extraction target each time it acquires an input image.
[0186] As described above, the control unit 11 extracts a character string of a specific amount (total amount) corresponding to the target item name (such as "total", "amount", etc.) from the image data P1 of the document (form). Specifically, the control unit 11 determines whether the form is a receipt or a general form based on the image data P1, extracts the item name related to the target item name from the image data P1, and in the search range (the same line, one line below, etc.) according to the form (receipt or general form) of the form based on the item name, extracts the amount corresponding to the item name, and outputs the specific amount based on the extracted amount.
[0187] When the form is a receipt, the control unit 11 extracts the amount in the search range of the same line as the first line where the item name is arranged or in the search range of the second line one line below the first line (see FIG. 34).
[0188] On the other hand, when the form is a general form, the control unit 11 extracts the amount in the search range on the right side of the first position where the item name is arranged or in the search range below the first position (see FIG. 30).
[0189] Further, when the control unit 11 extracts the item name related to the total from the image data P1 and extracts a character string including the number included in the search range and a specific character (such as "¥", "yen", etc.) representing the amount, the control unit 11 may output the number as the specific amount.
[0190] Further, when the control unit 11 does not extract the item name related to the total from the image data P1, extracts the item name related to the amount, and extracts a plurality of character strings including the numbers and the specific characters included in the search range, the control unit 11 outputs the calculated value calculated based on the plurality of numbers as the specific amount (steps S507 and S508 in FIG. 31).
[0191] In addition, if the control unit 11 does not extract the item names related to the total and amount from the image data P1 and extracts multiple strings of numbers from the image data P1, it may output a number among the multiple numbers that is identified based on the numbers related to consumption tax and change as the specific amount (steps S507 and S508 in Figure 31).
[0192] Furthermore, for example, the control unit 11 extracts keywords related to the total amount, such as "total" or "current total," from the recognized character strings in the character string information of image data P1 input as a receipt image. It then extracts numeric character strings located on the same line as the keyword or on the line immediately below as candidate amounts corresponding to the keywords. When extracting keywords, the control unit 11 searches for keywords using the recognition results of both the non-learning OCR (receipt mode OCR) and the general-purpose AI learning OCR. Meanwhile, when extracting candidate amounts, the control unit 11 prioritizes the recognition results of the general-purpose AI learning OCR to search for numeric character strings with an amount symbol (e.g., ¥, Yen) and extract candidate amounts. If the control unit 11 cannot find a keyword related to the total amount, it searches for keywords related to amounts other than the total amount to extract candidate amounts. When extracting candidate amounts, the control unit 11 assigns a higher score to candidate amounts extracted using keywords related to the total amount than to candidate amounts extracted using keywords other than the total amount. Furthermore, the control unit 11 adds a score based on indicators such as whether the items are on the same line, whether they have an amount symbol, whether they are the largest number, etc. The control unit 11 finally outputs the amount candidate with the largest score as the amount to be extracted (total amount).
[0193] Furthermore, if there is no character string of numbers with an amount symbol, the control unit 11 searches for character strings of numbers located as far away as possible on the same line to extract possible amounts.
[0194] Furthermore, if a keyword is not found, the control unit 11 extracts a candidate amount without a keyword. When extracting candidate amounts, the control unit 11 prioritizes searching for a string of numbers with a monetary symbol by using the recognition results of the AI learning OCR to extract the amount. If a string of numbers with a monetary symbol is not found, the control unit 11 also extracts a string of numbers without a monetary symbol. For example, if a keyword is not found, the control unit 11 may extract all strings containing numbers as candidate amounts, and if a string of numbers with a monetary symbol is found, all strings of numbers without a monetary symbol may be canceled. Furthermore, if the control unit 11 extracts only strings without a monetary symbol, it may cancel a string of numbers containing characters other than digits and commas. Furthermore, the control unit 11 may cancel a string of numbers extracted using keywords such as "cash" or "deposit." Furthermore, if there is a string of numbers that is the same as the amount calculated using a keyword and keywords such as "consumption tax" or "change," the control unit 11 may add a score to the string so that it becomes a candidate amount.
[0195] This configuration allows accurate extraction of the total amount of the extraction target for each general form and receipt. Furthermore, even if the layout of the receipt is different, accurate extraction of the total amount of the extraction target can be achieved.
[0196] [Company information extraction process] In the company information extraction process (step S9 in FIG. 4), the control unit 11 extracts company information such as the issuer, destination, registration number, etc. included in the form. Specifically, if the form is a receipt, the control unit 11 extracts the company information using receipt mode OCR processing, and if the form is a general form, the control unit 11 extracts the company information using general form mode OCR processing.
[0197] For example, the control unit 11 extracts character strings related to company information such as "store name," "corporation name," "name," and "registration number" through OCR processing according to the type of form, and determines the company information to be extracted based on an evaluation value according to the position of the character string in the image data P1. After step S9, the control unit 11 proceeds to step S10.
[0198] [Extraction result output process] After performing the date extraction process (FIG. 17), amount extraction process (FIG. 29, FIGS. 31 to 33), and company information extraction process described above, the control unit 11 outputs the extraction results in step S10 (see FIG. 4). Here, the control unit 11 (output processing unit 116) outputs the date (issue date, etc.), amount (total amount, etc.), and company information (issuer, etc.) of general documents (quotations, purchase orders, invoices, delivery notes, receipts), and outputs the date (issue date, etc.), amount (total amount, etc.), and company information (store name, etc.) of receipts.
[0199] 35 and 36, the output processing unit 116 displays an extraction result page P2 on the operation terminal 2. Fig. 35 shows the extraction result page P2 for a general form (quotation), and Fig. 36 shows the extraction result page P2 for a receipt.
[0200] The output processing unit 116 may also display an "OK" button K1 on the extraction result page P2. On the extraction result page P2, the user checks whether the results extracted from the image data P1 are correct, and presses the "OK" button K1 if they are determined to be correct. When the user presses the "OK" button K1, the control unit 11 registers the extraction results in a database.
[0201] The control unit 11 may also have a function of learning the corrections made when a rectangular area (character string rectangular area) is corrected. For example, the control unit 11 displays a "Learn" button K2 in a selectable manner when the user corrects the size and position of a rectangular area on the extraction result page P2, adds a new rectangular area, or corrects the extraction result. When the user presses the "Learn" button K2 on the extraction result page P2, the control unit 11 learns the corrections and corrects (re-learns) the learning model. This improves the accuracy of character string recognition (object recognition) for the image data P1.
[0202] In the above-described embodiment, the image processing device 1 alone corresponds to the information processing system according to the present disclosure, but the information processing system according to the present disclosure may be configured by the image processing device 1 and the operation terminal 2. For example, when the components of the image processing device 1 and the operation terminal 2 cooperate to share and execute the amount extraction process, a system including the multiple components that execute the process corresponds to the information processing system according to the present disclosure. Furthermore, when the operation terminal 2 executes the amount extraction process, the operation terminal 2 alone may constitute the information processing system according to the present disclosure.
[0203] In another embodiment of the present disclosure, the image processing device 1 may be configured to include at least one of the form determination process, the operation mode setting process, the date extraction process, the amount extraction process, and the company information extraction process. For example, the image processing device 1 may include the form determination process, the date extraction process, and the amount extraction process, as well as a conventional company information extraction process.
[0204] [Disclosure Supplement 1 (Document Judgment Processing)] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.
[0205] <Appendix 1> an acquisition processing unit that acquires a character string rectangle surrounding a character string included in a document based on image data of the document; a determination processing unit that determines the format of the document based on the layout density of the character string rectangles in an area corresponding to the size of the document; an extraction processing unit that extracts target characters corresponding to predetermined items included in the document from the image data in an operation mode according to a determination result of the determination processing unit; An information processing system comprising:
[0206] <Appendix 2> The document is a receipt-type document or a general form-type document other than the receipt. 10. The information processing system of claim 1.
[0207] <Appendix 3> the determination processing unit determines whether the document is a receipt-format document or a general form-format document other than the receipt-format document based on the arrangement density of the character string rectangles in the vertical direction of the area. 10. The information processing system of claim 2.
[0208] <Appendix 4> The determination processing unit calculating a total height obtained by adding up the vertical heights of the plurality of character string rectangles relative to a reference vertical height of the area as the placement density; If the placement density is equal to or greater than a threshold, the document is determined to be the receipt-format document, and if the placement density is less than the threshold, the document is determined to be the general form-format document. 10. The information processing system of claim 3.
[0209] <Appendix 5> The determination processing unit calculating a total height obtained by adding up the vertical heights of the plurality of character string rectangles relative to a reference vertical height of the area as the placement density; If the placement density is less than a threshold value, the document is determined to be a document in the general form format; If the layout density is equal to or greater than a threshold value, the aspect ratio of the area matches a predetermined value, and the estimated number of characters in the horizontal direction of the area is equal to or greater than a predetermined number of characters, the document is determined to be in the general form format; If the layout density is equal to or greater than a threshold value and the aspect ratio of the region does not match the predetermined value, or if the layout density is equal to or greater than a threshold value and the estimated number of characters is less than the predetermined number of characters, the document is determined to be a receipt-format document. 5. The information processing system according to claim 3 or 4.
[0210] <Appendix 6> the determination processing unit estimates the vertical and horizontal sizes of the region based on the maximum and minimum values of horizontal position coordinates and the maximum and minimum values of vertical position coordinates of each of the plurality of character string rectangles; An information processing system according to any one of Supplementary Notes 2 to 5.
[0211] <Appendix 7> determining the area based on a difference between the size of the image data and the size of an angle of view estimated from the arrangement positions of the plurality of character string rectangles; An information processing system according to any one of Supplementary notes 2 to 6.
[0212] <Appendix 8> the determination processing unit calculates, as the reference height, a length from a first character string rectangle located at the top position of the area to a second character string rectangle located at the bottom position of the area, excluding a height of a non-character string rectangle located between the first character string rectangle and the second character string rectangle; An information processing system according to any one of Supplementary Notes 2 to 7.
[0213] <Appendix 9> the determination processing unit calculates the placement densities corresponding to the character string rectangles included in divided regions obtained by dividing the region in the horizontal direction; An information processing system according to any one of appendices 1 to 8.
[0214] [Disclosure Note 2 (Operation Mode Setting Process)] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.
[0215] <Appendix 1> a determination processing unit that determines the format of the document based on image data of the document; a first extraction processing unit that extracts a category that is an item that represents a type of the document from among a plurality of items included in the document; a setting processing unit that sets an operation mode for extracting target characters included in the plurality of items to a first operation mode or a second operation mode based on the determination result of the determination processing unit and the classification extracted by the first extraction processing unit; a second extraction processing unit that extracts the target character from the image data in an operation mode set by the setting processing unit; An information processing system comprising:
[0216] <Appendix 2> the determination processing unit determines whether the document is a receipt-format document or a general form-format document other than the receipt format; the setting processing unit sets the operation mode for extracting the target characters of at least one of the date, amount, issuer, and destination included in the plurality of items to a general form mode or a receipt mode based on the determination result and the classification. 10. The information processing system of claim 1.
[0217] <Appendix 3> When the determination processing unit determines that the document is a document in the general form format and the first extraction processing unit extracts an estimate, an order form, a delivery note, an invoice, or a receipt as the classification, the setting processing unit sets the operation mode to the general form mode. 10. The information processing system of claim 2.
[0218] <Appendix 4> When the determination processing unit determines that the document is the receipt-format document and the first extraction processing unit extracts an estimate, an order form, a delivery note, or an invoice as the category, the setting processing unit sets the operation mode to the general form mode. 4. The information processing system according to claim 2 or 3.
[0219] <Appendix 5> When the determination processing unit determines that the document is the receipt-format document and the first extraction processing unit extracts a receipt, or an estimate, an invoice, an order form, a delivery note, or other documents other than receipts as the classification, the setting processing unit sets the operation mode to the receipt mode. An information processing system according to any one of appendixes 2 to 4.
[0220] <Appendix 6> When the determination processing unit determines that the document is a document in the general form format or the receipt format, and the first extraction processing unit extracts other documents other than estimates, invoices, purchase orders, delivery notes, and receipts as the classification, the setting processing unit sets the operation mode to the receipt mode. An information processing system according to any one of Supplementary Notes 2 to 5.
[0221] <Appendix 7> the second extraction processing unit executes a first process of extracting the target characters from the image data in the general form mode and a second process of extracting the target characters from the image data in the receipt mode, and determines an output result corresponding to the category based on the results of the first process and the second process; An information processing system according to any one of Supplementary notes 2 to 6.
[0222] <Appendix 8> When the determination processing unit determines that the document is the receipt format document, the first extraction processing unit extracts the category from the image data on which a first character recognition process corresponding to the receipt mode has been performed; When the determination processing unit determines that the document is a document in the general form format, the first extraction processing unit extracts the category from the image data on which a second character recognition process corresponding to the general form mode has been performed. An information processing system according to any one of Supplementary Notes 1 to 7.
[0223] [Disclosure Note 3 (Date Extraction Processing)] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.
[0224] <Appendix 1> an acquisition processing unit that acquires a result of a character recognition process executed on a character string included in an image of a document; a calculation processing unit that calculates an evaluation value according to the format of the document for each of a plurality of date character string candidates included in the result of the character recognition processing; an extraction processing unit that extracts a date string based on the evaluation value of each of the plurality of date string candidates calculated by the calculation processing unit; an output processing unit that outputs the date character string extracted by the extraction processing unit as the date of the document; An information processing system comprising:
[0225] <Appendix 2> The document is a receipt or a general document other than the receipt. 10. The information processing system of claim 1.
[0226] <Appendix 3> the calculation processing unit sets the evaluation value corresponding to the distance from the center position of the image to the position of the date character string candidate, to the date character string candidate. 10. The information processing system of claim 2.
[0227] <Appendix 4> If the document is the receipt, the calculation processing unit sets a higher evaluation value for the date string candidate as the distance becomes shorter; 4. The information processing system according to claim 2 or 3.
[0228] <Appendix 5> When the document is a general form, the calculation processing unit sets a higher evaluation value for the date string candidate as the distance increases; An information processing system according to any one of Supplementary notes 2 to 4.
[0229] <Appendix 6> the calculation processing unit adds a predetermined value to the evaluation value of the date string candidate when a character related to a date is included within a predetermined range from a position of the date string candidate in the image; An information processing system according to any one of Supplementary Notes 2 to 5.
[0230] <Appendix 7> the calculation processing unit subtracts a predetermined value from the evaluation value of the date string candidate when a character unrelated to a date is included within the predetermined range from the position of the date string candidate in the image; An information processing system according to any one of Supplementary notes 2 to 6.
[0231] <Appendix 8> the calculation processing unit sets the evaluation value corresponding to the position coordinates of the date string candidate in the image to the date string candidate. An information processing system according to any one of Supplementary Notes 2 to 7.
[0232] <Appendix 9> the output processing unit displays the plurality of date string candidates in descending order of the evaluation value; An information processing system according to any one of appendices 1 to 8.
[0233] [Disclosure Note 4 (Amount Extraction Processing)] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.
[0234] <Appendix 1> An information processing system that extracts a character string of a specific amount corresponding to a target item name to be extracted from image data of a document, a determination processing unit that determines whether the document is a first document in a first format or a second document in a second format based on the image data; an extraction processing unit that extracts item names related to the target item names from the image data and extracts amounts corresponding to the item names within a search range according to the document format based on the item names; an output processing unit that outputs the specific amount based on the amount extracted by the extraction processing unit; An information processing system comprising:
[0235] <Appendix 2> When the document of the image data is a receipt, the extraction processing unit extracts the amount in the search range of the same row as the first row in which the item name is arranged, or in the search range of the second row one row below the first row; 10. The information processing system of claim 1.
[0236] <Appendix 3> When the document of the image data is a general form other than the receipt, the extraction processing unit extracts the amount in the search range to the right of the first position where the item name is arranged or in the search range below the first position; 3. The information processing system according to claim 1 or 2.
[0237] <Appendix 4> When the extraction processing unit extracts the item name related to the total from the image data and extracts a character string including a number included in the search range and a specific character representing an amount, the output processing unit outputs the number as the specific amount. An information processing system according to any one of appendices 1 to 3.
[0238] <Appendix 5> When the extraction processing unit does not extract the item name related to the total from the image data, but extracts the item name related to the amount, and extracts a plurality of character strings including a number included in the search range and a specific character representing the amount, the output processing unit outputs a calculated value based on the plurality of numbers as the specific amount. An information processing system according to any one of appendices 1 to 4.
[0239] <Appendix 6> When the extraction processing unit does not extract the item names related to the total amount and the amount from the image data and extracts a plurality of character strings of numbers from the image data, the output processing unit outputs, as the specified amount, a number identified based on the numbers related to the consumption tax and the change from among the plurality of numbers. An information processing system according to any one of appendices 1 to 5. [Explanation of symbols]
[0240] 1: Image processing device 2: Operation terminal 10: Image processing system 11: Control section 12: Storage section 13: Operation display section 14: Communications Department 111: Acquisition processing unit 112: Judgment processing unit 113: Extraction processing unit 114: Setting processing section 115: Calculation processing unit 116: Output processing section
Claims
1. an acquisition processing unit that acquires a character string rectangle surrounding a character string included in a document based on image data of the document; a determination processing unit that determines the format of the document based on the layout density of the character string rectangles in an area corresponding to the size of the document; an extraction processing unit that extracts target characters corresponding to predetermined items included in the document from the image data in an operation mode according to a determination result of the determination processing unit; An information processing system comprising:
2. the determination processing unit determines whether the document is a receipt-format document or a general form-format document other than the receipt-format document based on the arrangement density of the character string rectangles in the vertical direction of the area. The information processing system according to claim 1 .
3. The determination processing unit calculating a total height obtained by adding up the vertical heights of the plurality of character string rectangles relative to a reference vertical height of the area as the placement density; If the placement density is equal to or greater than a threshold, the document is determined to be the receipt-format document, and if the placement density is less than the threshold, the document is determined to be the general form-format document. The information processing system according to claim 2 .
4. The determination processing unit calculating a total height obtained by adding up the vertical heights of the plurality of character string rectangles relative to a reference vertical height of the area as the placement density; If the placement density is less than a threshold value, the document is determined to be a document in the general form format; If the layout density is equal to or greater than a threshold value, the aspect ratio of the area matches a predetermined value, and the estimated number of characters in the horizontal direction of the area is equal to or greater than a predetermined number of characters, the document is determined to be in the general form format; If the layout density is equal to or greater than a threshold value and the aspect ratio of the region does not match the predetermined value, or if the layout density is equal to or greater than a threshold value and the estimated number of characters is less than the predetermined number of characters, the document is determined to be a receipt-format document. The information processing system according to claim 2 .
5. the determination processing unit estimates the vertical and horizontal sizes of the region based on the maximum and minimum values of horizontal position coordinates and the maximum and minimum values of vertical position coordinates of each of the plurality of character string rectangles; The information processing system according to claim 1 .
6. determining the area based on a difference between the size of the image data and the size of an angle of view estimated from the arrangement positions of the plurality of character string rectangles; The information processing system according to claim 1 .
7. the determination processing unit calculates, as the reference height, a length from a first character string rectangle located at the top position of the area to a second character string rectangle located at the bottom position of the area, excluding a height of a non-character string rectangle located between the first character string rectangle and the second character string rectangle; 5. The information processing system according to claim 3 or 4.
8. the determination processing unit calculates the placement densities corresponding to the character string rectangles included in divided regions obtained by dividing the region in the horizontal direction; The information processing system according to claim 1 .
9. acquiring a character string rectangle surrounding a character string included in the document based on image data of the document; determining a format of the document based on the density of the character string rectangles in an area corresponding to the size of the document; extracting target characters corresponding to predetermined items included in the document from the image data in an operation mode according to the determination result; An information processing method executed by one or more processors.
10. acquiring a character string rectangle surrounding a character string included in the document based on image data of the document; determining a format of the document based on the density of the character string rectangles in an area corresponding to the size of the document; extracting target characters corresponding to predetermined items included in the document from the image data in an operation mode according to the determination result; An information processing program for causing one or more processors to execute the above.
Citation Information
Patent Citations
Document reading system
JP2020181369A