Information processing system, information processing method, and information processing program

The information processing system enhances character recognition accuracy in documents with text and image objects by correcting confidence levels for overlapping rectangles, addressing the accuracy decrease in existing methods.

JP2026057911APending Publication Date: 2026-04-03SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Character recognition accuracy decreases when document data includes both text and image objects, as existing methods fail to distinguish and correct overlapping confidence levels effectively.

Method used

An information processing system with an extraction processing unit, calculation processing unit, and correction processing unit to extract and correct confidence levels for text and image objects, adjusting overlap corrections to enhance accuracy.

Benefits of technology

Improves character recognition accuracy in document data containing both text and image objects by correcting confidence levels when rectangles overlap, ensuring accurate output of candidate strings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026057911000001_ABST
    Figure 2026057911000001_ABST
Patent Text Reader

Abstract

This invention provides an information processing system, an information processing method, and an information processing program capable of improving the accuracy of character recognition in document data that includes embedded text objects and image objects. [Solution] The image processing device 1 includes an extraction processing unit 112 that extracts text information and image objects from the target items of document data; a calculation processing unit 114 that calculates a first confidence level for the text information and a second confidence level for the image objects, representing the confidence level of the target items; a correction processing unit 115 that corrects the first confidence level when the first string rectangle of the text information and the second string rectangle of the image object overlap each other; and an output processing unit 116 that outputs candidate strings for the target items based on the corrected first confidence level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] ,

[0006] , , , , , ,

[0005] , , , , , ,

[0001] This disclosure relates to a technique for performing image processing such as character recognition on an input image.

Background Art

[0002] Conventionally, a technique for performing character recognition (OCR processing) on document data such as forms has been known. For example, based on whether the document data is an electronic document generated by an application program while retaining text information or an electronic document generated by reading an image with a document reading device (such as a scanner), a technique for determining whether to perform optical character recognition processing on the electronic document is known (see, for example, Patent Document 1). <00,00010>

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Here, for example, a document generated by an application program may include a text object (embedded text) and an image object. In this case, the characters (embedded text) recognized from the text object are not necessarily correct, and there is a problem that the character recognition accuracy decreases with a method of uniformly extracting the embedded text.

[0005] An object of this disclosure is to provide an information processing system, an information processing method, and an information processing program capable of improving the character recognition accuracy of document data including a text object and an image object of embedded text.

Means for Solving the Problems

[0006] An information processing system according to one aspect of this disclosure comprises an extraction processing unit, a calculation processing unit, a correction processing unit, and an output processing unit. The extraction processing unit extracts text information and image objects from the items to be extracted from document data. The calculation processing unit calculates a first confidence level for the text information and a second confidence level for the image objects, representing the confidence level of the items to be extracted. The correction processing unit corrects the first confidence level when the first string rectangle of the text information and the second string rectangle of the image objects overlap each other. The output processing unit outputs candidate strings of the items to be extracted based on the first confidence level corrected by the correction processing unit.

[0007] An information processing method relating to another aspect of this disclosure is a method in which one or more processors perform the following actions: extracting text information and image objects from an item to be extracted from document data; calculating a first confidence level for the text information and a second confidence level for the image object, representing the confidence level of the item to be extracted; correcting the first confidence level when the first string rectangle of the text information and the second string rectangle of the image object overlap each other; and outputting candidate strings for the item to be extracted based on the corrected first confidence level.

[0008] An information processing program according to another aspect of this disclosure is a program that causes one or more processors to perform the following actions with respect to an item to be extracted from document data: extracting text information and image objects; calculating a first confidence level for the text information and a second confidence level for the image objects, which represent the confidence level of the item to be extracted; correcting the first confidence level when the first string rectangle of the text information and the second string rectangle of the image objects overlap each other; and outputting candidate strings for the item to be extracted based on the corrected first confidence level. [Effects of the Invention]

[0009] According to this disclosure, it is possible to provide an information processing system, an information processing method, and an information processing program that can improve the accuracy of character recognition in document data that includes embedded text objects and image objects. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 is a functional block diagram showing the configuration of an image processing system according to an embodiment of this disclosure. [Figure 2] Figure 2 shows an example of a document (PDF file) according to an embodiment of this disclosure. [Figure 3] Figure 3 is a flowchart showing an example of the procedure for character extraction processing performed in the image processing apparatus according to the embodiment of this disclosure. [Figure 4] Figure 4 is a flowchart showing an example of the procedure for character extraction processing performed in the image processing apparatus according to the embodiment of this disclosure. [Figure 5] Figure 5 shows a specific example of the confidence level calculated in the image processing apparatus according to the embodiment of this disclosure. [Figure 6] Figure 6 shows a specific example of the confidence level calculated in the image processing apparatus according to the embodiment of this disclosure. [Figure 7] Figure 7 shows a specific example of the confidence level calculated in the image processing apparatus according to the embodiment of this disclosure. [Figure 8] Figure 8 shows an example of displaying string information in an image processing apparatus according to the present disclosure. [Figure 9] Figure 9 shows an example of the display of string information in an image processing apparatus according to the present disclosure. [Figure 10] Figure 10 is a flowchart showing another example of the procedure for character extraction performed in the image processing apparatus according to the embodiment of this disclosure. [Figure 11] Figure 11 is a flowchart showing another example of the procedure for character extraction performed in the image processing apparatus according to the embodiment of this disclosure. [Modes for carrying out the invention]

[0011] The embodiments of this disclosure will be described below with reference to the attached drawings. Note that the following embodiments are merely examples of the embodiments of this disclosure and do not limit the technical scope of this disclosure.

[0012] Figure 1 is a block diagram showing the configuration of an image processing system 10 according to an embodiment of this disclosure. The image processing system 10 includes an image processing device 1 and an operation terminal 2. The image processing device 1 and the operation terminal 2 are connected to each other via a network N1 (e.g., the Internet, LAN, etc.). The image processing system 10 may include a plurality of operation terminals 2.

[0013] In the image processing system 10, the image processing device 1 acquires document data (image data) such as forms transmitted from the operation terminal 2 and extracts desired strings (strings to be managed) from the document data. For example, the operation terminal 2 transmits document data (such as a PDF file) generated by scanning paper forms such as invoices, quotations, delivery notes, purchase orders, receipts, and other documents to the image processing device 1. The operation terminal 2 also creates a document file of the form based on user operations using, for example, a document creation application, and transmits the document file as image data (for example, searchable PDF data (image + text data)) to the image processing device 1. When the image processing device 1 receives the document data transmitted from the operation terminal 2, it performs various processes described later on the document data to extract strings of managed items (items to be extracted) contained in the form. For example, the image processing device 1 extracts the category (type) of each form, the date of each form, the amount (total amount, etc.), company information (issuer, recipient, registration number, etc.), etc. Furthermore, the image processing device 1 registers the extracted strings in a predetermined database. For example, each time the image processing device 1 acquires image data of an invoice, it extracts strings related to the contents of the invoice (e.g., issue date, invoice amount, issuer, etc.) from the image data and registers them in a database that manages invoices. Similarly, each time the image processing device 1 acquires image data of a receipt, it extracts strings related to the contents of the receipt (e.g., issue date, total amount, issuer, etc.) from the image data and registers them in a database that manages receipts. This allows each document to be saved and managed as electronic data. The image processing device 1 also outputs the extracted strings to an operation terminal 2 or the like to present the character recognition results to the user.

[0014] Image processing system 10 is an example of an information processing system in this disclosure. Note that the information processing system in this disclosure may consist of the image processing device 1 alone.

[0015] [Image processing device 1] As shown in FIG. 1, the image processing apparatus 1 includes a control unit 11, a storage unit 12, an operation display unit 13, a communication unit 14, and the like. The image processing apparatus 1 may be one or more cloud servers, or one or more physical servers.

[0016] The communication unit 14 is a communication interface for connecting the image processing apparatus 1 to the network N1 by wire or wirelessly and performing data communication according to a predetermined communication protocol with the operation terminal 2 via the network N1. The network N1 is composed of, for example, the Internet, a LAN, or the like.

[0017] The operation display unit 13 is a user interface including a display unit such as a liquid crystal display or an organic EL display for displaying various information, and an operation unit such as a mouse, a keyboard, or a touch panel for receiving operations.

[0018] The storage unit 12 is a non-volatile storage unit such as a HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory for storing various information. The storage unit 12 stores a control program for causing the control unit 11 to execute the character extraction process described later. For example, the control program is non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, and is read by a reading device (not shown) such as a CD drive or a DVD drive provided in the image processing apparatus 1 and stored in the storage unit 12. Note that the control program may be distributed from a cloud server and stored in the storage unit 12.

[0019] In addition, the storage unit 12 stores document data (such as a PDF file) of documents such as forms acquired from the operation terminal 2.

[0020] Figure 2 shows an invoice as an example of a form (document data P1). As shown in Figure 2, the invoice includes strings such as the document type ("Invoice"), issue date, issuer's contact information (address, telephone number, fax number, person in charge), invoice amount, product name, quantity, standard price, discount amount, subtotal, consumption tax, and total amount. For example, a user uploads document data P1 (PDF file), which is an image (PDF) of a document created using a document creation application on the operating terminal 2, to the image processing device 1. When the control unit 11 acquires the document data P1 of the invoice, it stores it in the storage unit 12. The document data P1 in Figure 2 is data obtained by converting a document created by a document creation application into an image (PDF), and includes text information (character data, embedded text) and image objects (character image, seal image, etc.). Hereafter, the text information of document data P1 will also be referred to as embedded text.

[0021] In another embodiment, the control unit 11 may acquire the document file of the form created on the operation terminal 2 and store the document file in the storage unit 12.

[0022] The control unit 11 includes control devices such as a CPU, ROM, and RAM. The CPU is a processor that performs various arithmetic operations. The ROM stores control programs such as a BIOS and OS in advance to allow the CPU to perform various operations. The RAM stores various information and is used as a temporary storage memory (work area) for the various operations performed by the CPU. The control unit 11 controls the image processing device 1 by executing various control programs stored in advance in the ROM or storage unit 12 using the CPU.

[0023] As shown in Figure 1, the control unit 11 includes various processing units such as an acquisition processing unit 111, an extraction processing unit 112, a recognition processing unit 113, a calculation processing unit 114, a correction processing unit 115, and an output processing unit 116. The control unit 11 functions as these various processing units by executing various processes according to the control program. Some or all of the processing units included in the control unit 11 may be composed of electronic circuits. The control program may be a program that causes multiple processors to function as these various processing units.

[0024] The acquisition processing unit 111 acquires document data (document files). Specifically, the acquisition processing unit 111 acquires document data (PDF files, image files, etc.) from the operation terminal 2. In addition, the acquisition processing unit 111 acquires document data (image files) generated by scanning paper forms in an image forming apparatus (scanner, etc.).

[0025] The extraction processing unit 112 extracts string rectangles from the document data. Furthermore, if the document data contains embedded text and image objects, the extraction processing unit 112 extracts the string rectangle of the embedded text (first string rectangle) and the string rectangle of the image object (second string rectangle).

[0026] The recognition processing unit 113 performs character recognition processing on the image object. This character recognition processing includes OCR pre-processing, OCR processing, and OCR post-processing. For example, in OCR pre-processing, the recognition processing unit 113 performs processes such as vertical correction, tilt correction, background removal, and seal impression removal. In OCR processing, the recognition processing unit 113 recognizes characters and identifies the position and size of the character string rectangle. In OCR post-processing, the recognition processing unit 113 adjusts the size of the character string rectangle (the corrected second character string rectangle), corrects characters (corrects characters based on their relationship with surrounding character information, etc.). The character recognition processing can be performed using well-known techniques.

[0027] The calculation processing unit 114 calculates a confidence score (certainty score) for each recognized character, based on the recognized content, the position of the character, and the relationship between the character and surrounding characters, representing the likelihood of the extracted item. The calculation processing unit 114 calculates the confidence score for each character string rectangle. The calculation processing unit 114 also calculates the confidence score (second confidence score, second certainty score) for the character recognized by OCR processing for the second character string rectangle of the image object, and the confidence score (first confidence score, first certainty score) for the character recognized for the first character string rectangle of the embedded text. The calculation processing unit 114 also outputs character information, character position, size of the character string rectangle, and character confidence score. The size of the character string rectangle may be expressed as height and width, or the character position may be expressed as coordinates of the starting point (top left) and ending point (bottom right).

[0028] The correction processing unit 115 corrects the first confidence level corresponding to the string in the first string rectangle when the first string rectangle and the second string rectangle overlap each other. Specifically, the correction processing unit 115 corrects the first confidence level to a value greater than the second confidence level. The correction processing unit 115 may also correct the first confidence level when the area occupancy rate of the image object in the document data is less than a threshold.

[0029] The output processing unit 116 outputs the string information. Based on the first confidence level corrected by the correction processing unit 115, the output processing unit 116 outputs candidate strings for the items to be extracted. For example, the output processing unit 116 displays multiple string information in descending order of confidence level. The output processing unit 116 also displays the extraction results for the required items. The specific processing details of each processing unit will be described later.

[0030] [Character extraction process] In the character extraction process, the control unit 11 acquires document data P1 (document file), extracts text information and image objects from the target items of the document data P1, and calculates the confidence level of the target items: a first confidence level (first degree of certainty) for the text information and a second confidence level (second degree of certainty) for the image objects. Furthermore, the control unit 11 corrects the first confidence level of the text information if the string rectangle of the text information and the string rectangle of the image object overlap. Then, the control unit 11 outputs candidate strings for the target items based on the corrected first confidence level of the text information. Figure 3 shows an example of the procedure for the character extraction process.

[0031] This disclosure can be understood as a character extraction method that performs one or more steps included in the character extraction process. Furthermore, one or more steps included in the character extraction process described herein may be omitted as appropriate. In addition, the execution order of each step in the character extraction process may differ to the extent that similar effects are produced. Furthermore, although this description uses the case in which the control unit 11 of the image processing device 1 performs each step in the character extraction process as an example, in other embodiments, one or more processors may distribute and execute each step in the character extraction process. In addition, when the control unit 11 acquires document data P1 from each of the multiple operation terminals 2 (including scanners), it is possible to execute the character extraction process in parallel for each piece of document data P1.

[0032] In step S1, the control unit 11 (acquisition processing unit 111) acquires document data (PDF files, image files, etc.) from the operation terminal 2.

[0033] In step S2, the control unit 11 determines whether the acquired document data is a file to be processed (e.g., a PDF or image file). If the acquired document data is a file to be processed (S2: Yes), the control unit 11 proceeds to step S3. On the other hand, if the acquired document data is not a file to be processed (S2: No), the control unit 11 terminates the character extraction process.

[0034] In step S3, the control unit 11 determines whether the acquired document data is a PDF file or not. If the acquired document data is a PDF file (S3: Yes), the control unit 11 proceeds to step S4. On the other hand, if the acquired document data is not a PDF file, i.e., an image file (S3: No), the control unit 11 proceeds to step S5.

[0035] In step S4, the control unit 11 determines whether the area occupancy rate (area coverage rate) of the image objects included in the PDF file is equal to or greater than a threshold. For example, the control unit 11 determines whether the image objects occupy 95% or more of the total area of ​​the page of the document data P1 (see Figure 2). If the control unit 11 determines that the area occupancy rate of the image objects is equal to or greater than the threshold (S4: Yes), the process proceeds to step S5. On the other hand, if the control unit 11 determines that the area occupancy rate of the image objects is less than the threshold (S4: No), the process proceeds to step S11 (see Figure 4).

[0036] For example, in document data generated by scanning a paper form using an image forming apparatus (such as a scanner), the entire page is composed of a single image object. In some cases, text information recognized by OCR processing may be embedded in the document data generated by the scanner function. Since the area occupancy rate of the image object in the document data generated by these scanning functions (such as PDF files) exceeds a threshold, the control unit 11 proceeds to step S5.

[0037] In contrast, the document data generated by the document creation application on the operating terminal 2 may consist only of text data, or it may consist of text data and image objects. For example, the document data shown in Figure 2 consists of embedded text data and image objects such as characters and seals. In this case, the area occupancy rate of the image objects is less than the threshold, so the control unit 11 moves the processing to step S11.

[0038] In step S5, the control unit 11 (recognition processing unit 113) performs character recognition processing (OCR pre-processing, OCR processing, and OCR post-processing). Specifically, the recognition processing unit 113 performs OCR pre-processing such as vertical correction, tilt correction, background removal, and seal impression removal, then recognizes the characters (OCR processing), identifies the position and size of the character string rectangle, and then performs post-processing such as adjusting the size of the character string rectangle and correcting the characters.

[0039] In step S6, the control unit 11 performs item string determination processing. Specifically, first, the control unit 11 (extraction processing unit 112) extracts the strings of the necessary items (items to be extracted). For example, the control unit 11 extracts the type of document, date, amount (excluding tax / including tax), and issuer / issuer information (company name, address, telephone number, registration number). Next, the control unit 11 (calculation processing unit 114) calculates the confidence level (certainty) of the recognized character based on the content of the recognized character, the position of the character, and the relationship between the character and surrounding characters. Next, the control unit 11 outputs string information. Specifically, the control unit 11 outputs string information including character information, character position, size of the string rectangle, and character confidence level.

[0040] In step S7, the control unit 11 performs item string selection processing. Specifically, the control unit 11 (output processing unit 116) ranks the string information according to the confidence level, sets the string information in the order of first candidate, second candidate, ... from highest rank, and outputs it.

[0041] In step S8, the control unit 11 (output processing unit 116) displays the extraction results. Specifically, the control unit 11 displays the string information in candidate order and accepts user selection operations.

[0042] In step S9, the control unit 11 (output processing unit 116) outputs the extraction result. Specifically, the control unit 11 outputs the string information selected in a predetermined format according to the user's instructions.

[0043] As described above, if the document data is an image file (S3: No), or if the area occupancy rate of image objects contained in the PDF file of the document data is greater than or equal to the threshold (S4: Yes), the control unit 11 performs OCR processing on the PDF file to extract the string information of the necessary items. Conversely, if the area occupancy rate of image objects contained in the PDF file of the document data is less than the threshold (S4: No), the control unit 11 performs the following processing.

[0044] In step S11, the control unit 11 analyzes the PDF file. Specifically, the control unit 11 analyzes the PDF file and extracts embedded text and non-embedded text (image objects, etc.). If the document data is an IMG file (image file), it is output as image data, and embedded text is output as empty data.

[0045] In step S12, the control unit 11 performs rendering on the PDF file. Specifically, the control unit 11 converts the PDF file into an image to generate image data for character recognition.

[0046] In step S13, the control unit 11 performs character recognition processing. Specifically, the control unit 11 performs the above-described OCR pre-processing, OCR processing, and OCR post-processing to recognize characters and determine the position and size of the character string rectangle. Here, the control unit 11 performs OCR processing on the entire PDF file.

[0047] In step S14, the control unit 11 performs the second item string determination process. For example, the control unit 11 calculates the confidence level (accuracy) of a required item based on the content of the character recognized by the OCR process, the position of the character, and the relationship between the character and surrounding characters. The control unit 11 then outputs string information including the character information, the position of the character, the size of the string rectangle, and the confidence level of the character. After step S14, the control unit 11 moves the process to step S17.

[0048] In step S15, the control unit 11 extracts embedded text from the PDF file.

[0049] In step S16, the control unit 11 performs the first item string determination process. For example, the control unit 11 (calculation processing unit 114) calculates the confidence level (certainty) of a required item based on the content of the embedded text character, the position of the character, and the relationship between the character and surrounding characters. The control unit 11 then outputs string information including the character information, the character's position, the size of the string rectangle, and the character's confidence level. After step S16, the control unit 11 moves the process to step S17.

[0050] When the control unit 11 obtains the string information (S14), which is the character recognition result of the OCR process, and the string information (S16), it executes the following item string selection processes (S17~S21).

[0051] In step S17, the control unit 11 identifies the first string rectangle of embedded text and the second string rectangle recognized by OCR processing that are close together. For example, the control unit 11 identifies the first string rectangle and the second string rectangle that overlap each other in at least part.

[0052] In step S18, the control unit 11 determines whether the degree of overlap between the first string rectangle and the second string rectangle is greater than or equal to a threshold. Specifically, the control unit 11 determines whether the degree of overlap (overlap ratio) between the first string rectangle and the second string rectangle is 20% or more. If the degree of overlap is greater than or equal to the threshold (S18: Yes), the control unit 11 determines the process in step S19. On the other hand, if the degree of overlap is less than the threshold (S18: No), the control unit 11 determines the process in step S21.

[0053] In step S19, the control unit 11 compares the first confidence level (first accuracy) of the embedded text in the first character rectangle with the second confidence level (second accuracy) of the recognized character in the second character rectangle and determines whether the first confidence level is less than or equal to the second confidence level. If the control unit 11 determines that the first confidence level is less than or equal to the second confidence level (S19: Yes), it proceeds to step S20. On the other hand, if the control unit 11 determines that the first confidence level is greater than the second confidence level (S19: No), it proceeds to step S21.

[0054] In step S20, the control unit 11 (correction processing unit 115) corrects the first confidence level. Specifically, the control unit 11 corrects the first confidence level to a value greater than the second confidence level.

[0055] For example, Figure 5 shows the string recognized by OCR processing and the string of embedded text. Also, the degree of overlap of the string rectangles of each string is 20% or more. For example, if the calculation processing unit 114 calculates the confidence level of the recognized characters by OCR processing (second confidence level) as "95" and the confidence level of the recognized characters in the embedded text (first confidence level) as "90", the correction processing unit 115 corrects the first confidence level to a value of "100", which is greater than the second confidence level.

[0056] For example, Figure 6 shows the strings recognized by OCR processing (OCR1, OCR2) and the strings of the embedded text. The degree of overlap of the string rectangles of each string is 20% or more. For example, if the calculation processing unit 114 calculates the confidence level of the recognized characters by OCR processing 1 (second confidence level) as "90", the confidence level of the recognized characters by OCR processing 2 (second confidence level) as "82", and the confidence level of the recognized characters of the embedded text (first confidence level) as "85", the correction processing unit 115 corrects the first confidence level to "91", which is greater than the second confidence level.

[0057] In the example shown in Figure 7, the degree of overlap between the string recognized by OCR processing and the embedded text string is less than 20%. In this case, the correction processing unit 115 does not correct the first confidence score ("96") and the second confidence score ("92") calculated by the calculation processing unit 114.

[0058] In this manner, the correction processing unit 115 corrects the first confidence level when the degree of overlap between the first string rectangle and the second string rectangle is greater than or equal to a threshold. The correction processing unit 115 also corrects the first confidence level of the embedded text, which is text information, to a value greater than the second confidence level of the character recognition result obtained by OCR processing on the image object. After step S20, the control unit 11 moves the processing to step S21.

[0059] In step S21, the control unit 11 (output processing unit 116) outputs string information. Specifically, the control unit 11 outputs string information including character information, character position, size of the string rectangle, and character confidence level. The control unit 11 also ranks the string information according to the confidence level and outputs the string information in the order of first candidate, second candidate, and so on, from highest to lowest rank. After step S21, the control unit 11 moves the process to step S8 (see Figure 3).

[0060] In step S8, the output processing unit 116 displays string information including confidence level, recognized string, and recognition method ("embedded text", "OCR") in order of candidate ranking, as shown in Figure 8, for example.

[0061] In step S9, the control unit 11 (output processing unit 116) accepts the user's selection operation regarding the string information and outputs the string information selected by the user as the result of extracting the target item.

[0062] In another embodiment, as shown in Figure 9, the output processing unit 116 may display the recognized string with the first candidate rank. Alternatively, the output processing unit 116 may display a pull-down menu to allow the user to select string information.

[0063] As described above, the control unit 11 executes the character extraction process each time it acquires a document file (document data P1).

[0064] As described above, the image processing device 1 according to this embodiment extracts text information and image objects from the target items of document data (document file), calculates the confidence level (accuracy) of the target items as a first confidence level for the text information and a second confidence level for the image objects, and corrects the first confidence level when the first string rectangle of the text information and the second string rectangle of the image objects overlap each other. The image processing device 1 then outputs candidate strings for the target items based on the corrected first confidence level. Specifically, the image processing device 1 corrects the first confidence level when the degree of overlap between the first string rectangle and the second string rectangle is greater than or equal to a threshold. Also, the image processing device 1 corrects the first confidence level to a value greater than the second confidence level when the first confidence level is less than or equal to the second confidence level. The image processing device 1 then outputs candidate strings based on the corrected first and second confidence levels. For example, the image processing device 1 outputs the text information as the candidate string.

[0065] In this way, for a single extraction target item, the priority of each candidate character is determined based on the confidence level of each candidate character obtained from the results of multiple different methods (OCR processing, embedded text extraction). For example, the results of embedded text are output first. This allows for the output of appropriate candidate characters by correcting the priority of embedded text to a higher level, for example, when garbled candidate characters are included. Similarly, even when images such as seal impressions are included, appropriate candidate characters can be output by correcting the priority of embedded text to a higher level. Therefore, with the above configuration, it is possible to improve the accuracy of character recognition in document data that includes both embedded text objects and image objects.

[0066] [Other embodiments] The image processing system 10 of this disclosure is not limited to the embodiments described above, but may also be the following embodiments. Figures 10 and 11 show other examples of the procedure for character extraction processing. In the following, detailed explanations of the same processes as those shown in Figures 3 and 4 will be omitted as appropriate.

[0067] In step S51, the control unit 11 acquires document data (PDF files, image files, etc.) from the operation terminal 2.

[0068] In step S52, the control unit 11 determines whether the acquired document data is the document data to be processed (e.g., a form). If the acquired document data is the document data to be processed (S52: Yes), the control unit 11 proceeds to step S53. On the other hand, if the acquired document data is not the document data to be processed (S52: No), the control unit 11 terminates the character extraction process.

[0069] In step S53, the control unit 11 determines whether the acquired document data is a PDF file. If the acquired document data is a PDF file (S53: Yes), the control unit 11 proceeds to step S54. On the other hand, if the acquired document data is not a PDF file, i.e., an image file (S53: No), the control unit 11 proceeds to step S71. The processing in steps S71 to S74 is the same as the processing in steps S5, S6, S8, and S9 in Figure 3.

[0070] In step S54, the control unit 11 analyzes the PDF file. Specifically, the control unit 11 analyzes the PDF file and extracts embedded text and non-embedded text (image objects, etc.). If the document data is an IMG file (image file), it is output as image data, and embedded text is output as empty data.

[0071] In step S55, the control unit 11 performs rendering on the PDF file. Specifically, the control unit 11 converts the PDF file into an image to generate image data for character recognition.

[0072] In step S56, the control unit 11 performs character recognition processing. Specifically, the control unit 11 performs OCR pre-processing, OCR processing, and OCR post-processing to recognize characters and determine the position and size of the character string rectangle.

[0073] In step S57, the control unit 11 performs the second item string determination process. For example, the control unit 11 calculates the confidence level (accuracy) of a required item based on the content of the character recognized by the OCR process, the position of the character, and the relationship between the character and surrounding characters. The control unit 11 then outputs string information including the character information, the position of the character, the size of the string rectangle, and the confidence level of the character. After step S57, the control unit 11 moves the process to step S64 (see Figure 11).

[0074] In step S58, the control unit 11 extracts the embedded text.

[0075] In step S59, the control unit 11 performs the first item string determination process. For example, the control unit 11 calculates the confidence level (certainty) of a required item based on the content of the embedded text character, the position of the character, and the relationship between the character and surrounding characters. The control unit 11 then outputs string information including the character information, the character's position, the size of the string rectangle, and the character's confidence level. After step S59, the control unit 11 proceeds to step S60.

[0076] In step S60, the control unit 11 determines whether the area occupancy rate (area coverage rate) of the image objects included in the PDF file is equal to or greater than a threshold (e.g., 95%). If the control unit 11 determines that the area occupancy rate of the image objects is equal to or greater than the threshold (S60: Yes), the process proceeds to step S61 (see Figure 11). On the other hand, if the control unit 11 determines that the area occupancy rate of the image objects is less than the threshold (S60: No), the process proceeds to step S64 (see Figure 11).

[0077] In step S61, the control unit 11 determines whether the degree of overlap between the first string rectangle and the second string rectangle is greater than or equal to a threshold (e.g., 20%). If the degree of overlap is greater than or equal to the threshold (S61: Yes), the control unit 11 determines the process in step S62. On the other hand, if the degree of overlap is less than the threshold (S61: No), the control unit 11 determines the process in step S67.

[0078] In step S62, the control unit 11 compares the first confidence level (first accuracy) of the embedded text in the first character rectangle with the second confidence level (second accuracy) of the recognized character in the second character rectangle, and determines whether the first confidence level is less than or equal to the second confidence level. If the control unit 11 determines that the first confidence level is less than or equal to the second confidence level (S62: Yes), it proceeds to step S63. On the other hand, if the control unit 11 determines that the first confidence level is greater than the second confidence level (S62: No), it proceeds to step S64.

[0079] In step S63, the control unit 11 corrects the first confidence level. Specifically, the control unit 11 corrects the first confidence level to a value greater than the second confidence level.

[0080] In step S64, the control unit 11 outputs string information. Specifically, the control unit 11 outputs string information including character information, character position, string rectangle size, and character confidence level. The control unit 11 also ranks the string information according to the confidence level and outputs the string information in the order of first candidate, second candidate, and so on, from highest to lowest rank.

[0081] In step S65, the control unit 11 displays string information including confidence level, recognition string, and recognition method ("embedded text", "OCR") in order of candidate ranking.

[0082] In step S66, the control unit 11 accepts the user's selection operation regarding the string information and outputs the string information selected by the user as the result of extracting the target item.

[0083] If the degree of overlap in step S61 is less than the threshold (S61: No), in step S67, the control unit 11 executes a predetermined process. For example, the control unit 11 executes one of the following processes: (1) not to use embedded text, (2) to maintain the confidence level as is, or (3) to lower the confidence level. After step S67, the control unit 11 moves the process to step S64. The control unit 11 may execute the character extraction process as described above.

[0084] The control unit 11 of the image processing device 1 controls the entire image processing device 1. The control unit 11 realizes various functions by reading and executing various programs stored in the storage unit 12 (for example, storage or ROM). The control unit 11 may be realized by one or more control devices / arithmetic units (CPU (Central Processing Unit), SoC (System on a Chip)). The control unit 11 may also be composed of one or more control circuits (electronic circuits).

[0085] [Disclosure Note] The following is an overview of the disclosures extracted from the above-described embodiments. Note that each configuration and processing function described below can be selected and combined as desired.

[0086] <Note 1> Regarding the items to be extracted from the document data, an extraction processing unit extracts text information and image objects, A calculation processing unit calculates a first confidence level for the text information and a second confidence level for the image object, with respect to the confidence level of the extracted items. A correction processing unit that corrects the first accuracy when the first string rectangle of the text information and the second string rectangle of the image object overlap each other, An output processing unit outputs candidate strings for the extraction target item based on the first accuracy corrected by the correction processing unit, An information processing system equipped with the following features.

[0087] <Note 2> The correction processing unit corrects the first accuracy to a value greater than the second accuracy if the first accuracy is less than or equal to the second accuracy. The output processing unit outputs the candidate string based on the corrected first accuracy and the second accuracy. The information processing system described in Appendix 1.

[0088] <Note 3> The output processing unit outputs the text information as the candidate string. The information processing system described in Appendix 1 or 2.

[0089] <Note 4> The correction processing unit corrects the first accuracy when the degree of overlap between the first string rectangle and the second string rectangle is greater than or equal to a first threshold. An information processing system as described in any of the appendices 1 to 3.

[0090] <Note 5> The output processing unit displays the multiple candidate strings in order of increasing accuracy. An information processing system as described in any of the appendices 1 to 4.

[0091] <Note 6> The correction processing unit corrects the first accuracy of the embedded text, which is the text information, to a value greater than the second accuracy of the character recognition result obtained by OCR processing on the image object. An information processing system as described in any of the appendices 1 to 5.

[0092] <Note 7> The calculation processing unit calculates the accuracy of the recognized character based on the content of the recognized character, the position of the character, and the relationship between the character and surrounding characters. An information processing system as described in any of the appendices 1 to 6.

[0093] <Note 8> The correction processing unit corrects the first accuracy when the area occupancy rate of the image object in the document data is less than the second threshold. An information processing system as described in any of the appendices 1 to 7.

[0094] <Note 9> Regarding the items to be extracted from the document data, the process involves extracting text information and image objects. Regarding the confidence level of the extracted items, the first confidence level of the text information and the second confidence level of the image object are calculated. When the first string rectangle of the text information and the second string rectangle of the image object overlap, the first accuracy is corrected. Based on the corrected first accuracy, the candidate strings of the items to be extracted are output, An information processing method performed by one or more processors.

[0095] <Note 10> Regarding the items to be extracted from the document data, the process involves extracting text information and image objects. Regarding the confidence level of the extracted items, the first confidence level of the text information and the second confidence level of the image object are calculated. When the first string rectangle of the text information and the second string rectangle of the image object overlap, the first accuracy is corrected. Based on the corrected first accuracy, the candidate strings of the items to be extracted are output, An information processing program for causing one or more processors to execute, or a non-temporary computer-readable recording medium on which such information processing program is recorded. [Explanation of symbols]

[0096] 1: Image processing device 2: Operating terminal 10: Image Processing System 11: Control Unit 12: Storage section 13: Operation display section 14: Communications Department 111: Acquisition Processing Unit 112: Extraction Processing Unit 113: Recognition Processing Unit 114: Calculation Processing Unit 115: Correction Processing Unit 116: Output Processing Unit

Claims

1. Regarding the items to be extracted from the document data, an extraction processing unit extracts text information and image objects, A calculation processing unit calculates a first confidence level for the text information and a second confidence level for the image object, with respect to the confidence level of the extracted items. A correction processing unit that corrects the first accuracy when the first string rectangle of the text information and the second string rectangle of the image object overlap each other, An output processing unit outputs candidate strings for the extraction target items based on the first accuracy corrected by the correction processing unit, An information processing system equipped with the following features.

2. The correction processing unit corrects the first accuracy to a value greater than the second accuracy if the first accuracy is less than or equal to the second accuracy. The output processing unit outputs the candidate string based on the corrected first accuracy and the second accuracy. The information processing system according to claim 1.

3. The output processing unit outputs the text information as the candidate string. The information processing system according to claim 1.

4. The correction processing unit corrects the first accuracy when the degree of overlap between the first string rectangle and the second string rectangle is greater than or equal to a first threshold. The information processing system according to claim 1.

5. The output processing unit displays the multiple candidate strings in order of increasing accuracy. The information processing system according to claim 1.

6. The correction processing unit corrects the first accuracy of the embedded text, which is the text information, to a value greater than the second accuracy of the character recognition result obtained by OCR processing on the image object. The information processing system according to claim 1.

7. The calculation processing unit calculates the accuracy of the recognized character based on the content of the recognized character, the position of the character, and the relationship between the character and surrounding characters. The information processing system according to claim 1.

8. The correction processing unit corrects the first accuracy when the area occupancy rate of the image object in the document data is less than the second threshold. The information processing system according to claim 1.

9. Regarding the items to be extracted from the document data, the process involves extracting text information and image objects. Regarding the confidence level of the extracted items, the first confidence level of the text information and the second confidence level of the image object are calculated. When the first string rectangle of the text information and the second string rectangle of the image object overlap, the first accuracy is corrected. Based on the corrected first accuracy, the candidate strings of the items to be extracted are output, An information processing method performed by one or more processors.

10. Regarding the items to be extracted from the document data, the process involves extracting text information and image objects. Regarding the confidence level of the extracted items, the first confidence level of the text information and the second confidence level of the image object are calculated. When the first string rectangle of the text information and the second string rectangle of the image object overlap, the first accuracy is corrected. Based on the corrected first accuracy, the candidate strings of the items to be extracted are output, An information processing program that causes one or more processors to execute.

Citation Information

Patent Citations

  • Information processing device, information processing system, and program

    JP2023047133A