Information processing device, information processing method, and program

The information processing device addresses OCR accuracy issues on semi-standard forms by dividing and correcting form images to align them with a form type image, ensuring high-accuracy character recognition.

JP7680669B2Active Publication Date: 2025-05-21CANON MARKETING JAPAN INC +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021052848
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-26
Publication Date
2025-05-21
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing character recognition systems face accuracy issues when dealing with semi-standard forms where the positions and sizes of details vary slightly, leading to distortion and decreased OCR performance.

Method used

An information processing device that includes an acquisition means for form images, a division means for dividing the form image into areas, a correction means for aligning these areas with a form type image, and a character recognition means for performing OCR on the corrected image.

Benefits of technology

Enables high-accuracy OCR on semi-standard forms by accurately aligning and recognizing characters despite variations in position and size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680669000001
    Figure 0007680669000001
  • Figure 0007680669000002
    Figure 0007680669000002
  • Figure 0007680669000003
    Figure 0007680669000003
Patent Text Reader

Abstract

To provide a mechanism capable of accurately performing OCR (character recognition) even in a form such as a semi-standard form which is a form having slightly different positions and sizes of the details of the form.SOLUTION: A form image to be subjected OCR is acquired, and a form type image is divided according to a predetermined condition. For each area in the form image corresponding to a divided area in the divided form type image, the form image is corrected, and the OCR is performed on the corrected form image.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a technique for performing character recognition on semi-standard forms. [Background technology]

[0002] Automatic character recognition systems are used mainly in the financial and insurance industries to efficiently process multiple handwritten forms. These character recognition systems use OCR technology, but with OCR, the coordinates and character type information of the items to be read on the form are generally defined in advance (OCR definition information), and the target area is cut out from the input form image and subjected to character recognition processing.

[0003] As a pre-processing step for OCR, a method is often used in which the scanned form image is corrected to the coordinate position of a template form type image in order to process according to predefined OCR definition information. One method uses image features. Specifically, the features of the scanned form image and the template form type image are calculated, similar feature points are matched, and image transformation processing (for example, projective transformation or affine transformation) is performed so that the feature point coordinates are the same. (See Patent Document 1) [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2020-107272 A Summary of the Invention [Problem to be solved by the invention]

[0005] When the position correction of feature points described in Patent Document 1 is performed on a semi-standard form (a form in which the position and size of details of the form vary slightly), the scanned image may be distorted, and the OCR accuracy may decrease. In addition, the OCR accuracy may decrease due to a shift in the coordinates of the OCR area. In order to solve such problems, a method of registering all semi-standard forms as templates is conceivable, but registering many semi-standard forms as templates is time-consuming.

[0006] Therefore, the present invention provides a mechanism that enables accurate OCR (optical character recognition) even for semi-standard forms, which are forms in which the positions and sizes of details of the form vary slightly. [Means for solving the problem]

[0007] The information processing device of the present invention comprises an acquisition means for acquiring a form image to be subjected to OCR, a division means for dividing a form type image in accordance with predetermined conditions, a correction means for correcting the form image for each area in the form image that corresponds to a divided area in the form type image divided by the division means, and a character recognition means for performing OCR on the form image corrected by the correction means. Effect of the Invention

[0008] According to the present invention, it is possible to perform OCR (character recognition) with high accuracy even for input forms, such as semi-standard forms, in which the positions and sizes of details of the form vary slightly. [Brief description of the drawings]

[0009] [Figure 1] 1 is a diagram illustrating an example of a system configuration of an information processing system according to an embodiment of the present invention. [Diagram 2] 2 is a diagram illustrating an example of the hardware configuration of a document scanner, a client PC, and a server PC according to an embodiment of the present invention. [Diagram 3]2 is a diagram illustrating functional configurations of a document scanner, a client PC, and a server PC in the embodiment of the present invention. [Figure 4] 11 is a flowchart showing an example of a process for registering a form type according to an embodiment of the present invention. [Diagram 5] 5 is a flowchart showing an example of a process for OCR of a form in an embodiment of the present invention. [Figure 6] 11 is a flowchart showing an example of processing for performing position correction processing on an OCR-processed form in an embodiment of the present invention. [Figure 7] 1 is a flowchart illustrating an example of an image segmentation process using an image histogram in an embodiment of the present invention. [Figure 8] 11 is a flowchart showing an example of processing for integrating divided images in an embodiment of the present invention. [Figure 9] FIG. 13 is a diagram showing an example of a screen for setting an OCR definition for a form type. [Figure 10] FIG. 11 is a diagram illustrating an example of a coordinate information table for managing coordinate information of divided regions. [Figure 11] 13 is a diagram illustrating an example of a data table for managing OCR definition setting data. FIG. [Figure 12] 13 is a diagram showing an example of a feature point information table for managing feature point information of document-type images; FIG. [Figure 13] FIG. 2 illustrates an example of a storage that manages and stores various images. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0011] 1 is a diagram showing an example of the configuration of an information processing system according to the present invention. The information processing system including an information processing device according to the embodiment of the present invention is composed of a document scanner 101, a client PC 102, a server PC 103, and a LAN 104. The document scanner 101, the client PC 102, and the server PC 103 are connected via the LAN 104.

[0012] The document scanner 101 scans the image of the set document form. The client PC 102 starts the scan process of the document scanner 101 and instructs the server PC 103 to start OCR. When the search results are received from the server PC 103, the server PC 103 notifies the user. When the server PC 103 receives the instruction to start OCR from the client PC 102, it executes OCR processing. The LAN 104 is a network that connects the document scanner 101 with the client PC 102 and the server PC 103 to ensure communication.

[0013] The document scanner 101 may be any device capable of reading documents such as forms, questionnaires, etc. The client PC 102 may be a desktop personal computer, a notebook computer, a tablet terminal, a smartphone, etc.

[0014] Although an example of an information processing system according to an embodiment of the present invention has been described above, the configuration of this information processing system may of course be replaced with other devices, etc., as long as the system configuration can solve the problems of the present invention.

[0015] Next, a hardware configuration of an information processing apparatus applicable to the document scanner 101, the client PC 102, and the server PC 103 shown in Fig. 1 will be described with reference to Fig. 2. The hardware configuration of the server PC 103 will be described below.

[0016] FIG. 2 is a block diagram showing an example of a hardware configuration of the information processing device of the present invention.

[0017] As shown in FIG. 2, the information processing device includes a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an input controller 205, a video controller 206, a memory controller 207, and a communication I / F controller 208 connected via a system bus 204.

[0018] The CPU 201 comprehensively controls each device and controller connected to a system bus 204 .

[0019] ROM 202 or external memory 211 holds the BIOS (Basic Input / Output System) and OS (Operating System), which are control programs executed by CPU 201, computer-readable and executable programs for realizing this information processing method, and various necessary data (including data tables).

[0020] The RAM 203 functions as a main memory, a work area, etc., for the CPU 201. The CPU 201 loads programs and the like required for executing processes from the ROM 202 or the external memory 211 into the RAM 203, and executes the loaded programs to realize various operations.

[0021] The input controller 205 controls input from an input device such as an input device 209 such as a keyboard or a pointing device such as a mouse (not shown). If the input device is a touch panel, the user can give various instructions by pressing (touching with a finger or the like) icons, cursors, or buttons displayed on the touch panel.

[0022] The touch panel may be a touch panel capable of detecting positions touched by multiple fingers, such as a multi-touch screen.

[0023] The video controller 206 controls display on an external output device such as a display 210. The display also includes the display of a notebook computer integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. In addition, for devices capable of receiving the above-mentioned touch operation, an input device is also provided.

[0024] The video controller 206 is capable of controlling a video memory (VRAM) for display control, and can use part of the RAM 203 as the video memory area, or can provide a separate dedicated video memory.

[0025] The memory controller 207 controls access to the external memory 211. As the external memory, an external storage device (hard disk) for storing a boot program, various applications, font data, user files, edit files, and various data, a flexible disk (FD), or a compact flash (registered trademark) memory connected to a PCMCIA card slot via an adapter, etc., can be used.

[0026] The communication I / F controller 208 connects and communicates with external devices via a network, and executes communication control processing on the network. For example, communication using TCP / IP, telephone lines such as ISDN, and 3G lines for mobile phones are possible.

[0027] The CPU 201 enables display on the display 210 by, for example, executing a process of expanding (rasterizing) an outline font in a display information area in the RAM 203. The CPU 201 also enables a user to give instructions using a mouse cursor (not shown) on the display 210.

[0028] Various programs for implementing the present invention, which will be described later, are stored in the ROM 202 or the external memory 211, and are executed by the CPU 201 by being loaded into the RAM 203 as required.

[0029] Furthermore, definition files, various information, tables, etc. used when executing the above programs are also stored in the external memory 211, and a detailed description of these will be given later.

[0030] Next, the meanings of the terms used in the description of the embodiments of the present invention will be briefly explained. Here, the form type image, position correction, and semi-standard form will be explained.

[0031] In this explanation, the form type image refers to a blank form image that is registered in advance when performing form OCR. The combination of this form type image and OCR definition settings (settings related to the items and areas to be OCR-targeted and the type of characters to be entered) is called a form type template.

[0032] Next, we will explain position correction. The position, angle, and scale of an image (form image) obtained by scanning a filled-out form is misaligned compared to the form type image. Therefore, when applying the coordinate information of the form type template, the position, angle, and scale of the form image to be OCRed are corrected so that it is perfectly aligned with the form type image.

[0033] Next, a semi-standard form will be described. A semi-standard form is a form in which the contents and rough positions of the items to be read contained in the registered form type image are the same, but the size and position of the items are slightly different. For example, in a case where an original Excel file of a form is distributed to multiple business locations and used, each business location may make small changes to the file, resulting in the creation of a form with a slightly different layout. Such a form is called a semi-standard form in contrast to a registered form. This concludes the explanation of terms used in the description of the present invention.

[0034] Next, the functions of the document scanner 101, the client PC 102, and the server PC 103 in the embodiment of the present invention will be described with reference to FIG.

[0035] FIG. 3 is a block diagram showing a functional configuration of a system according to an embodiment of the present invention.

[0036] The document scanner 101 includes a scan image transmission unit 111 that scans and transmits an image of a set document. The scan image transmission unit transmits the scan image to an image data receiving unit 121 of the client PC .

[0037] The client PC 102 includes an image data receiving unit 121 , an image display unit 122 , an OCR setting information input unit 123 , a form type registration instruction unit 124 , a form OCR processing instruction unit 125 , and a processing result display unit 126 .

[0038] The image data receiving unit 121 is a receiving unit that receives a scanned image sent from the scanned image transmitting unit 111 of the document scanner 101. This scanned image is a form-type image. The image display unit 122 is a display unit that displays the image received by the image data receiving unit 121, i.e., the form-type image.

[0039] The OCR setting information input unit 123 is an input unit for inputting OCR definition setting information. OCR is performed based on the OCR definition settings inputted in the OCR setting information input unit 123.

[0040] The form type registration instruction unit 124 is an instruction unit that instructs the registration of a form type. A form type is information that identifies the type of a form. By registering the form type, it becomes possible to identify which of the registered form types a newly read form belongs to. When a registration instruction is given by the form type registration instruction unit 124, the instruction is sent to the server PC 103, and the instruction is accepted by the form type registration acceptance unit 131 of the server PC 103, and the form type is registered. When a registration instruction is given by the form type registration instruction unit, a form type image and OCR definition setting data are sent.

[0041] The form OCR processing instruction unit 125 is an instruction unit that instructs form OCR processing. When the form OCR instruction unit 125 issues an OCR processing instruction, the form OCR reception unit 134 of the server PC 103 receives the instruction.

[0042] The processing result display unit 126 is a display unit that receives the results of the OCR processing performed by the server PC 103 from the OCR result transmission unit 136 of the server PC 103 and displays them.

[0043] The server PC 103 includes a form type registration reception unit 131 , an image processing unit 132 , a form type data storage unit 133 , a form OCR reception unit 134 , an OCR processing unit 135 , and an OCR result transmission unit 136 .

[0044] The form type registration receiving unit 131 receives a form type registration instruction from the form type registration instructing unit 124 of the client PC 102. The image processing unit 132 processes the image for which the form type registration instruction is received. The form type data storage unit 133 stores data related to the image processed by the image processing unit 132 as form type data.

[0045] The form OCR reception unit 134 receives an instruction to execute OCR processing from the form OCR processing instruction unit of the client PC 102. The OCR processing unit 135 executes OCR processing when the form OCR reception unit 134 receives the OCR instruction. The OCR result transmission unit 136 transmits the processing result performed by the OCR processing unit 135 to the processing result display unit 126 of the client PC 102. This concludes the explanation of FIG. 3.

[0046] From here, each process in the embodiment of the present invention will be explained, but the overall process flow will be explained. The process in the embodiment of the present invention consists of a form type registration process and an OCR process. The form type registration process is a process that identifies and registers the type of the form related to the received input image.

[0047] In the OCR process, matching of image features, projective transformation process, and position correction process are performed to perform OCR. Specifically, the entire form type image is read, the image is divided according to a predetermined condition, the divided areas are integrated again according to the predetermined condition, position correction is performed for each integrated area, and the position-corrected area images are recombined to generate the entire form image. The embodiment of the present invention performs OCR on this combined form image. Detailed processing will be described in the next paragraph.

[0048] Next, a form type registration process in the embodiment of the present invention will be described with reference to Fig. 4. This process is a process for managing the type of a form in the server PC 103, based on an image (form) received by the client PC 102.

[0049] The client PC 102 is responsible for the processes from step S401 to step S404.

[0050] In step S401, the image data receiving unit 121 of the client PC 102 receives the document type image transmitted by the scanned image transmitting unit of the document scanner 101.

[0051] In step S402, the form type image received by image data receiving unit 121 of client PC 102 is displayed on image display unit 122, and OCR definition settings are accepted from the user. Specifically, the screen shown in Fig. 9 is displayed, and the areas and items to be OCRed are set by accepting the designation of an area for form type image 901 displayed on the left side and the designation of read items for read items 902 displayed on the right side.

[0052] The above method of setting the OCR definition for the form type image is merely an example, and any method may be used as long as the OCR definition can be set for the form type image. Of course, the position where the form type image is displayed and the position of the read item 902 may be set as long as the OCR definition can be set for the form type image, and they do not have to be arranged as described above.

[0053] In step S403, the OCR definition setting accepted in step S402 is acquired as OCR definition setting data. The acquired OCR definition setting data is stored in an OCR definition setting data table 1100 shown in FIG.

[0054] In step S404, the document type image data accepted in step S401 and the OCR definition setting data acquired in step S403 are sent to the document type registration acceptance unit 131 of the server PC 103, and an instruction is given to start the document type registration process.

[0055] The server PC 103 is responsible for the processing from this point on.

[0056] In step S405, the document type registration receiving unit 131 of the server PC 103 receives an instruction to start the document type registration process, and receives the document type image data and OCR definition setting data from the client PC .

[0057] In step S406, an image feature vector of the entire form-type image is calculated. A known algorithm is used as the image feature vector calculation algorithm. This known algorithm represents features of characteristic local regions (feature points) in the form image as vectors, and extracts an indefinite number of feature points in one image. The calculated feature vector is stored in the feature point information table 1200. The feature vector stored in step S406 is used in step S803.

[0058] In step S407, the form type image is used to call image segmentation processing using an image histogram. The image segmentation processing will be described with reference to FIG.

[0059] Here, the image division using the image histogram in step S407 will be described. An image histogram is a graph in which the horizontal axis represents pixel values ​​and the vertical axis represents the frequency with which the pixel values ​​appear in the image.

[0060] In step S701, the document type image received by the image data receiving unit 121 of the client PC 102 is read out.

[0061] In step S702, a histogram of black pixels in the vertical direction is calculated for the entire document type image. The reason for calculating the histogram of black pixels in the vertical direction is to determine points for dividing the image.

[0062] In step S703, the image is divided at points where the number of black pixels in the histogram calculated in step S702 becomes 0. Points where the value becomes 0 are points where components such as tables are separated when they form several clusters within the form area. For example, if there are two tables above and below on an A4 portrait form, the black pixels between the tables become 0 and the image is divided. The image is divided at points where the value becomes 0 in order to divide the image at positions that would become separations when the form image is scanned horizontally. Therefore, if there are multiple points where the value becomes 0, the image is divided into multiple parts. However, if there are consecutive points with the value 0, the image is considered to be divided once.

[0063] In step S704, one of the divided images is referenced. The image referenced here is the image divided in step S703. The divided image is referenced in order to divide objects such as tables and ruled lines contained in the form into smaller chunks. For example, in a case where the upper region of the form has ruled lines arranged horizontally, while the lower region of the form has different table objects divided into left and right, the image cannot be divided into left and right regions using only the histogram of the entire image, but it can be divided into left and right regions by judging only the lower region. In other words, the image is first divided vertically into an upper, middle, and lower row, and then the upper, middle, and lower rows are each divided horizontally into left and right rows.

[0064] In step S705, a histogram of black pixels in the horizontal direction is calculated for the divided image. The reason for calculating the histogram of black pixels in the horizontal direction is to determine the points at which to divide the image.

[0065] In step S706, the image is divided at locations where the number of black pixels in the histogram calculated in step S705 becomes 0. Locations where the value becomes 0 are locations within the reference area that are breaks in components such as tables. The image is divided at locations where the value becomes 0 in the form because the image is divided at locations that become breaks when the form image is scanned from above. Therefore, if there are multiple locations where the value becomes 0, the image is divided into multiple parts. However, if there are consecutive locations with the value 0, it is divided once.

[0066] In step S707, it is determined whether an unprocessed divided image exists in the OCR definition setting. If an unprocessed image exists (Yes), the process proceeds to step S704. If an unprocessed image does not exist (No), the process proceeds to step S708.

[0067] In step S708, the coordinate information of the areas divided up to step S707 is stored in a divided area coordinate information table 1000 shown in Fig. 10. Here, the coordinate information stored is a form type ID 1001 and area coordinates 1003. This concludes the explanation of Fig. 7.

[0068] Up to this point, the process of step S407 has been explained using FIG. 7. From here on, step S408, which is the next step after step S407, will be explained.

[0069] In step S408, the divided images processed in steps S407 and up are integrated using the image feature vector and the OCR definition setting. The integration process of the divided images is performed by the server PC 103. This process will be described with reference to FIG. 8.

[0070] From here, the process of integrating the divided images carried out in step S408 will be described with reference to FIG.

[0071] In step S801, coordinate information of the divided areas is read out from the divided area coordinate information table 1000 shown in Fig. 10. The coordinate information read out here is the document type ID 1001 and area coordinates 1003.

[0072] In step S802, OCR definition setting data is read from the OCR definition setting data table 1100 shown in FIG.

[0073] In step S803, feature vector data is read from the feature point information table 1200 of the form-type image shown in Fig. 12. The feature vector data read here is the feature vector data of the entire form-type image calculated in step S406. The feature vector data is information independent of the divided regions, and in this case, the feature vector data of the entire form-type image is read.

[0074] In step S804, the area information of one divided area is referenced, which is the coordinate information of the divided area shown in FIG.

[0075] In step S805, it is determined whether an OCR target item exists in the divided area based on the coordinate information of the divided area information read in step S801 and the OCR definition setting data read in step S802. In other words, it is determined whether the area specified by the area coordinates 1003 includes the area specified by 1103. If an OCR target area exists in the divided area (Yes), the process proceeds to step S806, and if an OCR target area does not exist in the divided area (No), the process proceeds to step S807. If it is determined in step S805 that an OCR target item exists, the OCR target item existing in the divided area is obtained.

[0076] In step S806, it is determined whether four or more feature points of the feature vector exist in the divided region. If four or more feature points exist, the process proceeds to step S809. If four or more feature points do not exist, the process proceeds to step S808. The process of step S806 is performed to calculate the number of feature points of the feature vector in the divided region by matching the coordinates of the divided region with the coordinate information of all feature vector data. In this embodiment, it is determined whether four or more feature points exist, but any other value may be used as long as it is four or more. The value four is set because four is the minimum number of coordinates required to perform projective transformation. Therefore, a value of four or more may be used as long as accurate position correction can be performed.

[0077] In step S809, a divided region having four (N) or more feature points is recorded as a “single convertible region.” When the process of step S809 ends, the process proceeds to step S810.

[0078] In step S808, a divided area that does not have four (N) or more feature points is recorded as an "independently unconvertible area." An "independently unconvertible area" is an area in which position correction is necessary because an OCR target item exists, but accurate position correction cannot be performed independently because the number of feature points is insufficient. After completing the process of step S808, the process proceeds to step S810.

[0079] In step S807, the divided area in which there is no OCR target item in the read area coordinates is recorded as a "conversion-unnecessary area." A conversion-unnecessary area is an area in which there is no OCR target item and therefore no need to perform position correction. After the process of step S807 is completed, the process proceeds to step S810.

[0080] In step S810, it is determined whether or not unprocessed divided region information exists. If unprocessed divided region information exists (Yes), the process proceeds to step S804. If unprocessed divided region information does not exist (No), the process proceeds to step S811.

[0081] In step S811, the area recorded as being independently unconvertible for the divided areas processed up to step S810 is referenced. In the processing from step S811 onwards, the surrounding areas are combined to make it possible to correct the position. When the processing of step S811 is completed, the process proceeds to step S812.

[0082] In step S812, the divided areas adjacent to the divided area recorded as not being independently convertible (the divided area to be processed) are referenced. The divided areas obtained by dividing the form image are adjacent to one or more different areas. When the process of step S812 is completed, the process proceeds to step S813.

[0083] In step S813, the divided areas adjacent to the divided area to be processed are referenced to determine whether or not a conversion-unnecessary area exists. If a conversion-unnecessary area exists (Yes), the process proceeds to step S814, and if a conversion-unnecessary area does not exist (No), the process proceeds to step S815.

[0084] In step S814, the divided area to be processed is combined with the adjacent unconvertible area. Based on the combination result, the coordinate information table of the divided area shown in Fig. 10 is updated. After the process of step S814 ends, the process proceeds to step S818.

[0085] In step S815, the process refers to the divided areas adjacent to the divided area to be processed and determines whether or not an independently unconvertible area exists. If an independently unconvertible area exists (Yes), the process proceeds to step S816, and if an independently unconvertible area does not exist (No), the process proceeds to step S817.

[0086] In step S816, the divided area to be processed is combined with the adjacent independently unconvertible area. Based on the combination result, the coordinate information table of the divided area shown in Fig. 10 is updated. After the process of step S816 ends, the process proceeds to step S818.

[0087] In step S817, the divided area to be processed is combined with the adjacent independently convertible area. Based on the combination result, the coordinate information table of the divided area shown in Fig. 10 is updated. When the process of step S817 ends, the process proceeds to step S818.

[0088] In step S818, it is determined whether the number of feature points in the region has become four (N) or more as a result of the combining processes up to the previous steps. If there are four (N) or more feature points, the process proceeds to step S820. If there are not four (N) or more feature points, the process proceeds to step S819.

[0089] In step S820, a divided area having four (N) or more feature points is recorded as an independently convertible area, and the coordinate information table of the divided area shown in Fig. 10 is updated. After the process of step S820 ends, the process proceeds to step S821.

[0090] In step S819, a divided area that does not have four (N) or more feature points is recorded as an independently unconvertible area, and the coordinate information table of the divided area shown in Fig. 10 is updated. When the process of step S819 ends, the process proceeds to step S821.

[0091] In step S821, it is determined whether there is an independently unconvertible area for which the processes in steps S811 to S820 have not been executed. If there is an unprocessed independently unconvertible area (Yes), the process proceeds to step S811, where the process is executed for the unprocessed independently unconvertible area. If there is no unprocessed independently unconvertible area (No), the process of this flowchart ends.

[0092] In this way, the process of step S408 is a process of integrating the divided areas taking into consideration the number of items of OCR definition setting data included in the divided areas, and the integrated areas are integrated so that they include one or more items of OCR definition setting data.

[0093] The process of integrating the divided images using the image feature vector and the OCR definition setting in step S408 has been described above with reference to Fig. 8. This concludes the detailed description of step S408. When the process in step S408 ends, the process proceeds to step S409.

[0094] In step S409, the area information integrated in step S408 is stored in the divided area coordinate information table 1000 shown in Fig. 10. The area information stored in step S409 includes the "number of included items 1004" obtained as a result of matching with the coordinates of the OCR definition setting data in the process of step S408. The area information stored in step S409 also includes an integrated area ID 1002, which is an ID that uniquely identifies the integrated coordinate information assigned in step S409. In this way, in the process of Fig. 4, the form type is registered and form type data required for position correction of the form image is stored. This concludes the explanation of the form type registration process using Fig. 4.

[0095] Next, the form OCR process in this embodiment will be described with reference to Fig. 5. The form OCR process in Fig. 5 is performed by the client PC 102 and the server PC 103.

[0096] The client PC 102 is responsible for the processing from this point on. In step S501, the image data receiving unit 121 of the client PC 102 receives the form image (input image file) transmitted by the scanned image transmitting unit of the document scanner 101.

[0097] In step S502, the form image received in step S501 is sent to the server PC 103, and an instruction to start OCR processing is issued. Up to this point, the subject of processing has been described as the client PC 102, but from the next step S505, the subject of processing becomes the server PC 103.

[0098] In step S505, the server PC 103 receives an instruction to start OCR processing from the client PC 102, and also receives the form image.

[0099] In step S506, it is determined which type of form the form image received in step S505 is from among the preregistered form types. The method for determining which type of form the image is, for example, may use form recognition using deep learning as described in Patent Document 1. After the process of step S506 is completed, the process proceeds to step S507.

[0100] In step S507, position correction processing is performed on the form image using the information related to the form type determined in step S506. The position correction processing will be described with reference to FIG.

[0101] The position correction process in step S507 will now be described with reference to Fig. 6. This process is performed by the server PC 103.

[0102] In step S601, an image feature vector is calculated for the entire form image. After the process of step S601 is completed, the process proceeds to step S602. The image feature vector is calculated for the entire form image in order to perform matching in step S604. In S604, matching points between feature points included in the divided areas of the form type image and feature points of the form image are determined, but since the coordinates of the feature points of the form image cannot be narrowed down in advance, a feature vector is calculated for the entire image.

[0103] In step S602, coordinate information of the divided area related to the document type identified in step S506 is obtained from the divided area coordinate information table 1000 shown in Fig. 10. After the process of step S602 ends, the process proceeds to step S603.

[0104] In step S603, one of the pieces of coordinate information of the acquired divided regions is referenced. After the process of step S603 ends, the process proceeds to step S604.

[0105] In step S604, the feature vectors of the feature points present in the divided area are obtained from the feature point information table 1200 of the form type image shown in Fig. 12, and the feature vectors of the feature points in the divided area are matched with the feature vectors of the form image. After the processing of step S604 is completed, the process proceeds to step S605.

[0106] The matching algorithm used in step S604 may be a known method such as Brute-Force matcher or FlannBased matching. Brute-Force matching calculates a feature descriptor for a feature point in the first image, matches it with the feature values ​​of all feature points in the second image based on any distance calculation, and returns the corresponding feature point with the smallest distance as the matching result. Whereas Brute-Force matcher searches the entire search space brute-force, FlannBased matching searches only the space close to the feature point to be searched and returns the matching result.

[0107] In step S605, a projective transformation matrix for two images is obtained from the matched feature vectors (feature point information). A projective transformation matrix (homography transformation matrix) is a 3×3 matrix that can be projected from the image coordinates of an original image to a transformed image when a projective transformation (enlargement, reduction, rotation, translation, etc.) is performed on an image. The projective transformation matrix for two images is obtained in order to project the form image to a position that matches the form type image. By applying the projective transformation matrix obtained for the matched feature points to the entire image, the entire image can be transformed into a coordinate system in which the feature points match. A known method is used to calculate the projective transformation matrix. After the processing of step S605 is completed, the process proceeds to step S606.

[0108] In step S606, the form image is transformed using the projection transformation matrix obtained in step S605, and an image projected onto the vector space of the form type image is obtained. This image is stored in memory. The form image is transformed using the projection transformation matrix in order to align the target area to be OCRed with the coordinate position of the predefined OCR definition setting.

[0109] In step S607, the area images of the divided area coordinates are cut out from the image that has been projected in step S606 and stored in memory. The divided area coordinates are cut out from the form image and stored in order to perform the combining process in step S609.

[0110] In step S608, it is determined whether or not there is unprocessed divided area coordinate information among the divided areas registered in Fig. 10. If there is unprocessed divided area coordinate information (Yes), the process proceeds to step S603 to execute processing on the unprocessed divided area, and if there is no unprocessed divided area (No), the process proceeds to step S609.

[0111] In step S609, the divided area images stored in step S607 are combined and stored as a converted form image in the storage shown in Fig. 13. When the process of step S609 ends, the position correction process ends. This is the end of the description of Fig. 6.

[0112] Up to this point, the position correction process in step S507 has been described with reference to FIG. 6. After the process in step S507 is completed, the process proceeds to step S508.

[0113] In step S508, the form image after the position correction conversion in step S507 is read into the memory.

[0114] In step S509, the OCR definition setting data related to the document type identified in step S506 is obtained from the OCR definition setting data table 1100 shown in FIG.

[0115] In step S510, the OCR definition setting is referenced, and one of the OCR reading target items is referenced. What is referenced here is one row of the OCR definition setting data table 1100, that is, one OCR target item.

[0116] In step S511, the area coordinates of the item information to be read by OCR are referenced.

[0117] In step S512, an image is cut out from the form image after position correction conversion based on the area coordinates acquired in step S511.

[0118] In step S513, OCR reading is performed on the extracted OCR target item image. The OCR reading process may be performed within the server, or may use an API service accessible via the Internet.

[0119] In step S514, the result of the OCR recognition is stored in memory as the recognition result of the OCR target item.

[0120] In step S515, it is determined whether there is an unprocessed OCR read target item in the OCR definition setting. If there is an unprocessed OCR read target item (Yes), the process proceeds to step S510, and if there is not (No), the process proceeds to step S516.

[0121] In step S516, the form image after position correction conversion and the recognition results for all items to be read by OCR are transmitted to the client PC 102. When the processing in step S516 is completed, the processing on the server PC 103 side is completed. The processing on the server PC 103 side has been described from step S505 to step S516, but from here on, the processing on the client PC 102 will be described.

[0122] In step S516, server PC 103 transmits the form image after position correction conversion and the recognition results for all items to be read by OCR to client PC 102, and in step S503 client PC 102 receives the form image after position correction conversion and the recognition results for all items to be read by OCR.

[0123] In step S504, the OCR processing result received in step S503 is displayed on the screen. When the processing up to step S504 is completed, the processing on the client PC 102 side is terminated. This concludes the description of the form OCR processing.

[0124] As explained so far, by dividing an image using an image histogram, performing position correction processing using the combined image created by dividing the images, and then performing OCR processing, it is possible to perform OCR (character recognition) with high accuracy even on semi-standard forms.

[0125] In other words, it is only necessary to register the base document type, and there is no need to register documents where the position or size of small details differs slightly each time they are read.By appropriately adjusting the size and position of the reading area of ​​the newly read filled-out document image to match the base document that has been registered in advance, it is possible to perform OCR (optical character recognition) with high accuracy.

[0126] Next, the OCR definition setting screen 900 for the form type in Fig. 9 will be described. The OCR definition setting screen for the form type is a screen for accepting settings for which areas to be OCR-targeted by accepting operational instructions such as dragging or range specification of the form displayed as shown in 901 among the forms that have been scanned and read. It is also possible to accept settings for OCR targets by setting the names of items to be OCR-targeted in the read item column 902. This concludes the description of Fig. 9.

[0127] Next, the coordinate information table 1000 of the divided areas in Fig. 10 will be described. The coordinate information table 1000 is composed of a form type ID 1001, an integrated area ID 1002, area coordinates 1003, and the number of included items 1004. The form type ID 1001 is information indicating the form that includes each integrated area generated by the processing in Fig. 8. The integrated area ID 1002 is information for uniquely identifying the integrated area generated by the processing in FIG. The area coordinates 1003 are coordinate information indicating the integrated area. The number of included items 1004 is the number of items of OCR definition setting data included in the integrated area specified by the area coordinates. This concludes the explanation of FIG.

[0128] Next, the OCR definition setting data table 1100 in Fig. 11 will be described. The OCR definition setting data table 1100 is composed of a form type ID 1101, an item ID 1102, area coordinates 1103, and a read character type 1104. The form type ID 1101 is information for identifying the form type to which the OCR definition setting data is applied. The item ID 1102 is the ID of the item to be read. The area coordinates 1103 are coordinate information indicating the area to be subjected to OCR. The read character type 1104 is information indicating the type of character to be read from the area specified by the area coordinates. This concludes the description of Fig. 11.

[0129] Next, a description will be given of the feature point information table 1200 of the form-type image in Fig. 12. The feature point information table 1200 is composed of a form type ID 1201 (information for identifying a form in which a feature point exists when related to the feature point ID), a feature point ID 1202 (information for uniquely identifying a feature point), coordinates 1203 (coordinate information indicating the position of a feature point specified by the feature point ID), and feature amount 1204 (feature amount of a feature point related to the feature point ID). This concludes the description of Fig. 12.

[0130] Next, a description will be given of the storage 1300 that manages and stores the various images in Fig. 13. The storage 1300 manages and stores the form type image, the form image, and the converted form image. This concludes the description of Fig. 13.

[0131] As described above, according to this embodiment, it is possible to perform OCR (character recognition) with high accuracy even for semi-standard forms, which are forms in which the positions and sizes of details of the form vary slightly.

[0132] Although the embodiment of the present invention has been described above, the present invention can be embodied, for example, as a system, an apparatus, a method, a program, a recording medium, etc. Specifically, the present invention may be applied to a system composed of multiple devices, or may be applied to an apparatus composed of a single device.

[0133] Moreover, the program of the present invention is a program that enables a computer to execute the processing method of the flowchart shown in Figures 4 and 5, and the storage medium of the present invention stores a program that enables a computer to execute the processing method of Figures 4 and 5. The program of the present invention may be a program for each processing method of each device in Figures 4 and 5.

[0134] As described above, it goes without saying that the object of the present invention can be achieved by supplying a recording medium on which a program that realizes the functions of the above-mentioned embodiments is recorded to a system or device, and having the computer (or CPU or MPU) of that system or device read and execute the program stored on the recording medium.

[0135] In this case, the program itself read out from the recording medium will realize the novel functions of the present invention, and the recording medium on which the program is recorded will constitute the present invention.

[0136] Examples of recording media for supplying the program include flexible disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, DVD-ROMs, magnetic tapes, non-volatile memory cards, ROMs, EEPROMs, silicon disks, etc.

[0137] Furthermore, it goes without saying that the functions of the above-mentioned embodiments are not only realized by the computer executing a program it has read, but also includes cases where an OS (operating system) running on the computer performs some or all of the actual processing based on the instructions of the program, and the functions of the above-mentioned embodiments are realized through that processing.

[0138] Furthermore, it goes without saying that this also includes cases where a program read from a recording medium is written into a memory provided on a function expansion board inserted into a computer or a function expansion unit connected to a computer, and then a CPU or the like provided on the function expansion board or function expansion unit performs some or all of the actual processing based on the instructions of the program code, thereby realizing the functions of the above-mentioned embodiments.

[0139] The present invention may be applied to a system consisting of multiple devices, or to an apparatus consisting of a single device. Needless to say, the present invention can also be applied to a case where the object is achieved by supplying a program to a system or apparatus. In this case, the effect of the present invention can be enjoyed by the system or apparatus by reading a recording medium storing a program for achieving the present invention into the system or apparatus.

[0140] Furthermore, by downloading and reading a program for achieving the present invention from a server, database, etc. on a network using a communication program, the system or device can enjoy the effects of the present invention. Note that the present invention also includes configurations that combine the above-mentioned embodiments and their modified examples. [Explanation of symbols]

[0141] Document Scanner 101 Client PC 102 Server PC103 LAN104

Claims

1. An acquisition means for acquiring an image that is a target for character recognition; a dividing means for dividing the image in a first direction at locations where a predetermined pixel does not exist; a correction means for performing correction for each of the regions divided by the dividing means; character recognition means for performing character recognition processing on the image corrected by the correction means; An information processing device comprising:

2. 2. The information processing apparatus according to claim 1, wherein the dividing means further divides the image divided in the first direction at a location where a predetermined pixel does not exist in the second direction at a location where a predetermined pixel does not exist.

3. 3. The information processing apparatus according to claim 1, wherein, when there are a plurality of locations where the predetermined pixel does not exist, the dividing means divides the image at the plurality of locations.

4. 4. The information processing device according to claim 1, wherein, when there are consecutive locations where the predetermined pixel does not exist, the dividing means divides the image at any one of the consecutive locations.

5. 5. The information processing apparatus according to claim 1, wherein the predetermined pixels are black pixels.

6. 6. The information processing apparatus according to claim 1, further comprising a combining unit that combines an area in which no recognition area to be subjected to character recognition processing exists and an area in which a recognition area exists within the area divided by the dividing unit.

7. 7. The information processing apparatus according to claim 6, wherein the combining means combines an area in which the recognition area does not exist with an area adjacent to the area in which the recognition area exists.

8. An acquisition step in which an acquisition means of the information processing device acquires an image to be subjected to character recognition; a division step in which a division means of the information processing device divides the image at a location where a predetermined pixel does not exist in a first direction of the image; a correction step in which a correction means of the information processing device performs correction for each area divided in the division step; a character recognition step in which a character recognition means of the information processing device performs character recognition processing on the image corrected by the correction means; An information processing method comprising:

9. A program for causing a computer to function as each of the means according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for collating document

    JP1998027208A

  • Form processing device, program for implementing it, and program for creating form format

    JP2004139484A

  • Information processing apparatus, method for controlling the same, and program

    JP2018185702A

  • Image analysis device and image analysis program

    JP2019036146A

  • Information processing apparatus, information processing method, and program

    JP2020107272A