Image processing apparatus, image processing method, and program

The image processing apparatus addresses errors in index extraction from scanned images by allowing user corrections and performing updates or new registrations, ensuring accurate document registration and index extraction even with changing document layouts.

JP7699952B2Active Publication Date: 2025-06-30CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021067973
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-13
Publication Date
2025-06-30
Estimated Expiration
2041-04-13

AI Technical Summary

Technical Problem

Existing methods for extracting indices from scanned images are prone to errors when document layouts change, leading to incorrect overwriting of registered document information, and may incorrectly identify similar documents, resulting in failed index extraction.

Method used

An image processing apparatus that acquires scanned images, determines similar document formats, specifies regions for property setting, allows user correction, and performs updates or new registrations based on user instructions, ensuring accurate index extraction and document registration.

Benefits of technology

The solution effectively updates information for setting properties of scanned images, reducing errors in index extraction and ensuring accurate document registration, even when document layouts change.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699952000001
    Figure 0007699952000001
  • Figure 0007699952000002
    Figure 0007699952000002
  • Figure 0007699952000003
    Figure 0007699952000003
Patent Text Reader

Abstract

To properly update information to be used in processing for setting a property of a scanned image.SOLUTION: An image processing apparatus includes: obtaining means configured to obtain a scanned image; first determination means configured to determine a document type of a document format similar to a document format of the scanned image based on information on each registered document type; extraction means configured to extract a character string included in the scanned image and corresponding to a predetermined item; second determination means configured to determine whether the document format indicated by the scanned image is similar to the document format of the document type determined by the first determination means when a user modifies the extracted character string, on the basis of a method higher in accuracy than that of the first determination means; and display control means configured to display a screen prompting the user to perform overwrite registration when the second determination means determines that the document format is similar.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for extracting an index included in an image.

Background Art

[0002] There is a method of pre-registering document information and determining whether a document corresponding to a scanned image is included in a group of registered documents. Further, when a document corresponding to a scanned image is specified, based on the position information of the character string associated with the specified document, there is a method of extracting a desired character string for setting a property from the scanned image and presenting it to the user.

[0003] Patent Document 1 describes a method of selecting a template by comparing all registered templates with a document reading result, and extracting a character string representing an attribute such as a claim amount from the reading result based on the selected template.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] When the user corrects the character string because the extracted character string is not the desired character string, a method of overwriting and registering the information of the registered document based on the scanned image can be considered. By overwriting and registering the information of the registered document, even when the layout of a document issued by a certain company is changed, next time, a desired character string can be extracted from a scanned image of a document of the same type as that document.

[0006] In addition, there may be a case where an unregistered document issued by a different company, whose layout is similar to that of a certain registered document, is scanned. In this case, even though the document corresponding to the scanned image should be determined as unregistered, the registered document may be identified as the document corresponding to the scanned image. If an incorrect document is identified, a string different from the desired string is extracted, and thus the extracted string is corrected by the user. For this reason, there is a risk that the information of the registered document may be erroneously overwritten and registered based on the scanned image. If the information of the registered document is erroneously overwritten and registered, when the document is scanned next and a process of extracting the desired string based on the scanned image is performed, there is a risk that the extraction of the desired string may fail.

[0007] The technology of the present disclosure aims to appropriately update information used in a process for setting properties of a scanned image.

Means for Solving the Problem

[0008] The image processing apparatus of the present disclosure includes an acquisition unit that acquires a scanned image obtained by scanning a document, a determination unit that determines a document format similar to the format of the document indicated by the scanned image from among pre-registered document formats, a reception unit that specifies information on a region in the scanned image for setting properties of the scanned image based on information registered in association with the document format determined by the determination unit and accepts correction of the information on the region, a display control unit that performs a display for accepting an instruction from a user to perform a process of updating the information registered in association with the determined document format based on the correction, or an instruction to perform a process of newly registering the format of the document indicated by the scanned image in association with the information based on the correction, and a processing unit that performs the updating process or the newly registering process based on the instruction received from the user. For a correction of a region in a scanned image obtained by scanning a first document similar to a predetermined document format as the document, a display prompting the user to give an instruction to perform the newly registering process is performed, and for a correction of a region in a scanned image obtained by scanning a second document more similar to the predetermined document format than the first document as the document, a display prompting the user to give an instruction to perform the updating process is performed.

Advantages of the Invention

[0009] According to the technology of the present disclosure, it is possible to appropriately update information used in a process for setting properties of a scanned image.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Modes for Carrying Out the Invention

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the technology of the present disclosure according to the claims, and not all combinations of the features described in this embodiment are essential for the solution means of the technology of the present disclosure.

[0012] <Embodiment 1> In this embodiment, a process of extracting a character string (hereinafter also referred to as an index) of a predetermined item included in a scanned image obtained by scanning a paper document (input document) such as a form with a reading device will be described. To extract an index from the content of the scanned image, the document is registered in advance, and an extraction area for extracting the index is set for each document. Then, after determining which registered document the input document corresponds to (or is similar to), the index is extracted by performing OCR processing partially on the corresponding extraction area in the scanned image.

[0013] Also, when the input document is an unregistered document, there is a function to register the input document as a new document. Also, there is a function to overwrite and register information regarding the corresponding registered document based on the input document. With these functions, the registered documents can be updated appropriately. Therefore, for example, it is possible to handle cases where the company that is the business partner changes, or when the business partner company changes the document format.

[0014] However, even if the input document is an unregistered document that is different from any of the registered document groups, it may be determined that the input document is similar to any of the registered document groups. In such a case, there is a risk that the information regarding the registered document will be overwritten based on the input document. If the information regarding the registered document is overwritten based on a document different from the registered document, then, next, when scanning a document of the same type as the registered document and extracting the index, the index extraction may fail.

[0015] Therefore, in this embodiment, a method of appropriately recommending new registration or overwrite registration to the user according to the input document will be described.

[0016] [System Configuration] FIG. 1 is a diagram showing the overall configuration of a system to which the present embodiment is applicable. The system 105 of the present embodiment includes an image forming apparatus 100 and a terminal 101. As shown in FIG. 1, the image forming apparatus 100 is connected to a LAN 102 and can communicate with terminals 101 such as a PC via the Internet 103 or the like. Note that in the present embodiment, the terminal 101 may be omitted, and only the configuration of the image forming apparatus 100 may be used.

[0017] The image forming apparatus 100 is a multi-functional peripheral (MFP) having a display / operation unit 123 (see FIG. 2), a scanner unit 122 (see FIG. 2), a printer unit 121 (see FIG. 2), and the like. The image forming apparatus 100 can be used as a scanning terminal that scans a document manuscript using the scanner unit 122. Further, it has a display / operation unit 123 such as a touch panel or hard buttons, and displays a user interface for displaying a file name or a recommendation result of a storage destination, and receiving an instruction from a user.

[0018] [Hardware Configuration of Image Forming Apparatus (100)] FIG. 2 is a block diagram showing the hardware configuration of the image forming apparatus 100. The image forming apparatus 100 of the present embodiment includes a display / operation unit 123, a scanner unit 122, a printer unit 121, and a control unit 110.

[0019] The control unit 110 includes a CPU 111, a storage device 112 (ROM 118, RAM 119, HDD 120), a printer I / F unit 113, a network I / F unit 114, a scanner I / F unit 115, and a display / operation I / F unit 116. In addition, in the control unit 110, these units are communicably connected to each other via a system bus 117. The control unit 110 controls the operation of the entire image forming apparatus 100.

[0020] The CPU 111 functions as means for executing each process such as reading control, image processing, and display control in the flowchart described later by reading and executing a control program stored in the storage device 112.

[0021] The memory device 112 stores and holds control programs, image data, metadata, setting data, processing result data, and the like. The memory device 112 includes a ROM 118 which is a non-volatile memory, a RAM 119 which is a volatile memory, and an HDD 120 which is a large-capacity storage area. The ROM 118 is a non-volatile memory that holds control programs and the like, and the CPU 111 reads out and controls the control programs. The RAM 119 is a volatile memory used as a main memory of the CPU 111, a temporary storage area such as a work area.

[0022] The network I / F unit 114 connects the control unit 110 (image forming apparatus 100) to the LAN 102 via the system bus 117. The network I / F unit 114 transmits image data to external devices on the LAN 102 or receives various information from external devices on the LAN 102.

[0023] The scanner I / F unit 115 connects the scanner unit 122 and the control unit 110 via the system bus 117. The scanner unit 122 reads a document manuscript and generates scan image data, and inputs the scan image data to the control unit 110 via the scanner I / F unit 115. Note that the scanner unit 122 is provided with a document feeder and can feed a plurality of manuscripts placed in the tray one by one and read them continuously.

[0024] The display / operation I / F unit 116 connects the display / operation unit 123 and the control unit 110 via the system bus 117. The display / operation unit 123 is provided with a liquid crystal display unit having a touch panel function, hard buttons, and the like.

[0025] The printer I / F unit 113 connects the printer unit 121 and the control unit 110 via the system bus 117. The printer unit 121 receives the image data generated by the CPU 111 via the printer I / F unit 113, and performs printing processing on the recording paper using the received image data. As described above, in the image forming apparatus 100 according to the present embodiment, the above hardware configuration enables the provision of an image processing function.

[0026] [Functional Configuration of Image Forming Apparatus] FIG. 3 is a block diagram showing the functional configuration of the image forming apparatus 100. In FIG. 3, among the various functions of the image forming apparatus 100, functions related to the processing until a document manuscript is scanned, digitized (filed), and saved are shown.

[0027] The display control unit 301 displays a user interface screen (UI screen) for receiving various user operations on the touch panel of the display / operation unit 123. Various user operations include, for example, scan settings, scan start instructions, index correction instructions, registration method instructions, file name settings, file save instructions, and the like.

[0028] The scan control unit 302 instructs the scan execution unit 303 to execute a scan process together with scan setting information in response to a user operation made on the UI screen (for example, pressing the "Scan Start" button). The scan execution unit 303 causes the scanner unit 122 to execute a reading operation of a document manuscript via the scanner I / F unit 115 in accordance with the scan process execution instruction from the scan control unit 302, and generates scan image data. The generated scan image data is stored in the HDD 120 by the scan image management unit 304.

[0029] The image processing unit 305 performs image analysis processes such as text block detection processing, OCR processing (character recognition processing), and similar document determination processing on the scan image data, as well as image processing such as rotation and inclination correction. The image forming apparatus 100 also functions as an image processing apparatus by the image processing unit 305. The character string area detected from the scan image is also called a "text block". Details of the image processing will be described later.

[0030] The functions of each part in FIG. 3 are realized by the CPU of the image forming apparatus 100 expanding and executing the program code stored in the ROM in the RAM. Alternatively, some or all of the functions of each part in FIG. 3 may be realized by hardware such as an ASIC or an electronic circuit.

[0031] [Flowchart of File Generation Process for Scan Image] The image forming apparatus 100 reads a document manuscript, performs image processing on the scan image of the first page of the document manuscript, generates a file name using the character string included in the scan image, and recommends it to the user through the display / operation unit 123. The entire process will be described below.

[0032] A series of processes shown in the flowchart of FIG. 4 are performed by the CPU of the image forming apparatus 100 expanding and executing the program code stored in the ROM into the RAM. Also, some or all of the functions of the steps in FIG. 4 may be realized by hardware such as an ASIC or an electronic circuit. Note that the symbol "S" in the description of each process means that it is a step in the flowchart, and the same applies to the subsequent flowcharts.

[0033] In S400, when the scan control unit 302 receives a scan instruction from the user via the display / operation unit 123, it causes the scan execution unit 303 to read (scan) a plurality of document manuscripts one by one from the tray of the document feeder of the scanner unit 122. Then, the scan control unit 302 acquires the image data of the image (referred to as a scan image or an input image) obtained as a result of the scan.

[0034] In S401, the image processing unit 305 analyzes the image data acquired in S400 and performs a process (index extraction process) of extracting an index included in the scan image based on an index extraction rule. An "index" is a character string representing a predetermined item such as a document title, a management number, or a company name. In the present embodiment, the index is used to set a part of the file name when saving the scan image or properties such as metadata. Details of the index extraction process in this step will be described later with reference to FIG. 5.

[0035] Note that the method of using the index is not limited to generating file names or extracting metadata. It may be used to set other properties such as folder paths. That is, file names and metadata are a type of information set as properties (attributes) related to scanned image data.

[0036] In S402, the display control unit 301 displays on the display / operation unit 123 a confirmation / modification screen 800 (see FIG. 8) that includes the index extracted in S401 and the file name and metadata generated using that index. Through the confirmation / modification screen 800, properties such as the index, file name, and metadata to be set for the scanned image are presented (recommended) to the user. Also, the display control unit 301 accepts input from the user to modify the extracted index. When the display control unit 301 accepts input for modification from the user via the display / operation unit 123, it presents the file name and metadata based on the modified index.

[0037] When the user's confirmation of the file name and metadata presented by the display control unit 301 is received, the presented file name is set as the file name of the scanned image. The confirmation / modification process will be described later.

[0038] In S403, the image processing unit 305 determines via the display / operation unit 123 whether an index has been input. For example, when the extracted index is modified by the user to another string, it is determined that an index has been input. Or, when a new unregistered document is scanned and no index is extracted, the user will input the string of each item as a string indicating the index. In this case as well, it is determined that an index has been input.

[0039] When the user inputs an index (S403 is YES), document registration processing is executed in S404. In the document registration processing, a process of overwriting and registering information regarding the document type included in the index extraction rules, or a new registration process of registering an unregistered document type by registering a new document is performed. The document registration processing will be described later.

[0040] When the user does not input an index (S403 is NO), or when the processing of S404 is completed, the process proceeds to S405. In S405, the image processing unit 305 creates a file from the image data acquired in S400 and sets the file name which is the property determined in S402. In this embodiment, as an example, it will be described that the scan image is saved in the PDF (Portable Document Format) as the file format. In the case of PDF, it is possible to save the scan image by dividing it into pages. When the scan images of a plurality of document manuscripts are acquired in S400, the image data corresponding to each document manuscript is saved as separate pages in one file.

[0041] In S406, the scan image management unit 304 transmits the file created in S405 to a predetermined destination through the LAN 102.

[0042] [Regarding the Index Extraction Process (S401)] FIG. 5 is a flowchart showing the index extraction process of S401. The details of the index extraction process will be described with reference to FIG. 5. In the index extraction process, for one page of the scan image, a process of correcting the orientation, specifying the document type, and performing index extraction according to the document type is performed.

[0043] In S500, the image processing unit 305 detects the angle of inclination of the scanned image from the image data, and corrects the inclination of the scanned image by rotating the image in the opposite direction by the detected inclination. The inclination to be corrected is caused, for example, by the wear of the rollers in the document feeder of the scanner unit 122 when scanning a document manuscript, resulting in the document manuscript not being read straight. Or it occurs because the scanned document manuscript is not printed straight during printing.

[0044] As a method for detecting the angle of inclination, first, objects included in the image data are detected, and object groups adjacent in the horizontal direction or the vertical direction are connected. Then, the angle formed by connecting the center positions of the connected object groups is derived to obtain the inclination, that is, to determine how much it is inclined from the horizontal direction or the vertical direction. Note that the method for detecting the inclination is not limited to this method. For example, other methods may include obtaining the center coordinates of the objects included in the image data, rotating the center coordinate group in units of 0.1 degrees, and obtaining the angle at which the center coordinate group is arranged most densely in the horizontal direction or the vertical direction as the inclination of the scanned image. By correcting the inclination of the scanned image, the processing accuracy of each of the subsequent rotation correction, block selection processing, and OCR processing can be improved.

[0045] In S501, the image processing unit 305 performs rotation correction on the scanned image after inclination correction obtained as a result of the processing in S500 in units of 90 degrees so that the characters in the image are in an upright orientation. The method of rotation correction is, for example, using the scanned image after inclination correction as a reference image, preparing four images: the reference image, the image obtained by rotating the reference image by 90 degrees, the image obtained by rotating the reference image by 180 degrees, and the image obtained by rotating the reference image by 270 degrees. Then, for each image, a simple OCR processing that can be processed at high speed is executed, and there is a method of using the image with the largest number of characters recognized with a confidence level of a certain value or more as the image after rotation correction. However, the method of rotation correction is not limited to this method. Hereinafter, the scanned image refers to the scanned image corrected in S500 and S501 unless otherwise specified.

[0046] In S502, the image processing unit 305 performs block selection processing on the scanned image. The block selection processing is a process of classifying the image into a foreground region and a background region, then dividing the foreground region into text blocks and other blocks, and detecting the text blocks.

[0047] Specifically, for the scanned image binarized into black and white, contour tracing is performed to extract a block of pixels surrounded by a black pixel contour. Then, for a block of black pixels with an area larger than a predetermined size, contour tracing is also performed on the white pixels inside it to extract a block of white pixels, and further, a block of black pixels is recursively extracted from inside a block of white pixels with an area of a certain size or more. The block of black pixels thus obtained is determined as the foreground region. The determined foreground region is classified into regions with different attributes according to size and shape. For example, a foreground region with an aspect ratio close to 1 and within a certain size range is regarded as a pixel block corresponding to a character, and a region where adjacent characters can be grouped neatly is determined as a character string region (TEXT). A flat pixel block is determined as a line region (LINE). The range occupied by a block of black pixels that encloses a block of white pixels with a rectangular shape and a size of a certain size or more in an aligned manner is determined as a table region (TABLE). A region where irregular pixel blocks are scattered is determined as a photo region (PHOTO). And a pixel block with other shapes is determined as a picture region (PICTURE). Then, from among those region-divided by object attribute, the foreground region (TEXT) determined to have a character attribute is detected as a text block.

[0048] FIG. 6 is a diagram showing an example of the result of the block selection processing. FIG. 6(a) shows the scanned image after rotation correction. FIG. 6(b) shows the result of the block selection processing on the scanned image of FIG. 6(a), and the rectangle indicated by the dotted line represents the foreground region. Note that in FIG. 6(b), the attributes of all foreground regions are determined, but only the attributes of some foreground regions are shown. The information of each text block detected in this step (information indicating the attribute, position, and size of each block) is used in subsequent processes such as OCR processing and similarity calculation.

[0049] In the block selection process of this step, only text blocks are detected. The reason is that the positions of character strings well represent the structure of the scanned image and are closely related to the index information. It does not exclude using the information of blocks determined to have other attributes such as photo areas and table areas in subsequent processing.

[0050] In S503, the image processing unit 305 acquires the index extraction rules from the HDD 120 and expands them in the RAM 119.

[0051] FIG. 7 is a diagram showing a part of the index extraction rules (hereinafter simply referred to as extraction rules). In the extraction rules, for each document corresponding to the registered document type, each data of "document ID", "registered name", "scanned image", "document identification information", and "index information" is associated in record units. The extraction rules hold these combinations (records) for the number of registered document types. FIG. 7 shows a record in which information regarding the document format of the document type to which "0001" is assigned as the document ID among the registered document types included in the extraction rules is held.

[0052] The "document ID" holds a unique value representing the type of the document. The "registered name" holds a name representing the type of the document. The "scanned image" holds the scanned image of the document corresponding to the registered document type. Note that the image held in the "scanned image" only needs to hold information that allows the user to understand the content of the document. For example, an image with a resolution reduced to about 150 dpi may be held.

[0053] "Document identification information" holds the position and size of text blocks obtained as a result of performing block selection processing on the scanned image of a document of the document type registered in the record. Document identification information is information for determining the document type of the input document and is used in document matching described later. Note that the document identification information only needs to be information necessary for specifying the document type and is not limited to information on the position and size of text blocks. For example, the position and shape of ruled lines included in the document may be retained, or combinations of character strings appearing in the document may be utilized.

[0054] "Index information" holds index information for extracting indexes. An index is a character string used to set the properties of the scanned image as described above. Specifically, the index information includes the coordinates and size information of text blocks containing the character strings (indexes) of each item in the document of the registered document type. The image 701 of "index information" in FIG. 7 is illustrated by arranging the positions and sizes of text blocks containing each item at the coordinates on the image for explanation. In addition, the index information includes information indicating the indexes used to generate the file name and their order, and information for attaching as metadata.

[0055] The "file name rule" of the index information shows that the indexes of title, sender company name, and document number are connected in this order with an underscore as a separator to generate a file name. Also, the "metadata" shows that the index of the total price is used as metadata. That is, by extracting the indexes of predetermined items, the file name recommended to the user and the metadata can be set.

[0056] In this embodiment, an example is shown in which the extracted index is used as a file name or metadata. However, rules for determining folder information of the file destination, which is other property information, may be retained. Also in this case, the property information generated using the index is recommended to the user in S402, and the property information corrected or confirmed by the user in S405 is set in the scan image file. Further, the items to be extracted are not limited to the title, the issuing company name, the document number, and the total amount.

[0057] Information regarding the document type registered in the extraction rules is set based on a scan image obtained by scanning a document of that document type. It can be said that information regarding the document is registered in the extraction rules. Therefore, in the following description, the document type registered in the extraction rules may sometimes be simply described as the registered document.

[0058] In S504, the image processing unit 305 performs document matching on the scan image. In document matching, a process of determining a document corresponding to the scan image is performed from the document group registered in the extraction rules.

[0059] In the document matching process, first, the text blocks of the scan image and the text blocks of each registered document are compared one-to-one, and the similarity representing how similar the shapes and arrangements of the text blocks are is calculated. As a method for calculating the similarity, for example, alignment is performed between the entire text block of the scan image and the entire text block of the registered document. Then, the square of the sum of the areas where each text block of the scan image overlaps with each text block of the registered document (referred to as value A) is obtained. Further, the product of the sum of the areas of the text blocks of the scan image and the sum of the areas of the text blocks of the registered document (referred to as value B) is obtained. And there is a method of using the value obtained by dividing value A by value B as the similarity. This calculation of the similarity is performed between the scan image and all the documents registered in the extraction rules.

[0060] Among the documents registered in the extraction rules, the document with a similarity degree of at least the threshold value TH1 and the highest similarity degree is determined as the document (type) with a document format similar to the document format indicated by the scan image. That is, a document similar to the document indicated by the scan image can be determined from the documents registered in the extraction rules.

[0061] In S505, the image processing unit 305 determines whether a document similar to the document indicated by the scan image has been determined in the extraction rules based on the result of the document matching executed in S504.

[0062] If there is no document in the documents registered in the extraction rules whose similarity degree with the document format of the scan image is at least the threshold value TH1, it is determined that a document similar to the document indicated by the scan image could not be determined. If a document similar to the document indicated by the scan image could not be determined (S505 is NO), the processing of this flowchart ends. In this case, in S402 which is the next step of the flowchart in FIG. 4, the file name and metadata are not recommended to the user. For this reason, the display control unit 301 accepts an index input from the user. Then, in S403, it is determined that the index input has been performed and the process proceeds to S404.

[0063] If a document similar to the document indicated by the scan image has been determined (S505 is YES), the process proceeds to S506. That is, if a document with a document format whose similarity degree with the document format indicated by the scan image is at least the threshold value TH1 has been determined from the documents registered in the extraction rules, the process proceeds to S506. In S506, the image processing unit 305 assigns a value indicating the document ID associated with the document determined in S504 to the scan image.

[0064] In S507, the image processing unit 305 executes an index block determination process for determining a text block of an item to be extracted within the scanned image based on the information of the extraction rule associated with the document ID assigned in S506. A text block containing a character string (index) indicating items such as a title, an issuing company name, and a form number may be referred to as an index block.

[0065] To determine the index block, first, overall alignment is performed between the text block group of the scanned image obtained in S502 and the text block group of the registered document determined to be similar to the document indicated by the scanned image in S504. Then, the text block within the scanned image with the highest degree of overlap with the index block of the registered document is determined as the index block in the scanned image. Note that the method for determining the index block is not limited to this method. For example, a partial layout composed of an index block and its surrounding text blocks may be extracted from the text block group of the document determined to be similar. Then, local alignment may be performed on the text block group of the scanned image after overall alignment obtained in S502 using the extracted partial layout to determine the index block within the scanned image. The local alignment may be performed by executing pattern matching using the partial layout within a preset search range.

[0066] In S508, the image processing unit 305 executes partial OCR processing on each item's index block group within the scanned image determined in S507, and extracts the character string corresponding to each index block as the index of each item.

[0067] [Regarding the confirmation / correction process (S402)] FIG. 8 is a diagram showing an example of a confirmation / correction screen 800 displayed on the display / operation unit 123. The details of the confirmation / correction process (S402) will be described with reference to FIG. 8.

[0068] The preview area 820 is an area where the scan image acquired in S400 is displayed as a preview image. A rectangle representing the position and size of the index block is highlighted and displayed superimposed on the scan image. When a document similar to the document indicated by the scan image is determined in S504, the index block determined in S507 is pre-displayed in the preview area 820.

[0069] Button 801 is a button for instructing an enlargement of the display magnification of the preview image, and button 802 is a button for instructing a reduction of the display magnification. Button 803 is a button for instructing an enlargement or reduction of the size of the preview image so that the preview image fits the width or height of the preview area 820.

[0070] Text fields 804 and 805 are areas where a file name and metadata generated by combining indexes are displayed.

[0071] The index area 806 is composed of index areas 806a to 806d for each item. Each of the index areas 806a to 806d includes index names 807 to 810, partial preview areas 811 to 814, and text fields 815 to 818. In the partial preview areas 811 to 814, respective images corresponding to the index blocks are partially cut out from the scan image and displayed. In the text fields 815 to 818, strings obtained as a result of performing OCR processing on the index blocks are respectively displayed. When a document similar to the document indicated by the scan image is determined in S504, the extracted indexes are pre-displayed in the text fields 815 to 818.

[0072] For items for which the index block could not be determined in S507, the respective item names are displayed in index names 807 to 810, but the partial preview areas 811 to 814 and the text fields 815 to 818 are displayed in an empty state. If no document similar to the document indicated by the scanned image could be determined in S504, the index blocks for all items of title, issuing company name, form number, and total amount are treated as those for which the index blocks could not be determined. For this reason, all the partial preview areas 811 to 814 and the text fields 815 to 818 are displayed in an empty state.

[0073] The procedure for correcting the index block (correcting the index) when the position of the index block is incorrectly determined or when the index block could not be determined will be described. The case of correcting the index block for the item of form number (number) will be used as an example for the explanation. First, the user selects the index area 806c of the item to be corrected, which is “(3) number”. For example, it is selected by clicking on any position within the index area 806c. Subsequently, the text block containing the character string of the item to be corrected on the preview image in the preview area 820 is selected by clicking or the like. In accordance with the selected text block, a partial image of the selected text block is displayed in the partial preview area 813, and the text field 817 displays the character string obtained by performing OCR processing on the selected text block. The position of the text block thus selected is corrected as the position of the index block for that item. The position information of the corrected index block is used to update the index information of the document to be overwritten in the overwriting registration described later. Or the position information of the corrected index block is used to register the index information of a new document in the new registration.

[0074] When the user has completed the modification of the index block and finished checking the currently displayed index, the user presses the decision button 830. When the pressing of the decision button 830 is accepted, the index, file name, and metadata are finalized, and the process proceeds to S403.

[0075] [Regarding the document registration process (S404)] Figure 9 is a flowchart showing the document registration process of S404. The details of the document registration process will be described with reference to Figure 9. In the document registration process, a process of updating information regarding the documents registered in the extraction rules (overwriting registration) or a process of registering new information in the extraction rules (new registration) is performed.

[0076] In S900, the image processing unit 305 determines whether a document ID has been assigned to the scanned image in S506. When a document similar to the document shown in the scanned image can be determined from the extraction rules in the document matching of S504 in Figure 5, a document ID is assigned to the scanned image. Therefore, in this step, it is determined whether a document similar to the document shown in the scanned image has been determined in the document matching of S504 in the index extraction process.

[0077] If a document ID has been assigned to the scanned image (S900 is YES), the process proceeds to S901. In S901, the image processing unit 305 acquires the document ID assigned in S506 of Figure 5.

[0078] In S902, the image processing unit 305 determines whether the document associated with the document ID assigned in S506 among the documents registered in the extraction rules is similar to the document shown in the scanned image. In S902, it is determined whether they are similar by a method with higher accuracy than the method in the document matching of S504. The document associated with the document ID assigned in S506 is the registered document similar to the document shown in the scanned image determined in S504 of the index extraction process.

[0079] In S902 of the present embodiment, first, the image processing unit 305 acquires the similarity between the document indicated by the scanned image and the document determined to be similar to the document indicated by the scanned image in S504. The similarity may be calculated in the same manner as the document matching in S504, or the value calculated in the document matching in S504 may be acquired.

[0080] Then, in S902 of the present embodiment, the image processing unit 305 uses a threshold TH2 different from the threshold TH1 used in the document matching in S505. When the similarity between the document indicated by the scanned image and the document determined in S504 is greater than the threshold TH2, the image processing unit 305 determines that the document indicated by the scanned image is similar to the document corresponding to the scanned image determined in S504.

[0081] FIG. 10 is a diagram showing an example of a document registered in the extraction rule and an input document to be scanned. The document in FIG. 10(a) represents a document of the document type registered in the extraction rule with the document ID "0001" shown in FIG. 7.

[0082] The input document 1 in FIG. 10(b) is an example of an input document that has the same issuing company as the document in FIG. 10(a) and is of the same type, but with some layout changes.

[0083] The input document 2 in FIG. 10(c) is an example of a document that has a layout similar to the document in FIG. 10(a) but is issued by a different company and is of a different document type from the document in FIG. 10(a).

[0084] The input document 3 in FIG. 10(d) is an example of a document with a different document type because it is issued by a company different from the issuing company of the document in FIG. 10(a) and has a different layout.

[0085] FIG. 11 is a diagram for explaining the threshold TH2 used in S902 of the present embodiment. FIG. 11(a) shows the relationship between the threshold TH1 used in the document matching of S504 and the threshold TH2 used in S902. In the document matching of S504, if the document corresponding to the scanned image is not determined, the process cannot proceed to S506 to S508 to execute the process for extracting the index. Therefore, in order to increase the frequency of transitioning to S506 to S508, a certain degree of difference between the document indicated by the scanned image and the registered document is allowed, and the threshold TH1 is set so that the document type of the input document can be determined.

[0086] On the other hand, the threshold TH2 used in the determination of S902 is a threshold for determining whether the document determined to correspond to the scanned image in S504 is really the same type of document as the input document. Therefore, the threshold TH2 is set in advance to a value larger than the threshold TH1.

[0087] FIG. 11(b) is a diagram showing the similarity of the document formats of input document 1, input document 2, and input document 3 to the document format of the registered document in FIG. 10(a). The process of S902 will be described with reference to FIGS. 10 and 11.

[0088] The similarity of input document 1 is greater than the threshold TH1 and the threshold TH2. Therefore, in the document matching of S504 in FIG. 5, it is determined that the document indicated by the scanned image of input document 1 is similar to the document in FIG. 10(a). Therefore, it is determined that a document ID is assigned in S900 and the process proceeds to S902. In S902, since it is determined that the similarity is greater than the threshold TH2, it is also determined in S902 that the document indicated by the scanned image of input document 1 is similar to the document in FIG. 10(a). Therefore, it can be determined that input document 1 is the same type of document as the document used in the index extraction process of FIG. 10(a).

[0089] The similarity of Input Document 2 is greater than threshold TH1 but less than threshold TH2. Therefore, in the document matching of S504 in FIG. 5, it is determined that the document shown in the scanned image of Input Document 2 is similar to the document in FIG. 10(a). For this reason, it is determined that a document ID is assigned in S900 and the process proceeds to S902. However, in S902, since it is determined that the similarity is less than threshold TH2, it is determined that the document shown in the scanned image of Input Document 2 is not similar to the document in FIG. 10(a). Input Document 2 is an example of a document that is determined to be similar in the document matching of S504 because it has a similar layout although it is a different type of document from the document in FIG. 10(a).

[0090] The similarity of Input Document 3 is less than threshold TH2 and threshold TH1. Therefore, in the document matching of S504 in FIG. 5, it is not determined that it is similar to the document in FIG. 10(a). For this reason, in S900, it is not determined that a document ID is assigned, and S902 is not transitioned to.

[0091] The next S903 is a step for switching the process according to the processing result of S902. If it is determined that the document shown in the scanned image is similar to the document determined in S504 (S903 is YES), the process proceeds to S904.

[0092] In S904, the display control unit 301 performs a process of recommending to the user an overwriting registration of information regarding the document registered in the extraction rule.

[0093] For example, in the input document 1 of FIG. 10(b), a text block 1000 regarding the expiration date that does not exist in the registered document of FIG. 10(a) is inserted, and the position of the index block 1001 of the form number is shifted downward. Thus, even for a scanned image of a document of the same type as the registered document, if the positions of the registered indexes are partially different, the extraction of the indexes may fail. In this case, in the confirmation / correction process of S402, the index block is corrected according to the user's instruction. Also, the user can update the registered document of FIG. 10(a) based on the scanned image of the input document 1 by instructing overwriting registration. Therefore, by performing overwriting registration, when a document similar to the input document 1 is scanned next time, it is possible to suppress the failure of index extraction.

[0094] On the other hand, when a document similar to the document indicated by the scanned image in S504 of FIG. 5 cannot be determined from the extraction rules, it is determined in S900 that the document ID is not assigned. When the document ID is not assigned (S900 is NO), the process proceeds to S905. Also, when it is determined in the process of S902 that the document indicated by the scanned image is not similar to the document determined in S504 (S903 is NO), the process also proceeds to S905.

[0095] In S905, the display control unit 301 performs a process of recommending new registration to the user based on the scanned image. For example, like the input document 2 of FIG. 10(c), although it is a document of an unregistered type, since the layout of the document is similar to that of the registered type of document, in the document matching of S504, it may be determined that there is a similar document. Since the input document 2 is of an unregistered document type, it is desirable to be newly registered. Therefore, even when it is determined to be similar in the document matching of S504, by performing document matching again in S902, it is possible to recommend new registration when the input document is of an unregistered type.

[0096] Thus, in this embodiment, it is determined whether the document used to extract the index and the document indicated by the scanned image are similar in a different method with higher accuracy than during the index extraction process. Therefore, it is possible to appropriately switch between recommending overwriting registration and recommending new registration.

[0097] FIGS. 12(a) to 12(c) are diagrams showing an example of a registration confirmation screen 1200, which is a screen for a user to indicate whether to perform overwriting registration or new registration. The registration confirmation screen 1200 will be described with reference to FIG. 12(a).

[0098] Radio buttons 1201 to 1203 are provided corresponding to "Overwriting Registration", "New Registration", and "Do Not Register", and are set to be in a state where any one of the radio buttons 1201 to 1203 is selected. The text field 1204 is an area for displaying the registration name of the document registered in the extraction rule to be overwritten. The thumbnail area 1205 is an area for displaying the scanned image of the document to be overwritten as a thumbnail.

[0099] The detailed confirmation / change button 1206 is a button for transitioning to a document registration screen 1300 (see FIG. 13), which is a screen for detailed confirmation of the document to be overwritten or for changing the document to be overwritten to another document. The text field 1207 is an area for receiving the document name from the user when performing new registration.

[0100] The decision button 1208 is a button for the user to instruct processing according to the selected radio buttons 1201 to 1202. By pressing the decision button 1208 in a state where the radio button 1201 for "Overwriting Registration" is selected, the user can instruct overwriting registration. Similarly, by selecting the radio buttons 1202 or 1203 for "New Registration" or "Do Not Register" and pressing the decision button 1208, the user can instruct processing for new registration or not registering.

[0101] FIG. 12(a) is an example of a registration confirmation screen 1200 for recommending overwriting registration to the user. That is, it is an example of the registration confirmation screen 1200 displayed by the display control unit 301 in S904. For example, when the input document 1 in FIG. 10(b) is scanned and the index extracted from the scanned image is corrected by the user, in the document registration process, it transitions to S904 and the registration confirmation screen in FIG. 12(a) is displayed. Since overwriting registration is recommended in S904, the registration confirmation screen 1200 is displayed with the radio button 1201 for the user to instruct overwriting registration selected.

[0102] When overwriting registration is recommended in S904, as the document to be overwritten, the document determined to be similar in the document matching of S504 in FIG. 5 is recommended. For example, FIG. 12(a) is an example where the registered document in FIG. 7 is recommended as the document to be overwritten. Therefore, in the text field 1204 of the registration name in the registration confirmation screen 1200 in FIG. 12(a), "Estimate_ABC" is displayed, and in the thumbnail, the scanned image in FIG. 7 is displayed. Therefore, when the user instructs overwriting registration, it becomes possible to directly select, as the document to be overwritten, the document with the highest similarity to the document indicated by the scanned image in the document matching of S504 in FIG. 5.

[0103] FIG. 12(b) is an example of a registration confirmation screen 1200 for recommending new registration. That is, it is an example of the registration confirmation screen 1200 displayed by the display control unit 301 in S905. For example, when the input document 2 in FIG. 10(c) is scanned and the index extracted from the scanned image is corrected by the user, in the document registration process, it transitions to S905 and the registration confirmation screen in FIG. 12(b) is displayed. Since new registration is recommended in S905, the registration confirmation screen 1200 is displayed with the radio button 1202 for instructing new registration selected.

[0104] In the text field 1204 and the thumbnail area 1205 of FIG. 12(b) as well, the registered name and the thumbnail of the document determined to be similar in the document matching of S504 are also displayed. That is, even when the user instructs overwriting registration instead of new registration, it is possible to select, as the document to be overwritten, the document having the highest similarity to the document indicated by the scanned image in the document matching of S504 in FIG. 5.

[0105] FIG. 12(c) is an example of a registration confirmation screen 1200 for recommending new registration, which is displayed by the display control unit 301 when proceeding to S905 after determining that no document ID is assigned in S900. For example, it is a registration confirmation screen when the input document 3 in FIG. 10(d) is scanned. Similar to FIG. 12(b), the registration confirmation screen 1200 is displayed with the radio button 1202 for instructing new registration selected. However, different from FIG. 12(b), since no similar document has been determined in the document matching of S504 in FIG. 5, the information of the document to be overwritten is not displayed in the text field 1204 and the thumbnail area 1205. When instructing overwriting registration, the user presses the detailed confirmation / change button 1206 and selects the document to be overwritten on the document registration screen 1300 (see FIG. 13 or FIG. 14).

[0106] As described above, in this embodiment, it is switched whether to recommend the document determined to be similar in the index extraction process as the document to be overwritten or not to recommend the document to be overwritten. Therefore, when instructing overwriting registration, the labor of the user for selecting a document can be reduced.

[0107] FIGS. 13 and 14 are diagrams showing an example of the document registration screen 1300. The document registration screen 1300 is a screen that is displayed when the detailed confirmation / change button 1206 is pressed on the registration confirmation screen 1200. The document registration screen 1300 will be described with reference to FIG. 13.

[0108] The preview area 1301 is composed of a preview area 1301a where an image of a registered document to be overwritten is displayed and a preview area 1301b where a scanned image of the input document is displayed. Therefore, the user can visually compare the input document with the document to be overwritten.

[0109] In the list 1302, a list of the registered names of the documents registered in the extraction rules is displayed. The user can change the document to be overwritten by selecting a document from the list 1302. Also, the image of the document selected from the list 1302 is displayed in the preview area 1301a.

[0110] The sort instruction button 1303 is a button for instructing the sorting of the registered names of the documents displayed in the list 1302, and can, for example, give an instruction for ascending or descending order with respect to similarity, registration date and time, or usage date and time.

[0111] The filter instruction button 1304 is a button for instructing the narrowing down of the documents displayed in the list 1302. For example, only the documents whose similarity of the documents indicated by the scanned image is equal to or higher than a certain value can be displayed in the list 1302, or filtering can be performed by the registered name and the documents can be displayed in the list 1302. Thereby, the user can, for example, give an instruction to display the documents in descending order of similarity by the sort instruction and further give an instruction to display only the documents whose similarity with the documents indicated by the scanned image is equal to or higher than a certain value by the filter instruction. Therefore, the user can select the document to be compared and displayed in the preview area 1301a from among the documents whose similarity is equal to or higher than a certain value, and thus can reduce the trouble of selecting the document to be overwritten. Also, the sort instruction and the filter instruction may be applied by default and the document registration screen 1300 may be displayed.

[0112] Thus, in this embodiment, when the user selects a document to be overwritten and registered, the display order can be switched based on the similarity between the document registered in the extraction rule and the document indicated by the scanned image. In addition, the displayed documents can be filtered to display a list of documents that can be selected by the user. Therefore, the user can easily select the document to be overwritten and registered.

[0113] Radio buttons 1305 and 1306 are buttons for selecting information to be superimposed on the image of the document displayed in the preview area 1301. When radio button 1305 is selected, as shown by the dotted rectangle in the preview area 1301 of FIG. 13, the position of the index block can be highlighted and displayed. When radio button 1306 is selected, as shown in the preview area 1301b on the document registration screen 1300 of FIG. 14, the areas 1400 and 1401 where there are differences between the comparison target document and the input document are highlighted and displayed. The method for determining the different areas is as follows: First, the overall alignment is performed based on the text block groups of the comparison document and the input document, and the blocks with a high degree of overlap in the text block groups are determined as corresponding blocks. The text blocks for which no corresponding blocks are found are determined as the different areas. Alternatively, as a result of comparing the corresponding blocks of the comparison document and the input document, text blocks with a large size difference are determined as the different areas. Note that the method for determining the different areas is not limited to this method. For example, the overall alignment may be performed using the scanned images of the comparison document and the input document respectively, the difference in the average luminance value for each text block may be derived, and text blocks with a difference of a certain value or more may be determined as the different areas.

[0114] Thus, in this embodiment, the difference in the position of the index block and the difference in the document in the scanned image are displayed. Therefore, a screen for comparing the document selected as the target for overwriting and registration with the document indicated by the scanned image can be displayed, and the user can reduce the effort required for making a judgment on whether to instruct any of the processes of overwriting and registration, new registration, or not registering.

[0115] Returning to FIG. 13, the description of the document registration screen 1300 will be continued. The radio buttons 1307 to 1309 and the text fields 1310, 1311 have the same functions as the radio buttons 1201 to 1203 and the text fields 1204, 1207 on the registration confirmation screen 1200 in FIG. 12. Currently, the registration name of the document selected from the list 1302 is displayed in the text field 1310. When the OK button 1312 is pressed, the process selected by the radio buttons 1307 to 1309 is instructed. That is, on the document registration screen 1300, after comparing the image of the document registered in the extraction rule with the scanned image of the input document and comparing the positions of the index blocks, the user can instruct any of the processes of overwriting registration, new registration, and not registering.

[0116] Returning to FIG. 9, the description of the document registration process will be continued. In S906, the image processing unit 305 receives the user's instruction via the registration confirmation screen 1200 or the document registration screen 1300. For example, after the recommendation is made in S904 or S905, if the OK button 1208 on the registration confirmation screen 1200 or the OK button 1312 on the document registration screen 1300 is pressed, the instruction is received in this step.

[0117] In S907, the image processing unit 305 switches the process based on the instruction received in S906. When the instruction not to register is received, the document registration process is terminated. When the instruction to overwrite and register is received, the process proceeds to S908.

[0118] In S908, when the decision button 1208 on the registration confirmation screen 1200 is pressed, the image processing unit 305 executes an overwriting registration process using the document with the registration name displayed in the text field 1204 as the document to be overwritten. Or, when the decision button 1312 on the document registration screen 1300 is pressed, the image processing unit 305 executes an overwriting registration process using the document selected from the list 1302 as the document to be overwritten. As for the method of overwriting registration, for example, among the information held in the extraction rule associated with the document ID of the document to be overwritten, the "scan image" is updated with an image based on the scan image acquired in S400. The "document identification information" is updated with the information of the text block detected in S502. Furthermore, the "index information" is updated based on the position of the index block input in S402.

[0119] If it is determined in S907 that a new registration is to be performed, the process proceeds to S909. In S909, the image processing unit 305 generates a new unique value as the document ID. Then, for the "scan image", an image based on the scan image acquired in S400 is set. For the "document identification information", the information of the text block detected in S502 is set. Furthermore, for the "index information", the position information of the index block input in S402 is set. The set information is newly registered in the extraction rule in association with the generated document ID.

[0120] As described above, in this embodiment, when the user modifies the index, control is performed to display a screen for receiving an instruction to either perform an overwriting registration, a new registration, or no registration. Therefore, according to this embodiment, when an unregistered type of document is scanned, or when a document that is of the same type as a document type already registered by the issuer but has been partially modified is scanned, the effort of instructing an overwriting registration or a new registration can be reduced.

[0121] Moreover, simply presenting the user with a simple option of whether to perform overwriting registration or new registration as a separate document may not enable the user to determine which option to choose. In this embodiment, a process is performed to determine whether the document used in the index extraction process is of the same type as the document indicated by the scan image, and based on the result, the recommendation for either overwriting registration or new registration is switched. Therefore, according to this embodiment, the user can easily determine which of overwriting registration or new registration should be instructed.

[0122] As described above, according to this embodiment, by clearly presenting the user with the options of overwriting registration or new registration as a separate document, the information registered in the extraction rules can be appropriately updated. Therefore, the properties of the scan image can be appropriately set.

[0123] Note that the threshold value TH1 and the threshold value TH2 may be fixed common values in the document indicated by the scan image for which the similarity is calculated or the document group registered in the extraction rules. Alternatively, the values of the threshold values TH1 and TH2 may be changed according to the document indicated by the scan image or the registered document, and the change may be made during operation. For example, assume that there is a case where, although the similarity between the document indicated by the scan image and a certain registered document is higher than the threshold value TH1, new registration has been instructed by the user a certain number of times. In this case, the value of the threshold value TH1 for determining whether it is similar to the registered document may be increased. By changing the threshold value TH1 according to the document in this way, the frequency of erroneously recommending an index can be reduced.

[0124] Also, when new registration is instructed by the user even though the similarity between the document indicated by the scan image and a certain registered document is higher than the threshold value TH2, the threshold value TH2 for determining whether it is similar to the registered document may be increased in the same way.

[0125] Also, when no correction is made by the user in the confirmation and correction process of S402, the document determined in the document matching of S504 is considered to be of the same type as the input document. Therefore, the threshold value TH2 may be updated based on the average value and variance of the similarity when no index is input by the user in S402.

[0126] <Embodiment 2> In Embodiment 1, a method of determining whether the document determined in the index extraction process is similar to the document indicated by the scan image was described using a threshold value TH2 different from the threshold value TH1 used in the document matching in the index extraction process. However, if the threshold value TH2 cannot be set appropriately, an incorrect determination may be made. Therefore, in this embodiment, a method of determining whether the document determined in the index extraction process is similar to the document indicated by the scan image will be described based on the character string type. Note that this embodiment will be described mainly focusing on the differences from Embodiment 1. For parts not specifically mentioned, the configuration and processing are the same as those in Embodiment 1.

[0127] FIG. 15 is a flowchart of the file generation process of the scan image in this embodiment. The content of the process will be described mainly focusing on the differences from the file generation process of the scan image in Embodiment 1 (FIG. 4). S1500 to S1502 are the same as S400 to S402, so the description will be omitted. Also, S1504 is the same process as S404, and S1506 to S1507 are the same processes as S405 to S406, so the description will be omitted.

[0128] When it is determined in S1503 that the user has not input an index (S1503 is NO), it means that the index has been extracted appropriately. In this case, the process proceeds to S1505, where the image processing unit 305 performs a process of determining the character string type of the index in the document registered in the extraction rule. Specifically, based on the index extracted from the scan image acquired in S1500, a process for determining the character string type representing the characteristics of the index character string of each item of the document registered in the extraction rule is performed.

[0129] FIG. 16 is a diagram for explaining a method for determining a string type of an index of a certain document among the documents registered in the extraction rule. Table 1600 is a table for determining the string type of each index of the document with the document ID "0001" shown in FIG. 7. In this way, tables corresponding to the documents registered in the extraction rule are stored respectively. In the row group 1601 in Table 1600, strings indicating the indexes of the respective items extracted from the scan images determined to be similar to the document in FIG. 7 in the index extraction process have been held so far.

[0130] The details of the process of S1505 will be described. The image processing unit 305 acquires the table 1600 corresponding to the document determined to be similar to the document indicated by the scan image by document matching in the index extraction process. Then, a row is added to the row group 1601 of the table 1600, and the strings indicating the indexes extracted in the index extraction process are transcribed into the columns corresponding to the respective items. When the number of rows included in the row group 1601 exceeds a certain number, that is, when the number of scan images in which the indexes are transcribed exceeds a certain number, the image processing unit 305 determines the string type for each item.

[0131] Row 1602 is a row that holds the string type determined for each item. Row 1603 is a row for holding the details of the string type held in row 1602. The types of string types include, for example, fixed string type, numerical type, and presumptive type. Note that the determined string type is not limited to the fixed string type, numerical type, and presumptive type described above.

[0132] The fixed string type is, for example, in the index extraction process, in a scanned image determined to have a document and document format similar to each other, a string type when the string indicating the index is fixed. In Table 1600, as shown in row 1602, the items of title and sender are determined to be of the fixed string type. This is because when looking at the strings held in the "title" column in row group 1601, there are no strings other than "Qotation". Therefore, since the string is fixed, it is determined to be of the fixed string type. Thus, as shown by the string in the "title" column of row 1603, the string extracted as the index of the title is determined and held to be "Qotation" in any scanned image. Similarly, in the case of the sender, the string "ABC.Co" is held in row 1603.

[0133] The presumptive type is not the fixed string type. For example, it is a string type when a string conforming to a specific pattern is extracted as an index. In Table 1600, as shown in row 1602, the item of number is determined to be of the presumptive type. This is because when looking at the strings held in the "number" column in row group 1601, although all the strings are different, they are all composed of four-digit number sequences, so it is determined to be of the presumptive type. Also, as shown in row 1603, as the details of the string type of "number", "#" indicating that it is composed of four number sequences is held.

[0134] The numerical type is not the fixed string type or the presumptive type. It is a string type when the string indicating the index is extracted as a variable-length string composed only of numbers, commas, and dots. In Table 1600, as shown in row 1602, the item of total_price is determined to be of the numerical type. This is because when looking at the strings held in the "total_price" column in row group 1601, although all the strings are different, they are composed of at least one of numbers, commas, and dots.

[0135] In S902 of the present embodiment, it is determined whether the document indicated by the scanned image is similar to the document determined to be similar in the index extraction process, using the string type of each item determined in S1505.

[0136] FIG. 17 is a diagram for explaining the determination process in S902 of the present embodiment. In the table of FIG. 17, row 1701 in which the string type is held and row 1702 in which the details are held are the same as row 1602 and row 1603 in FIG. 16, respectively. That is, it is the string type determined as a result of performing the file generation process of FIG. 15 on a plurality of scanned images obtained by past scanning, and is the string type of each item of the document with the document ID of "0001" shown in FIG. 7.

[0137] In row 1703 of the table in FIG. 17, a string indicating the index of each item extracted from the scanned image of input document 1 in FIG. 10(b) is held. Assume that the document indicated by the scanned image of input document 1 is determined to be similar to the document with the document ID of "0001" shown in FIG. 7 as a result of the document matching in S504. Therefore, in S902, the determination of whether it is similar to the document with the document ID of "0001" is made based on the string type.

[0138] In FIG. 17, for the title and sender, which are fixed string type items, the string held in the details of line 1702 matches the extracted string held in line 1703. Also, for the number item, which is a presumed type item, since the extracted string held in line 1703 is composed of four number sequences, it matches the details of line 1702. Further, for the total_price item, which is a numerical type item, since the extracted string held in line 1703 is composed of a string with at least one of a number, a comma, and a dot, the string types match. Therefore, for all items, it is determined that the string type of the index extracted from the scan image matches the string type of the document determined to be similar in the index extraction process. Thus, in S902, it is determined that the document indicated by the scan image is similar to the document determined in S504 of the index extraction process. That is, it is determined that the document type of input document 1 is the same as the document type registered in the extraction rule.

[0139] On the other hand, in line 1704 of the table in FIG. 17, strings indicating the indexes of the respective items extracted from the scan image of input document 2 in FIG. 10(c) are held. Similar to input document 1, in S902, it is assumed that the determination as to whether the document with document ID "0001" is similar is made based on the string type.

[0140] In FIG. 17, for the issuer company name (sender), which is a fixed string type item, the string of the registered document held in the details of line 1703 is "ABC Co.". On the other hand, the string extracted from the scanned image of input document 2 is "LMN Co." as shown in cell 1705. Therefore, although it is a fixed string type, since the strings do not match, it is determined that the string types of the issuer company name (sender) do not match. Also, for the form number (number), which is a presumptive item, the extracted string as shown in cell 1706 is not composed of four number sequences, so it does not match the details of line 1702. Therefore, it is determined that the string types of the form number (number) do not match.

[0141] Thus, when there is a string that does not match the string type of the document determined to be similar in the index extraction process, in S902 of this embodiment, it is determined that the document indicated by the scanned image is not similar to the document determined in the index extraction process. That is, it is determined that the document type of input document 2 is a type not included in the document types registered in the extraction rules.

[0142] Note that as described above, if there is even one item with a non-matching string type, it may be determined as not similar, or if the number of non-matching string types is equal to or more than a predetermined threshold, it may be determined as not similar.

[0143] As described above, in this embodiment, similarity determination is performed between the document determined in the index extraction process and the document indicated by the scanned image using the determined string type of the index. Therefore, according to this embodiment, even when there is a document that is different from the input document but has a similar document layout to the input document, instead of overwriting registration, new registration can be recommended.

[0144] <Other Embodiments> In the above-described embodiment, an example has been described in which the image forming apparatus 100 performs the processes of each step of the flowchart in FIG. 4 or FIG. 15 alone. Alternatively, all or part of these processes may be performed by another image processing apparatus on the system 105 having the functions shown in FIG. 3.

[0145] For example, the scanning process is executed by the image forming apparatus 100, and the scanned image is transmitted to the terminal 101 via the network. If the terminal 101 has the same functions as the image processing unit 305, the terminal 101 may execute the index extraction process. In this case, the terminal 101 returns the index extraction result to the image forming apparatus 100, and the image forming apparatus 100 generates and transmits a file based on the acquired index extraction result.

[0146] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

Explanation of Reference Numerals

[0147] 100 Image forming apparatus 305 Image processing unit 111 CPU

Claims

1. An acquisition means for acquiring a scanned image obtained by scanning a document; A determination means for determining a document format similar to the format of the document indicated by the scanned image from among pre-registered document formats; Based on information registered in association with the document format determined by the determination means, identifying information on an area within the scanned image for setting properties of the scanned image, and a reception means for receiving correction of the information on the area; A display control means for performing a display for receiving from the user an instruction to perform a process of updating the information registered in association with the determined document format based on the correction, or an instruction to perform a process of newly registering by associating the format of the document indicated by the scanned image with the information based on the correction; A processing means for performing the update process or the newly registration process based on an instruction from the received user; having For correction of an area within a scanned image obtained by scanning, as the document, a first document similar to a predetermined document format, a display prompting the user to give an instruction to perform the newly registration process is performed; For correction of an area within a scanned image obtained by scanning, as the document, a second document more similar to the predetermined document format than the first document, a display prompting the user to give an instruction to perform the update process is performed An image processing apparatus characterized by the above.

2. A determination means for determining, based on a method with higher accuracy than the method by the determination means, whether the format of the document indicated by the scanned image is similar to the document format determined by the determination means, based on receiving correction of the information on the area; further comprising When the determination means determines that the format of the document indicated by the scanned image is not similar to the document format determined by the determination means, a display prompting the user to give an instruction to perform the newly registration process is performed; When the determination means determines that the format of the document indicated by the scanned image is similar to the document format determined by the determination means, a display prompting the user to give an instruction to perform the update process is performed The image processing apparatus according to claim 1, characterized by the above.

3. The predetermined document format is a document format determined as a document format similar to the format of the first document by the determining means, The second document is a document in a format determined by the determining means to be similar to the predetermined document format The image processing apparatus according to claim 2, characterized in that

4. The determining means determines, as the similar document format, a document format from among the registered document formats, the similarity of which to the format of the document indicated by the scanned image is greater than a first threshold value and which has the greatest similarity The image processing apparatus according to claim 2 or 3, characterized in that

5. The determining means determines that the format of the document indicated by the scanned image is similar to the document format determined by the determining means when the similarity between the format of the document indicated by the scanned image and the document format determined by the determining means is greater than a second threshold value that is greater than the first threshold value The image processing apparatus according to claim 4, characterized in that

6. A character string type representing a feature of a character string corresponding to a setting item of the property is registered in association with each of the registered document formats, The determining means determines that the format of the document indicated by the scanned image is similar to the document format determined by the determining means when the feature of the character string included in the specified area matches the character string type associated with the determined document format The image processing apparatus according to claim 2 or 3, characterized in that

7. The display control means displays, when the user makes a predetermined input, a list for the user to select a document type corresponding to the document format to be updated The image processing apparatus according to any one of claims 1 to 6, characterized in that

8. The display control means is configured to be able to control to display, in the list, a document type corresponding to a document format narrowed down based on the similarity between the format of the document indicated by the scanned image and each of the registered document formats The image processing apparatus according to claim 7, characterized in that

9. The display control means configured to be able to control to sort by the degree of similarity between the format of the document indicated by the scan image and each of the registered document formats, and display the document type corresponding to the document format in the list The image processing apparatus according to claim 7 or 8, characterized in that.

10. The display control means As a display prompting the user to give an instruction to perform the update process, a screen in a state where the document type corresponding to the document format determined by the determination means is selected as the document type corresponding to the document format of the update target is displayed The image processing apparatus according to any one of claims 1 to 7, characterized in that.

11. Even when the determination means cannot determine a document format similar to the format of the document indicated by the scan image, a display prompting the user to give an instruction to perform the newly registering process is performed The image processing apparatus according to claim 2, characterized in that.

12. The display control means Displays a confirmation screen on which the user can at least select whether to give an instruction to perform the update process or an instruction to perform the newly registering process, As a display prompting the user to give an instruction to perform the update process, in the confirmation screen, it is displayed in a state where a selection to give an instruction to perform the update process is automatically set The image processing apparatus according to claim 1, characterized in that.

13. In the registered document format, position information of an area for setting the property is registered in association therewith, The processing means performs a process of updating the position information of the area registered in association with the determined document format to the position information of the area specified by the user in the correction, or a process of newly registering the format of the document indicated by the scan image in association with the position information of the area specified by the user in the correction The image processing apparatus according to any one of claims 1 to 12, characterized in that.

14. Detection means for detecting an area including a character string in the scan image, and Area specifying means for specifying an area in the scan image for setting a property for the scan image The image processing apparatus according to any one of claims 1 to 13, further comprising.

15. An acquisition means for acquiring a character string in the scanned image, which is obtained by performing an optical character recognition process on a region in the scanned image for setting properties of the scanned image; The image processing apparatus according to any one of claims 1 to 14, further comprising the acquisition means. **Claim 16** The reception means further presents the acquired character string and receives correction of the character string. The image processing apparatus according to claim 15, characterized in that. **Claim 17** When the user corrects the presented character string, it is determined whether the format of the document indicated by the scanned image is similar to the document format determined by the determination means. The image processing apparatus according to claim 16, characterized in that. **Claim 18** An acquisition step of acquiring a scanned image obtained by scanning a document; A determination step of determining a document format similar to the format of the document indicated by the scanned image from among pre-registered document formats; Based on information registered in association with the document format determined in the determination step, specifying information on a region in the scanned image for setting properties of the scanned image, and receiving correction of the information on the region; A display control step of performing a display for receiving an instruction from the user to perform a process of updating the information registered in association with the determined document format based on the correction, or an instruction to perform a process of newly registering the format of the document indicated by the scanned image in association with the information based on the correction; A process step of performing the update process or the newly registration process based on the instruction received from the user; Comprising For a correction of a region in a scanned image obtained by scanning a first document similar to a predetermined document format as the document, a display prompting the user to give an instruction to perform the newly registration process is performed. For a correction of a region in a scanned image obtained by scanning a second document more similar to the predetermined document format than the first document as the document, a display prompting the user to give an instruction to perform the update process is performed. An image processing method characterized by that. **Claim 19** A program for causing a computer to function as each means of the image processing apparatus according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Method for discriminating document and method for registering document

    JP2002324236A

  • Information processing device and control method thereof

    JP2016091375A

  • Apparatus for setting filename for scan image, control method thereof, and program

    JP2019040250A

  • Computer and template management method

    JP2019159898A

  • Information processing apparatus, information processing method, and program

    JP2020107272A