Information processing apparatus, method, recording medium, and computer program product

By combining edge detection and string shape analysis with key item search, the problem of inaccurate original image extraction when the scanning device is not using black background paper is solved, and accurate original area extraction and cutting are achieved to adapt to different types of originals.

CN113259533BActive Publication Date: 2025-10-10FUJIFILM BUSINESS INNOVATION CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010903091.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-10
Filing Date
2020-09-01
Publication Date
2025-10-10
Estimated Expiration
2040-09-01

AI Technical Summary

Technical Problem

When a conventional scanning device does not use black background paper, it is difficult to accurately extract the edges of multiple original images, resulting in reduced cutting accuracy or the original images being merged into one large image, affecting the effect of the multi-cropping function.

Method used

The manuscript area is estimated by combining edge detection and character string shape analysis with key item search. Key items are determined based on the manuscript type, and formal estimation is performed to extract the accurate manuscript area.

Benefits of technology

Improves the cutting accuracy of the original image, ensures that each original image is correctly segmented, reduces the detection processing load, and adapts to different types of originals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113259533B_ABST
    Figure CN113259533B_ABST
Patent Text Reader

Abstract

The present application provides an information processing apparatus, a recording medium, and an information processing method, which can appropriately extract a document region from an input image photographed to a document, as compared with a case where a document region is not extracted from the input image according to a project of the document. The information processing apparatus is characterized by including a processor that performs processing of: receiving an input image including images of a plurality of documents; detecting one or more projects predetermined as projects included in the documents from the input image; and performing output processing of extracting and outputting an image of each document from the input image, based on the detected one or more projects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing device, a recording medium, and an information processing method. Background Art

[0002] Some scanners and multifunction peripherals (i.e., devices that combine scanner, printer, and copier functions) have a function that reads multiple documents placed on a document table (also called a platen), cuts out images of each document from the read images, and converts them into data. This function is called a multi-cropping function.

[0003] Conventional devices use a method such as covering a plurality of manuscripts placed on a manuscript table with black background paper to increase the contrast between the peripheral edges of the manuscripts and the background, thereby improving the cropping accuracy of each manuscript image.

[0004] However, it is common to forget to cover the top of the manuscript group with black background paper. Multifunction machines and the like are equipped with a manuscript cover that can be opened and closed relative to the manuscript table (in most cases, it has a built-in automatic manuscript feeding device), and the surface of the manuscript cover facing the manuscript table is usually white. If you forget to cover the multiple manuscripts on the manuscript table with black background paper and close the manuscript cover as usual to read them, the image of the read result will show a state where multiple white manuscripts are arranged on a white background. In most cases, the edges of the manuscripts will not be clearly shown in the image of the read result. When the edges of the manuscripts are not clear, the cutting accuracy of each manuscript image will be reduced. For example, errors such as cutting multiple separate manuscripts into one large manuscript may occur.

[0005] Furthermore, even when using a black background paper to improve cutting accuracy, errors can sometimes occur during cutting. For example, if multiple documents are neatly placed on the document table with no gaps or slightly overlapping, the multi-cut function may cut these documents into a single document.

[0006] The device described in Patent Document 1 acquires a regional image representing an area including a document placed in a reading area, inverts or rotates the regional image so that the arrangement of the document image included in the regional image matches the arrangement of the reading area when viewed from a predetermined direction, and then outputs the inverted or rotated regional image.

[0007] Patent Document 1: Japanese Patent Application Publication No. 2019-080166 Summary of the Invention

[0008] An object of the present invention is to appropriately extract a document area from an input image that captures a document, compared to a case where the document area is extracted from the input image not based on document items.

[0009] The invention described in Option 1 is an information processing device, characterized in that it is equipped with a processor, which performs the following processing: receiving an input image including images of a plurality of originals, detecting one or more items predetermined as items included in the originals from the input image, and performing output processing of extracting and outputting the images of each original from the input image based on the one or more items detected.

[0010] The invention described in Option 2 is characterized in that, in the information processing device described in Option 1, the one or more items predetermined as items included in the original include a plurality of items, and in the output processing, an image of a continuous area including all the plurality of items is extracted from the input image and output as an image of the original.

[0011] The invention described in Scheme 3 is characterized in that, in the information processing device described in Scheme 2, the processor performs a temporary estimation of the area of ​​each original included in the input image, and performs the detection and output processing on the image of each of the area parts obtained by the temporary estimation.

[0012] The invention described in Scheme 4 is an information processing device described in Scheme 3. In the output processing, for each of the areas obtained by the temporary estimation, for each continuous part that includes all the multiple items in sequence from one end of the area to the other end, the image of the part is extracted and output as an image of the original.

[0013] The invention described in Scheme 5 is an information processing device described in Scheme 3. In a case where patterns of a plurality of the areas are obtained in the temporary estimation, in the output processing, for each area belonging to a pattern adopted from the plurality of the patterns, an image of a continuous portion within the area including the plurality of items and divided by the boundaries between the areas in the patterns not adopted from the plurality of the patterns is extracted and output as an image of an original.

[0014] The invention described in Scheme 6 is characterized in that, in the information processing device described in any one of Schemes 1 to 5, the processor performs the following processing: obtaining type information representing the type of the original document included in the input image, and in the detection, detecting from the input image one or more items predetermined in correspondence with the type represented by the acquired type information.

[0015] The invention described in claim 7 is characterized in that, in the information processing apparatus described in claim 6, the processor performs processing for receiving selection of the one or more items from the user for each type of the document.

[0016] The invention described in Scheme 8 is a recording medium, which records a program for causing a computer to perform the following processing: receiving an input image including images of a plurality of originals, performing output processing of detecting one or more items predetermined as items included in the originals from the input image, and extracting and outputting images of each original from the input image based on the one or more items detected.

[0017] The invention described in Option 9 is an information processing method, characterized in that it includes the following steps: receiving an input image including images of a plurality of originals; performing detection of one or more items predetermined as items included in the originals from the input image; and performing output processing of extracting and outputting the images of each original from the input image based on the one or more items detected.

[0018] Effects of the Invention

[0019] According to the first, second, eighth, or ninth aspect of the present invention, the document area can be appropriately extracted from the input image capturing the document, compared to the case where the document area is extracted from the input image not based on the document's items.

[0020] According to the third or fourth aspect of the present invention, the processing load for detection can be reduced compared to the case where one or more items predetermined as items included in a document are detected from the entire input image.

[0021] According to the fifth aspect of the present invention, compared to a case where the region information of the pattern of the unselected provisional estimation result is not used, it is possible to extract an accurate region corresponding to the document.

[0022] According to the sixth aspect of the present invention, it is possible to extract a document area corresponding to the document type.

[0023] According to the seventh aspect of the present invention, it is possible to receive from the user a selection of items used when extracting a document area, depending on the document type. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Embodiments of the present invention will be described in detail with reference to the following drawings.

[0025] Figure 1 This is a diagram for explaining conventional multi-cropping processing using a black background sheet;

[0026] Figure 2 This is a diagram for explaining problems in conventional multi-cropping processing using a black background sheet;

[0027] Figure 3 FIG. 1 is a diagram illustrating a scanned image when a black background sheet is not used;

[0028] Figure 4 This figure is used to explain the problem of multi-cropping processing when a black background sheet is not used;

[0029] Figure 5 A diagram for explaining an outline of a method of an embodiment;

[0030] Figure 6 is a diagram illustrating the content of key project management information;

[0031] Figure 7 is a diagram illustrating a hardware configuration of an information processing device;

[0032] Figure 8 is a diagram illustrating the overall processing flow of the method according to the embodiment;

[0033] Figure 9 This is a diagram showing an example of the processing flow of the formal estimation process;

[0034] Figure 10 It is used for Figure 9 A diagram illustrating the formal presumption of the process;

[0035] Figure 11 This is a diagram showing another example of the processing flow of the formal estimation process;

[0036] Figure 12 It is used for Figure 11 A diagram illustrating the formal presumption of the process;

[0037] Figure 13 It is used for Figure 11 A diagram illustrating the formal presumption of the process;

[0038] Figure 14 This is a diagram for explaining how a plurality of original region patterns are obtained by provisional estimation and one of them is used as a provisional estimation result;

[0039] Figure 15 FIG. 1 is a diagram illustrating a characteristic portion of a formal estimation process using information on a document region of a pattern not adopted in a provisional estimation;

[0040] Figure 16 This is a diagram showing an example of a setting screen for the document determination method;

[0041] Figure 17 This is a diagram showing another example of a setting screen for the document determination method;

[0042] Figure 18 is a diagram showing an example of a detailed setting screen;

[0043] Figure 19 is a diagram showing an example of an estimation result screen that displays an official estimation result;

[0044] Figure 20 is a diagram showing another example of an estimation result screen that displays an official estimation result;

[0045] Figure 21 is a diagram showing another example of an estimation result screen that displays an official estimation result.

[0046] Symbol explanation

[0047] 10A to 10C - scanned image, 12a to 12f - original image, 14a to 14h - original area, 102 - processor, 104 - memory, 106 - auxiliary storage device, 108 - input / output device, 110 - network interface, 112 - bus, 114 - scanner control circuit, 116 - printer control circuit, 118 - facsimile device. DETAILED DESCRIPTION

[0048] <Multi-cropping processing and its problems>

[0049] Multi-cropping processing refers to a process of photographing a surface on which a plurality of original documents are arranged, and automatically extracting an image of each original document from an image obtained by the photographing and individually converting the image into a file.

[0050] Multi-cropping processing has been developed as a technology for a scanned image obtained by scanning with a scanner, a copier, or a multifunction peripheral (i.e., a device having functions of a scanner, a printer, a copier, and a facsimile device). Hereinafter, a scanner built in a separate scanner device, a copier, and a multifunction peripheral will be collectively referred to as a scanner. However, the technology of the present embodiment described below can be applied not only to a scanned image based on a scanner but also to an image photographed by various imaging devices (e.g., a smartphone, a digital camera).

[0051] Reference Figure 1 An example of the conventional multi-cropping processing will be described. In a case where multi-cropping processing is performed on original documents with backgrounds such as invoices and business cards being white, a group of original documents placed on a platen of a scanner is scanned with a black background sheet covering the group. A scanned image 10A obtained thereby includes original images 12a and 12b including characters or images in a white background within a black background. In Figure 1 In the example of the scanned image illustrated in

[0052] In Figure 2 is illustrated in the example of the scanned image Figure 1 A scan image 10B when two original documents of the same example are gaplessly arranged and scanned with their two side edges arranged on approximately the same straight line. Two original images 12a and 12b within the scan image 10B become one body to form a rectangle, and the edges between these two images are very shallow and are not detected by edge detection. In this case, in the conventional multiple cropping processing, an image within a region 14c of the smallest rectangle circumscribing the periphery of these two original images 12a and 12b, i.e., one original image 12c, is extracted instead of separately extracting the two original images 12a and 12b.

[0053] Figure 2 The example shown in FIG. 10C shows a case where an accurate original region is not extracted when scanned using a black background sheet.

[0054] On the other hand, an original cover that generally covers the scanner platen from the back is white. In order to perform multiple cropping processing, a special operation is required in which, instead of using the original cover, a black background sheet prepared separately is used to cover the top of the platen. Sometimes the user is reluctant to do this operation or does not know the necessity thereof and performs scanning as usual with the original cover covering the platen. In this case, even if the original documents are separated to some extent from each other, the images of these original documents are sometimes not separated but extracted as one image. Figure 3 and Figure 4 This example is shown in FIG. 10D.

[0055] Figure 3 The scan image 10C shown as an example shows a state in which the two original images 12a and 12b are arranged slightly apart from each other on the white background of the original cover. In this example, since the background and the background color of the original images 12a and 12b are the same white color, a clear edge does not easily appear between them. Therefore, the original images 12a and 12b are not easily extracted by edge detection.

[0056] Therefore, as exemplified in Figure 4 It is also possible to consider a method in which string shape analysis is performed on the scan image, and the region 14d of the original image 12d is extracted based on the information of the image objects such as strings obtained thereby.

[0057] In character string shape analysis, for example, through layout analysis or line clipping, a preprocessing step for character recognition in OCR (Optical Character Recognition) technology, lines 15 of character strings, etc., included in scanned image 10C are determined. Once the lines are determined, a coordinate system can be established using the line direction and directions perpendicular thereto as the x and y directions, and the coordinates of the character strings in each line within this coordinate system can be determined (for example, the coordinates of the circumscribed rectangle of the character strings in the line). By examining the x coordinates of the left ends of each line of character strings starting from the previous line (when writing from left to right), intervals with approximately the same x coordinates are determined to be within the same manuscript area. However, if the distance between adjacent lines is greater than a predetermined threshold, the intervals belonging to the previous line and the intervals belonging to the next line can be determined to be separate manuscript areas.

[0058] Furthermore, as another example of document area estimation processing using character string shape analysis, there is a process described in the specification, claims, and drawings of Japanese Patent No. 2019-229599 filed by the present applicant on December 19, 2019. In this process, the area of ​​the pixel group of the foreground (i.e., characters or images on a white background) is determined by sequentially applying expansion filters and contraction filters to the scanned image. The foreground areas belonging to the same document are determined based on the distance between these foreground areas or the area of ​​the gaps between these areas. Then, the set of foreground areas belonging to the same document is integrated into the area of ​​a single document image.

[0059] In character string shape analysis, the document area can be determined not only by the aforementioned analysis method based on character recognition results, but also by taking into account edge information extracted from the scanned image. Even if the extracted edges are thin or incomplete, by combining the results of the aforementioned analysis method based on character recognition results, the document area can be extracted with higher accuracy than using either the edge-based method or the character recognition-based analysis method alone.

[0060] In the estimation of the manuscript area based on the character string shape analysis, as shown in FIG. Figure 4 As illustrated in FIG, when the gap between the bottommost line 15 of the document image 12a and the topmost line 15 of the document image 12b is not sufficiently large as a result of two documents being placed close to each other, they are extracted as one document image 12d.

[0061] As described above, regardless of whether a black background sheet is used, there is a possibility that images of a plurality of documents that should be extracted separately may be extracted as one image.

[0062] <Overview of the Solution>

[0063] refer to Figure 5An overview of processing executed by the information processing apparatus of this embodiment to cope with such a situation will be described.

[0064] This process is performed on the document image 12d within the document area 14d, which is estimated by the estimation process using the aforementioned edge detection or character string shape analysis. In this process, the document image 12d is searched for words or phrases corresponding to items (hereinafter referred to as "key items") that are predetermined as items included in the document ("key item search" process in the figure).

[0065] Documents include various items such as name, company name, address, phone number, email address, product name, total amount, and credit card payment information. For documents of the same type, some items are assumed to be mandatory. These items are called key items. For example, in the case of an invoice, the issuer's company name, address, and total amount are examples of key items, while in the case of a business card, contact information such as name, company name, address, and phone number are examples of key items. For each document type, one or more key items are set.

[0066] exist Figure 6The example in Figure 2 illustrates key item management information for a business card-type document stored in an information processing device. This example includes columns such as an item ID, a detected flag, specific items, and a determination condition. The item ID uniquely identifies each key item. In this example, a name representing the item's meaning is used as the item ID, but this is merely a convenient example prioritized for ease of understanding. The detected flag indicates whether the sentence corresponding to the key item was detected in the document image. It is used to record the key item detected in the processing flow described below. If not detected, the flag's value is "OFF," and if detected, it is "ON." The specific item is the specific item corresponding to the key item. In particular, finding a sentence corresponding to at least one of multiple specific items confirms that the key item has been found. In other words, the multiple specific items included in a key item can be considered an OR condition used to determine if a key item has been found. For example, finding only a phone number, only an email address, or both constitutes the equivalent of finding the key item "contact information." The determination condition specifies, for each specific item, the conditions that the sentence corresponding to that item must meet. For example, among the conditions for the sentence corresponding to the key item "company name," there are conditions for sentences registered in a pre-prepared company / organization name database, conditions including a specified character string such as "stock company" or "(shares)," and the like. A sentence that satisfies at least one of these listed conditions is determined to correspond to the key item "company name" (i.e., an OR condition). In the determination condition column, it is possible to set a plurality of individual conditions defined by logical expressions including OR conditions, AND conditions, and the like. Furthermore, in the case where there are a plurality of specific items belonging to the key item, the determination condition for each specific item is set in the determination condition column.

[0067] Furthermore, when one key item includes a plurality of specific items constituting an OR condition, the key item management information may include a searched flag for the specific items in addition to the searched flag for the key item.

[0068] exist Figure 5In the example shown, for the document type "invoice," three key items are identified: company name, address, and total amount. During the key item search process, the character strings of each line are sequentially checked from the top or bottom of the document image 12d downward or upward to determine whether these character strings include a single sentence corresponding to the key item set for the document type. Then, a continuous interval in the document image row arrangement direction that includes sentences corresponding to these three key items and does not include different sentences corresponding to the same key item is inferred to be a document area. In the example shown, for example, the three key items: company name "FX Store," address "Roppongi xxx, Minato-ku, Tokyo," and total amount "Total ¥4,200" are found in sequence from the top of area 14d. If this third item is found, the interval from the top of area 14d, or the top of the first key item found, to the bottom of the third key item is determined to be area 14e of the first document image 12e. Then, the search continues downward, finding the three key items: company name "YMM cafe," address "Minato Mirai xxx, Nishi-ku, Yokohama, Kanagawa Prefecture," and total amount "2100 yen." The section from the top of the first key item, "YMM cafe," to the bottom of the last key item, "2100 yen," is determined to be region 14f of the second document image 12f.

[0069] As described above, in the present embodiment, the content of the document image 12 d is checked, and the regions 14 e and 14 f of the document images 12 e and 12 f are divided based on the key items included in the content.

[0070] In the example described below, a two-stage estimation process is performed: each document image region is estimated using edge detection or character string shape analysis, and a more rigorous document region is estimated from the estimated region using a search for key items. The former is referred to as provisional estimation, and the latter is referred to as formal estimation.

[0071] Hereinafter, an example of the hardware configuration of the information processing device according to the present embodiment and a specific example of the processing executed by the information processing device will be described.

[0072] <Hardware Structure>

[0073] The hardware configuration of the information processing device of this embodiment is shown in FIG. Figure 7 middle. Figure 7 The example described above is a case where the information processing apparatus is a so-called multifunction peripheral. A multifunction peripheral may also have a function of receiving requests from a client such as a personal computer via a network such as a local area network, or communicating with a server on the Internet.

[0074] For example, Figure 7As shown, the information processing device has the following circuit configuration: a processor 102 (hardware), a memory (primary storage device) 104 such as random access memory (RAM), a non-volatile storage device (auxiliary storage device) 106 such as flash memory, an SSD (solid state drive), or an HDD (hard disk drive), interfaces with various input / output devices 108, and a network interface 110 for controlling connection to a network such as a local area network, all connected via a data transmission path such as a bus 112. The input / output devices 108 include, for example, a display device / input device such as a touch panel, a voice output device such as a speaker, and a card reader for user authentication. The circuit configuration described above can be similar to that of a general-purpose computer.

[0075] The information processing device also includes a scanner control circuit 114, a printer control circuit 116, and a facsimile device 118, all connected to the computer portion via a bus 112 or the like. These devices are used for various functions of the information processing device (in this example, a multifunction peripheral). The scanner control circuit 114 controls the scanner or automatic document feeder built into the multifunction peripheral, while the printer control circuit 116 controls the printer built into the multifunction peripheral. Furthermore, the facsimile device 118 is a device that performs the facsimile transmission and reception functions of the multifunction peripheral.

[0076] The computer portion of the information processing device performs information processing for UI (user interface) processing, control of data exchange via a network, and control of various functional elements such as scanners, printers, and facsimile devices. Programs describing these various information processing functions are installed into the computer via a network or other means and stored in auxiliary storage device 106. The information processing device of this embodiment is implemented by processor 102 executing the programs stored in auxiliary storage device 106 using memory 104.

[0077] Here, the processor 102 refers to a processor in a broad sense, including general-purpose processors (such as CPU: Central Processing Unit, etc.), special-purpose processors (such as GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic devices, etc.).

[0078] Furthermore, the operations of the processor 102 may be performed not only by a single processor 102 but also by cooperation of multiple processors 102 located in physically separate locations. Furthermore, the operations of the processor 102 are not limited to the order described in the following embodiments and may be modified as appropriate.

[0079] The processing in this embodiment is performed on images captured by an imaging mechanism (e.g., a scanner connected to the scanner control circuit 114) included in the information processing device. Therefore, the information processing device may not include a printer, the printer control circuit 116 that controls it, or a facsimile device 118. The following description primarily uses a multifunction peripheral as an example, but this is merely an example. The information processing device may be any device with an imaging mechanism, such as a scanner, kiosk, smartphone, tablet, or personal computer.

[0080] <Overall Processing Flow>

[0081] refer to Figure 8 The overall processing flow of the method of this embodiment executed by the processor 102 of the information processing device will be described.

[0082] If the user places one or more originals on the table of the scanner attached to the information processing device and instructs the information processing device to start the "multi-cropping" process, the process will start. According to the instruction, the scanner of the information processing device performs scanning. The image obtained by the scanning (hereinafter referred to as the scanned image) is an image of the size of the entire surface of the table, which includes one or more original images. The scanned image becomes Figure 8 The object of the processing flow.

[0083] When instructing to start multi-cropping, the processor 102 may ask the user to specify the document type (eg, business card or invoice).

[0084] exist Figure 8 In the processing flow, processor 102 first performs background classification determination (S10). This determination is a process for determining whether the background of the scanned image is black (i.e., using a black background sheet) or white. In this determination, if the cumulative value or average value of the density of the pixel group in the peripheral portion of the scanned image is greater than a threshold value, the background is determined to be black; otherwise, the background is determined to be white.

[0085] If the result of S10 is "yes," that is, if the background is black, processor 102 tentatively estimates the document area based on edge detection (S14). This tentative estimation based on edge detection can be performed using known techniques. If the result of S10 is "no," processor 102 tentatively estimates the document area based on the aforementioned character string shape analysis (S16).

[0086] After provisionally estimating the document area (S14 or S16), the processor 102 performs a formal estimation of the document area (S18). The processor 102 then displays information on the estimation result obtained by the formal estimation on a display device connected to the information processing device (S19).

[0087] Example 1 of formal presumption

[0088] Will Figure 8 An example of the specific processing of S18 of the process is shown in Figure 9 . Figure 9 The process is executed for each document area temporarily estimated in S14 or S16.

[0089] In this process, processor 102 first performs character recognition on the document image within the document area of ​​the provisional estimation result (S20). If character recognition preprocessing has been completed during the provisional estimation (S14 or S16), character recognition is performed using the preprocessing results. Furthermore, in addition to character recognition, recognition of company logos, etc., can also be performed in S20. To this end, a database of registered company logos is prepared. For example, it is sufficient to determine whether an image that is not a character within the document image matches a logo in the database.

[0090] Then, the processor 102 sets the height of the upper end of the document area to be processed in the variable “area upper end height” ( S22 ).

[0091] Here, reference Figure 10 The coordinate system used in this processing is explained. In this example, the direction in which the rows in the manuscript area 14d to be processed extend (i.e., the direction from left to right in the figure) is set as the x-direction, and the direction perpendicular to it, i.e., the direction in which multiple rows are arranged in parallel, is set as the y-direction. The y-direction is the "height" direction. For ease of understanding, in the following description, the upper and lower directions are set as the directions shown in the figure. From another point of view, the direction indicated by the arrow in the y-direction shown in the figure is the "downward" direction, and the direction opposite to it is the "upward" direction. In addition, in this example, the vertex in the upper left corner of the manuscript area 14d is set as the origin of the coordinate system, but this is only an example.

[0092] The “area upper end height” (expressed as “y s”) is a variable that holds the y coordinate of the upper end of the document area 14d, which is the formal estimation result.

[0093] Then, based on the recognition result of S20 , the processor 102 sets the character string (or image such as a logo) in the row immediately below the upper end height of the region to the target object ( S24 ).

[0094] Next, the processor 102 determines whether the target object includes a statement corresponding to a key item (S26). In this determination, it is determined whether the statement satisfies the key item management information (reference Figure 6 ) judgment condition. If any judgment condition is satisfied, the judgment result of S26 is "Yes." Furthermore, processor 102 then identifies the key items and specific items corresponding to the judgment condition satisfied by the sentence. Furthermore, in this judgment, a model such as a neural network learned to judge the key items and specific items corresponding to the input sentence can be used instead of preparing explicit judgment conditions.

[0095] If the result of the determination in S26 is "No", the processor 102 proceeds to S38 and determines whether the next line of the target object exists in the document image 12d. If the result of the determination is "No", the processing is completed until the end of the document area 14d. Therefore, for example, the range from the upper end height of the area to the lower end of the document area 14d is extracted as one document area (S39), and then the processing is terminated. Figure 9 If the determination result of S38 is "Yes", the processor 102 changes the target object to the next row and repeats the processing after S26.

[0096] If the determination result of S26 is "Yes", the detected flag of the key item included in the target object identified in S26 is set to ON (S28). Figure 9 At the start of the process, the detected flags of all key items are turned off. Next, the processor 102 refers to the detected flags in the key item management information to determine whether all key items corresponding to the document type have been detected (S30).

[0097] If the result of the determination in S30 is "No", the processor 102 proceeds to S38, and if there is a next row of the target object, the target object is changed to the next row, and the processing after S26 is repeated.

[0098] When the judgment result of S30 is "yes", the processor 102 determines whether the specific item identified in S26 is the same as the specific item detected before S26 (S31). When the judgment result is "yes", all the key items that a manuscript image should include have been detected at this moment, and the specific items of the key items found this time are of the same type as the specific items of the key items included in the manuscript image. This means that the area of ​​one manuscript image has been checked, and the first row of the area of ​​the next manuscript image (= current target object) has also been found. In this case, the processor 102 extracts the range from the height of the upper end of the area in the y direction to the height of the upper end of the target object in the manuscript area 14d of the temporary estimation result as a manuscript area (S32a). The manuscript area extracted at this time is one of the formal estimation results. Then, the processor 102 changes the height of the upper end of the area to the height of the upper end of the current target object (= y coordinate) (S34). Then, the manuscript area extends downward from the height of the upper end of the area.

[0099] exist Figure 10 In the example case, in the process of detecting in sequence from the upper end of the manuscript area 14d, the sentences "FX store", "xxx, Roppongi, Minato-ku, Tokyo", and "a total of ¥4,200" corresponding to the key items and specific items "company name", "address", and "total amount" are found in sequence. The subsequent sentence "YMM cafe" corresponds to the key item and specific item "company name". Therefore, the judgment result of S31 is "yes" for the sentence "YMM cafe". Therefore, in S32a, the range from the upper end of the manuscript area 14d to the upper end of the area of ​​the sentence "YMM cafe" (for example, the smallest rectangle circumscribed with these character strings) is extracted in the manuscript area 14d as the first manuscript area. Then, in S34, the upper end of the area of ​​the sentence "YMM cafe" is set as the area upper end height of the next manuscript area.

[0100] After S34, the processor 102 resets the retrieved flags of all items in the key item management information to OFF (S36), and proceeds to the process of S38.

[0101] According to the above-mentioned formal presumption process, Figure 10 In the example shown, two document areas, document 1 and document 2 shown in the figure, are obtained as the final estimation results from the document area 14 d of the provisional estimation result.

[0102] Example 2 of Formal Presumption

[0103] refer to Figure 11 Another example of the process flow for formal presumption is described below. Figure 11 In the processing flow, Figure 9 The same steps in the processing flow are marked with the same symbols and repeated descriptions are omitted.

[0104] exist Figure 11 In the processing flow, Figure 9 S32a of the process is replaced with S32b, and S42 is inserted between S24 or S40 and S26. In this process, after S24 or S40, processor 102 sets the lower end of the target object to the variable "region lower end height" (S42). That is, each time the next line is found, the region lower end height is updated to the lower end of the found line. Then, if the result of S31 is "yes", processor 102 extracts the range from the region upper end height to the region lower end height at that moment in the manuscript area of ​​the provisional estimation result as a manuscript area (S32b).

[0105] According to this process, if Figure 12 As shown, first, the upper end of the manuscript area 14d of the provisional estimation result is set at the area upper end height y s (S22), then, the bottom of the top row of "FX Store" is set at the bottom of the area height y e (S42). Then, as the sentences of each line are checked sequentially from above, the height of the lower end of the area decreases line by line. Then, when the processing reaches the line of the sentence "YMMcafe", the judgment result of S31 is "yes". At this moment, the height of the lower end of the area is set to the height of the lower end of the previous line "Total ¥4200". Therefore, in S32b, the range from the upper end of the manuscript area 14d to the lower end of "Total ¥4200" is extracted as the first manuscript area. Then, the second manuscript area is extracted in the same way.

[0106] And, according to Figure 11 In the process of extracting the document area including the latter items when the document image includes items that are not key items after all key items. Figure 13 In the document image 12 shown in the example, after the three key items "FX store", "Tokyo Minato-ku Roppongi xxx", and "Total ¥4200" in the document 1, an item indicating a card number, "xyz card **********1234", which is not a key item, is included. Figure 11 The process extracts the document area from the upper end height of the area to the lower end of the item as the document area of ​​document 1.

[0107] Example 3 of Formal Presumption

[0108] This example is based on the following premise in the temporary estimation (S14 and S16): a plurality of patterns are obtained as the pattern of the manuscript area, and the best one of these plurality of patterns is selected as the temporary estimation result. For example, in the previous extraction process of the manuscript area using edge detection, for each pattern of the manuscript area obtained, a score representing the accuracy of the pattern is calculated. Then, the pattern with the highest score is automatically adopted, and the image of each manuscript area shown by the pattern is extracted and output. In the case of this method, when it becomes Figure 9 or Figure 11 The manuscript area of ​​the temporary estimation result of the object of the process may include a plurality of manuscript areas of another pattern that is not adopted. Figure 14 In the example shown, in addition to the pattern consisting of only the adopted document area 14 d , there is also a non-adopted pattern consisting of two document areas 14 g and 14 h . The processor 102 stores information on such non-adopted patterns in the memory 104 .

[0109] The characteristic part of the process of this example is shown in Figure 15 In. Figure 15 In the process shown, replace Figure 9 or Figure 11 The step group between S24 and S38 in the process.

[0110] In this process, a variable for the previous target object, ie, the previous object, is prepared. If the determination result of S26 is "No," the processor 102 sets the current target object as the variable for the previous object (S44), and proceeds to S38.

[0111] If the determination result of S26 is "Yes", the processor 102 executes the processes of S28 and S30. If the determination result of S30 is "No", the processor 102 executes the above-mentioned S44 and then proceeds to S38.

[0112] If the result of S30 is "yes," processor 102 determines whether there is a boundary between adjacent document areas in the unused pattern between the lower end of the previous object and the upper end of the target object (S46). If this result is "yes," processor 102 extracts the area from the height of the area top to the boundary within the document area being processed, i.e., the provisional estimated result, as the document area for the final estimated result (S48). Processor 102 changes the height of the area top to the height of the boundary (S50), clears the previous object (S52), and then proceeds to S38.

[0113] In addition, in S46, two lines may be detected as the boundary between the regions of the manuscript that do not adopt the pattern. Figure 14In the example shown in Figure 1, between "Total ¥4200" and "YMM cafe" within the unused pattern, there is a line at the bottom of manuscript area 14g and a line at the top of manuscript area 14h. In this case, in S48, the manuscript area is extracted from the height of the region's top edge to the upper line of these two lines. Then, in S50, the lower line of these two lines is set at the height of the region's top edge.

[0114] In the above description, based on Figure 9 or Figure 11 While the example of the process described above explains the search from the top to the bottom of the manuscript area of ​​the provisional estimation result, this search directionality is not essential for the method of Example 3. In this method, the boundaries between manuscript areas of unused patterns are used as the boundaries of the manuscript area of ​​the formal estimation result. Therefore, knowing the positions of each extracted key item is sufficient. There is no need to find information on the order in which each key item is executed.

[0115] As described above, in the method of this example, when extracting the original manuscript area included in the manuscript area of ​​the temporary estimation result, information of the manuscript area of ​​the pattern that was not adopted, that is, not selected as the temporary estimation result, is used. In Examples 1 and 2 of the above-mentioned formal estimation, the manuscript area is divided into line units of the character recognition result. Therefore, the manuscript area of ​​the formal estimation result does not include the white paper portion or the portion other than the key items included in the original manuscript image, or, on the contrary, includes the blank portion between the original manuscript images. In contrast, in the pattern that is not adopted as the temporary estimation result, sometimes a portion of the manuscript area that is not adopted in the comprehensive evaluation but includes a portion close to the periphery of the original manuscript image is included. In the method of this Example 3, by adopting the boundary of the manuscript area of ​​the pattern that is not adopted, it is possible to estimate the manuscript area more accurately than in the above-mentioned Examples 1 and 2.

[0116] <Settings screen example>

[0117] exist Figure 16 In the example of FIG. 1 , the information processing apparatus of this embodiment provides the user with a setting screen 200 for the document determination method in the multi-cropping process. In this setting screen 200 , a document type selection is received as information for determining the document determination method.

[0118] In the illustrated settings screen 200, two document types, "Invoice / Receipt" and "Business Card," are selected as alternatives. In settings screen 200, the judgment method for document type "Invoice / Receipt" is displayed with the following description: "※ The area is judged based on company name, address, and total amount." This indicates that the company name, address, and total amount are used as key items for formal estimation. Furthermore, settings screen 200 indicates that for document type "Business Card," company name, full name, address, and telephone number are used as key items for formal estimation.

[0119] Before instructing the start of the multi-cropping process, the user selects the document type to be processed this time on the setting screen 200 .

[0120] exist Figure 17 On the setting screen 200 illustrated in FIG. 1 , a button 202 for setting details for each selectable document type is displayed. When the user presses the button 202, the processor 102 displays a screen 220 (see FIG. 2 ). Figure 18 ), screen 220 receives detailed settings for the corresponding document type determination method. For example, if the user presses button 202 corresponding to document type "Business Card," screen 220 displays "Business Card" in the determination method name field 222. Screen 220 lists selectable items as key items, with a checkbox 224 to the left of each item indicating whether it is selected. Items with checkbox 224 displayed in black are selected as key items, while items with checkbox 224 displayed in white are selected as key items. If any of the selected items are unnecessary, the user can, for example, touch them to deselect them. Furthermore, if any of the deselected items are necessary as key items, they can be touched to select them. After selecting the desired key items, the user presses confirm button 226. This causes processor 102 to return to displaying setting screen 200. On the currently displayed setting screen 200, the selected item group is listed in the description field for the "Business Card" determination method on screen 220.

[0121] <Example of the display screen for the official estimation results>

[0122] Right Figure 8 An example of the estimation result screen 300 displayed on the display device included in the information processing device in S19 of the process will be described.

[0123] Figure 19The estimation result screen 300 shown in the figure displays document areas 14a and 14b of the formal estimation result superimposed on the scanned image 10. Document images 12a and 12b are displayed within the scanned image 10. In the illustrated example, the document areas 14a and 14b are in the form of frames that surround the key item groups within the corresponding document images 12a and 12b, respectively.

[0124] However, the display form of the frame line shown in the figure is just an example.

[0125] Figure 19 The arrangement of the document images 12a and 12b and the document areas 14a and 14b in the estimation result screen 300 shown in FIG. 1 is a mirror image arrangement when the user views the platen from above. Figure 19 In the arrangement of , it may be difficult for the user to understand the relationship between each document placed on the platen and each document area 14 a and 14 b in the estimation result screen 300 .

[0126] Therefore, in Figure 20 The estimation result screen 300 illustrated in FIG. 1 shows, within a background image 30 representing the range of the platen, images obtained by converting the document regions 14 a and 14 b obtained from the scanned image 10 into mirror-image arrangements as document region images 17 a and 17 b . Figure 20 The arrangement of the document area images 17a and 17b in the estimation result screen 300 corresponds to the arrangement of the two documents on the platen, so the user can easily understand the correspondence between the two. Figure 20 The estimation result screen 300 illustrated in FIG does not display the image contents of the document corresponding to the respective document region images 17 a and 17 b .

[0127] Therefore, in Figure 21 The estimation result screen 300 shown in FIG. Figure 20 In the estimation result screen 300 shown in the example, the corresponding original images 19a and 19b are displayed in the original area images 17a and 17b respectively. In this example, the original images 19a and 19b are aligned with the corresponding original area images 17a and 17b by rotating the images in the original areas 14a and 14b in the scanned image 10 in the same plane, so that the user can easily understand them intuitively. An image in which the frame lines of the original area images 17a and 17b are superimposed on an image obtained by mirror-transforming the scanned image 10 (i.e., with the inside and outside reversed) can be displayed on the estimation result screen. However, compared with this, Figure 21 The shown image makes it easier for the user to intuitively understand which area corresponds to which original.

[0128] The above describes the structure and processing of the embodiment. However, the above structure and processing examples are merely illustrative. Various modifications and improvements are possible within the scope of the present invention. For example, in the processing examples described above, processing is performed from the top to the bottom of the document area of ​​the provisional estimation result, but processing can naturally be performed from the bottom to the top.

[0129] The above-described embodiments of the present invention are provided for the purpose of illustration and explanation. In addition, the embodiments of the present invention do not fully encompass the present invention, and do not limit the present invention to the disclosed embodiments. It is obvious that various modifications and variations are self-evident to those skilled in the art to which the present invention belongs. The present embodiment is selected and described in order to most easily explain the principles of the present invention and its application. Thus, other persons skilled in the art will be able to understand the present invention through various modifications optimized for specific uses of the assumed various embodiments. The scope of the present invention is defined by the following claims and their equivalents.

Claims

1. An information processing device comprising a processor, The processor performs the following processing: receiving an input image including images of a plurality of originals; detecting, from the input image, one or more items predetermined as items included in the document; The information processing apparatus is characterized in that, based on the one or more detected items, an output process of extracting and outputting an image of each document from the input image is executed. The one or more items predetermined as items included in the manuscript include a plurality of items. In the output process, an image of a continuous region including all of the plurality of items is extracted from the input image and output as one original image.

2. The information processing device according to claim 1, wherein The processor performs a temporary estimation of an area of ​​each document included in the input image, The detection and output processing are performed on the image of each of the region portions obtained by the provisional estimation.

3. The information processing device according to claim 2, wherein In the output process, for each of the regions obtained by the provisional estimation, an image of each continuous portion including all of the plurality of items in sequence from one end toward the other end of the region is extracted and output as one original image.

4. The information processing device according to claim 2, wherein When patterns of a plurality of the regions are obtained in the provisional estimation, In the output processing, for each area belonging to a pattern adopted from the plurality of patterns, an image of a continuous portion within the area including the plurality of items and divided by boundaries between the areas in the plurality of patterns not adopted is extracted and output as an image of the original.

5. The information processing device according to any one of claims 1 to 4, characterized in that The processor performs the following processing: acquiring type information indicating a type of the document included in the input image; In the detection, the one or more items predetermined in association with the category indicated by the acquired category information are detected from the input image.

6. The information processing device according to claim 5, wherein: The processor performs the following processing: A selection of the one or more items is received from a user for each type of the document.

7. A recording medium having recorded thereon a program for causing a computer to perform the following processing: receiving an input image including images of a plurality of originals; detecting, from the input image, one or more items predetermined as items included in the document; Based on the one or more detected items, an output process of extracting and outputting an image of each document from the input image is executed, wherein the recording medium is characterized in that: The one or more items predetermined as items included in the manuscript include a plurality of items. In the output process, an image of a continuous region including all of the plurality of items is extracted from the input image and output as one original image.

8. An information processing method comprising the following steps: receiving an input image including images of a plurality of originals; detecting one or more items predetermined as items included in a document from the input image; and Based on the one or more detected items, an output process of extracting and outputting an image of each document from the input image is executed, wherein the information processing method is characterized in that: The one or more items predetermined as items included in the manuscript include a plurality of items. In the output process, an image of a continuous region including all of the plurality of items is extracted from the input image and output as one original image.

9. A computer program product configured to cause a computer to perform the following functions: receiving an input image including images of a plurality of originals; detecting, from the input image, one or more items predetermined as items included in the document; The computer program product is characterized in that, based on the one or more detected items, an output process of extracting and outputting an image of each document from the input image is executed. The one or more items predetermined as items included in the manuscript include a plurality of items. In the output process, an image of a continuous region including all of the plurality of items is extracted from the input image and output as one original image.

Citation Information

Patent Citations

  • Image processing device and control method

    JP2019080166A

  • Image processing device and image processing program

    JP7447472B2

  • Document processing system using full image scanning

    US20030059098A1