Information processing device, image orientation determination method, information processing system, area determination method, and program
The information processing device enhances image orientation determination accuracy by using predetermined format images and determination features to match image regions, addressing the challenges of conventional methods in accurately determining image orientation.
Patent Information
- Application Number
- JP2021072545
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-28
- Filing Date
- 2021-04-22
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2041-04-22
AI Technical Summary
Conventional image orientation determination methods often fail to accurately determine the orientation of scanned documents and images, especially when features are difficult to recognize or high accuracy is required, leading to incorrect image orientation.
An information processing device that uses predetermined format images and determination features to determine image uprightness by matching features between partial regions in the image and a reference image, incorporating learning images to enhance accuracy.
Improves the accuracy of image orientation determination by using predetermined format images and determination features, enabling precise upright determination of scanned documents and images.
Smart Images

Figure 0007784237000001 
Figure 0007784237000002 
Figure 0007784237000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to techniques for determining image orientation. [Background technology]
[0002] A conventional method has been proposed for determining the up-down direction of an input image more accurately without incorporating special equipment for detecting tilt into an image acquisition device, in which an object candidate detection means detects object candidates from the input image and the angles of the object candidates in the input image, a similarity calculation means calculates the similarity between each detected object candidate and each object stored in advance, an input image angle calculation means determines the up-down direction of the input image based on the calculated similarity of each object candidate and the angle in the input image, and the input image angle calculation means, for example, weights the angle in the input image of each object candidate based on the similarity and calculates the tilt angle with respect to the up-down direction of the input image using the weighted angle (see Patent Document 1).
[0003] Furthermore, a method has been proposed in the past that includes an extraction unit that extracts groups of feature points that are commonly included in at least two of a plurality of first images from a plurality of first images, each of which contains orientation information indicating the top-to-bottom orientation of the image; a detection unit that detects, from the groups of feature points extracted by the extraction unit, groups of feature points that are distributed in a unique positional relationship with the top-to-bottom orientation as the reference in at least two of the first images; a search unit that searches for the groups of feature points detected by the detection unit in a second image that does not have orientation information; and a determination unit that determines the orientation of the second image based on a comparison between the positional relationship of the groups of feature points found by the search unit and the unique positional relationship corresponding to the group of feature points (see Patent Document 2).
[0004] An image recognition device has been proposed in which an area division unit divides the binarized image data output from an image signal processing unit into multiple areas, a reliability determination unit calculates the reliability of each divided area when used for top-bottom recognition, and an top-bottom recognition unit extracts character data from the area with the highest reliability value and performs top-bottom recognition processing (see Patent Document 3).
[0005] A form dictionary generation device has been proposed that includes a feature extraction means that extracts feature information indicating the features of a form from each of multiple template images of forms designated as the same type; a common information generation means that generates common information indicating the features of common cells formed by common lines in multiple template images based on the feature information extracted for each template image; an ignored area determination means that, for each common cell, determines, based on the common cell information and the feature information of each template image, if there is a cell in each template image that corresponds to the common cell and has features different from the common cell, determine the area of the common cell as an ignored area in form identification; and a dictionary generation means that generates data including the common information and information indicating the ignored area as dictionary data for identifying forms (see Patent Document 4). [Prior art documents] [Patent documents]
[0006] [Patent Document 1] International Publication No. 2007 / 142227 [Patent Document 2] Japanese Patent Application Laid-Open No. 2014-134963 [Patent Document 3] Japanese Patent Application Laid-Open No. 2000-032247 [Patent Document 4] Japanese Patent Application Laid-Open No. 2010-262578 Summary of the Invention [Problem to be solved by the invention]
[0007] Conventionally, when scanning documents such as forms, the resulting scanned image may not be upright. This requires users to visually check the orientation of the scanned image and manually rotate the misoriented image to the desired orientation. To streamline this process, a technology has been proposed that performs optical character recognition (OCR) processing on scanned images to determine whether the characters are upright. In addition to characters, other objects, such as human faces and two-dimensional barcodes, are also targets. It is common for these objects, including all of these, to be used to determine whether the recognized objects are upright. However, in cases where the image has features that are difficult to recognize with the previously anticipated recognition method, or when extremely high accuracy is required, the previously anticipated recognition method often fails, resulting in incorrect image orientation determination.
[0008] In view of the above-mentioned problems, an object of the present disclosure is to improve the accuracy of image orientation determination. [Means for solving the problem]
[0009] An example of the present disclosure is an image having a predetermined format. A predetermined format image and an upright position determination means for determining whether or not a feature corresponding to the determination feature exists at the position of an input image to be determined, thereby determining whether or not the image to be determined is upright. The determination features include feature points and feature amounts, and the upright determination means performs matching using the determination features between the partial region in the image in the predetermined format and a region related to the position in the image to be determined, thereby determining whether or not a feature corresponding to the determination features exists at the position in the image to be determined. It is an information processing device. Furthermore, one example of the present disclosure is an information processing device comprising: a determination information storage means for storing determination features relating to a predetermined partial region in a predetermined format image, which is an image having a predetermined format, and the position of the partial region when the predetermined format image is upright; an upright determination means for determining whether an input determination target image is upright by determining whether a feature corresponding to the determination feature exists at the position of the input determination target image; an image accepting means for accepting input of a learning image having the predetermined format; a candidate extraction means for extracting, from the learning image in an upright state, a region that is likely to be included in common with images having the predetermined format, and setting the extracted region as a candidate for the partial region; and a designation accepting means for accepting designation of the partial region in the learning image, wherein the designation accepting means accepts designation of the partial region by a user who refers to the extracted candidate, and the determination information storage means stores the determination features relating to the partial region accepted by the designation accepting means and the position of the partial region when the learning image is upright, as the determination features and the position. Furthermore, one example of the present disclosure is an information processing system that determines a partial region to be used for uprightness determination, which determines whether an input image to be determined is upright, based on determination features related to a partial region in an image having a predetermined style and the position of the partial region when the image having the predetermined style is upright. The information processing system includes: an image accepting means that accepts input of a plurality of learning images having the predetermined style; a common region extraction means that extracts one or more regions common to the plurality of learning images in the upright state as partial region candidates that are candidates for partial regions to be used for the uprightness determination; a reliability determining means that determines a common region reliability for each of the one or more partial region candidates, which indicates the degree to which the partial region candidate is suitable as a partial region to be used for the uprightness determination; and a region determining means that determines one or more partial regions to be used for the uprightness determination from the one or more partial region candidates based on the common region reliability.
[0010] The present disclosure can be understood as an information processing device, a system, a method executed by a computer, or a program executed by a computer. The present disclosure can also be understood as such a program recorded on a recording medium readable by a computer, other device, machine, etc. Here, a recording medium readable by a computer, etc. refers to a recording medium that stores information such as data and programs by electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer, etc. [Effects of the Invention]
[0011] According to the present disclosure, it is possible to improve the accuracy of determining the orientation of an image. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a schematic diagram illustrating a configuration of a system according to an embodiment. [Figure 2] 1 is a diagram illustrating an outline of a functional configuration of an information processing apparatus according to a first embodiment. [Figure 3] FIG. 10 is a diagram illustrating an example of a feature region before expansion according to the embodiment. [Figure 4] FIG. 10 is a diagram illustrating an example of a feature region after expansion according to the embodiment. [Figure 5] 10A and 10B are diagrams illustrating an example of an upright orientation determination process (with matching process) for a determination target image (rotation angle: 0 degrees) according to an embodiment. [Figure 6] 10A and 10B are diagrams illustrating an example of an upright orientation determination process (with matching process) for a determination target image (rotation angle: 90 degrees) according to an embodiment. [Figure 7] 10A and 10B are diagrams illustrating an example of an upright orientation determination process (without matching process) for a determination target image (rotation angle: 180 degrees) according to an embodiment. [Figure 8] 10A and 10B are diagrams illustrating an example of an upright orientation determination process (with matching process) for a determination target image (rotation angle: 270 degrees) according to an embodiment. [Figure 9] 10A and 10B are diagrams illustrating examples of relative positions of matching points related to correct determination according to the embodiment. [Figure 10] 10A and 10B are diagrams illustrating examples of relative positions of matching points related to erroneous determination according to the embodiment. [Figure 11] 10A and 10B are diagrams illustrating an example of changing the order of feature regions used when performing the upright orientation determination process according to the embodiment. [Figure 12] 10 is a flowchart showing an outline of the flow of determination information registration processing according to the first embodiment. [Figure 13] FIG. 10 is a schematic diagram illustrating an example of a registration screen for rotating a learning image according to the embodiment. [Figure 14] FIG. 10 is a schematic diagram illustrating an example of a registration screen for determining a feature region according to the embodiment. [Figure 15] 1 is a flowchart (1) showing an outline of the flow of an upright orientation determination process according to an embodiment. [Figure 16] 10 is a flowchart (2) showing an outline of the flow of the upright orientation determination process according to the embodiment. [Figure 17] FIG. 10 is a diagram illustrating an outline of the functional configuration of an information processing apparatus according to a second embodiment. [Figure 18] FIG. 10 is a diagram illustrating an example of a matching result according to the embodiment. [Figure 19] 10A and 10B are diagrams illustrating an example of a method for generating paired feature region candidates according to the embodiment; [Figure 20] 10A to 10C are diagrams illustrating an example of a method for generating feature region candidates according to the embodiment. [Figure 21] 10A and 10B are diagrams illustrating examples of feature region candidates according to the embodiment; [Figure 22] 10 is a flowchart showing an outline of the flow of determination information registration processing according to the second embodiment. [Figure 23] 1 is a flowchart (1) showing an overview of the flow of a common area reliability determination process according to the embodiment. [Figure 24] 10 is a flowchart (2) showing an outline of the flow of the common area reliability determination process according to the embodiment. [Figure 25] FIG. 10 is a diagram showing an outline of the functional configuration of an information processing device according to first to seventh modifications of the embodiment. [Figure 26]FIG. 13 is a diagram illustrating an outline of the functional configuration of an information processing system according to an eighth modification of the embodiment. [Figure 27] FIG. 10 is a diagram illustrating an outline of the functional configuration of an information processing device according to a third embodiment. [Figure 28] 13 is a flowchart showing an outline of the flow of determination information registration processing according to the third embodiment. [Figure 29] 13 is a flowchart showing an outline of the flow of determination information registration processing according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of an information processing device, an image orientation determination method, an information processing system, a region determination method, and a program according to the present disclosure will be described with reference to the drawings. However, the embodiments described below are merely examples, and the information processing device, image orientation determination method, information processing system, region determination method, and program according to the present disclosure are not limited to the specific configurations described below. When implementing the present disclosure, a specific configuration may be appropriately adopted depending on the implementation mode, and various improvements and modifications may be made. In the present embodiment, an embodiment will be described in which the information processing device, image orientation determination method, information processing system, region determination method, and program according to the present disclosure are implemented in a system that determines the orientation of a scanned document image. However, the information processing device, image orientation determination method, information processing system, region determination method, and program according to the present disclosure can be widely used in technologies for determining the orientation of a captured image, and the application of the present disclosure is not limited to the examples shown in the present embodiment.
[0014] [First embodiment] <System configuration> 1 is a diagram showing an outline of the configuration of a system according to this embodiment. The system according to this embodiment includes an information processing device 1 and an image acquisition device 9 that are connected to a network and can communicate with each other.
[0015] The information processing device 1 is a computer including a central processing unit (CPU) 11, a read-only memory (ROM) 12, a random access memory (RAM) 13, a storage device 14 such as an electrically erasable and programmable read-only memory (EEPROM) or a hard disk drive (HDD), a communication unit such as a network interface card (NIC) 15, an input device 16 such as a keyboard or a mouse, and an output device 17 such as a display. However, the specific hardware configuration of the information processing device 1 may be omitted, replaced, or added as appropriate depending on the embodiment. Furthermore, the information processing device 1 is not limited to a device consisting of a single housing. The information processing device 1 may be realized by multiple devices using so-called cloud or distributed computing technology, etc.
[0016] The image acquisition device 9 is a device that acquires images, and examples thereof include scanners and multifunction devices that acquire document images by reading documents such as forms, and imaging devices such as digital cameras and smartphones that capture images of people, landscapes, etc.
[0017] 2 is a diagram showing an outline of the functional configuration of the information processing device 1 according to this embodiment. The information processing device 1 functions as an information processing device including a determination information storage unit (determination information database) 21, an image reception unit 22, a designation reception unit 23, a feature extraction unit 24, a candidate extraction unit 25, a rotation unit 26, an upright position determination unit 27, an orientation correction unit 28, and a display unit 29, by a program recorded in the storage device 14 being read into the RAM 13 and executed by the CPU 11, which controls each piece of hardware included in the information processing device 1. Note that in this embodiment and other embodiments described below, each function included in the information processing device 1 is executed by the CPU 11, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.
[0018] The information processing device 1 according to this embodiment is implemented on the cloud, for example, and receives learning images having a predetermined format from the image acquisition device 9, thereby learning information for determining the orientation of the image having the predetermined format (determination features related to characteristic regions, positions of the characteristic regions in an upright state), and storing the information in the determination information database 21. Then, when the information processing device 1 receives a determination target image (an image whose upright orientation is to be determined) from the image acquisition device 9, it determines whether a feature corresponding to the determination feature exists in a position corresponding to the characteristic region in the determination target image, thereby determining whether the determination target image is upright.
[0019] The determination information storage unit (determination information database) 21 stores (registers) information (determination information) for determining the orientation of an image having a predetermined format (image in a predetermined format). The image in a predetermined format is an image having a predetermined image at a predetermined position in an upright state, such as an image of a document having a predetermined format, such as a form, or other captured image (camera image, etc.). In this embodiment, the determination information storage unit 21 stores, as the determination information, determination features related to a characteristic region (region having a characteristic), which is a predetermined partial region of the image in the predetermined format, and the position of the characteristic region when the image in the predetermined format is upright. Specifically, when a characteristic region is designated in a training image having a predetermined format, the determination information storage unit 21 stores the determination feature related to the designated characteristic region and the position of the characteristic region when the training image is upright, as the determination feature and the position related to the image in the predetermined format.
[0020] For example, an area including a company name, logo, document title, or notes is selected as a characteristic area, and information for determining the similarity between the image included in the area and a comparison image, such as image data, feature points, and feature amounts, related to the characteristic area, is stored as a determination feature. The determination information storage unit 21 stores determination features related to one or more images in a predetermined format. In addition, in this embodiment, one characteristic area is registered for one image in a predetermined format, but this is not limited thereto, and multiple characteristic areas may be registered for one image in a predetermined format.
[0021] Furthermore, the determination information storage unit 21 may store determination features by defining an area (extended area) obtained by adding a peripheral area to an area specified (selected) by the user as a characteristic area. In this embodiment, the user selects an area in the learning image to determine the characteristic area. However, depending on the feature point extraction method, it may not be possible to extract feature points within a predetermined range inside the edge of the selected area. Therefore, in this embodiment, the extended area is defined as the characteristic area so that feature points related to the entire area specified by the user are detected. However, the characteristic area is not limited to this extended area, and the area specified by the user itself may be defined as the characteristic area.
[0022] In addition, when the peripheral area exceeds the edge of the specified format image, that is, when the expanded area extends outside the edge of the document, the judgment information storage unit 21 may add a margin area as a peripheral area for the part that exceeds the edge and store judgment features.
[0023] Fig. 3 is a diagram showing an example of a feature region before expansion according to this embodiment. As shown in Fig. 3, in the feature region before expansion, feature points at the edge of a character, such as the top, may not be detected. Therefore, in this embodiment, for example, a region expanded by a predetermined width Xpx (e.g., 30px) on all four sides of the region designated by the user is set as the feature region.
[0024] Fig. 4 is a diagram showing an example of an expanded feature region according to this embodiment. As shown in Fig. 4, in the expanded feature region, it is possible to detect feature points in the entire region (entire character) specified by the user.
[0025] In the example of Figure 4, the area specified by the user is the top area of the document, and the peripheral area located above the specified area extends beyond the edge of the image, so a margin area is added as the peripheral area. In this case, the margin area added from the edge of the document may be adjusted to a width (e.g., 15px) smaller than the width (Xpx) of the peripheral area to be added so that black areas at the edge of the document (thick black lines (bands) or U-shaped or L-shaped areas) are not detected as feature points. For example, in Figure 4, because the user specified the top area of the document, a margin area (peripheral area) with a width of 15px is added to the top of the specified area, and peripheral areas with a width of 30px are added to the bottom, right, and left.
[0026] Furthermore, the determination information storage unit 21 may store the determination features at a resolution adjusted according to the size of the feature region. In a site where a large number of scans are performed, it is necessary to quickly complete the scan process, including determining the image's uprightness. However, if the size of the feature region to be registered is large, the amount of data related to the determination features used for the uprightness determination increases, which may result in a long time required for the uprightness determination (scanning process). Therefore, in this embodiment, the determination features are stored at a resolution adjusted according to the size of the feature region to prevent a decrease in processing speed while extracting feature points and feature amounts related to the feature region. In this case, the resolution is adjusted so that the larger the size of the feature region, the lower the resolution.
[0027] For example, for a small feature region such as 0.3 x 0.3 inches, the determination features are stored at a predetermined resolution (e.g., 300 dpi), and for a large feature region such as 2.0 x 2.0 inches, the determination features are stored at a resolution lower than the predetermined value (e.g., 100 dpi). By adjusting the resolution according to the size of the feature region in this way, it becomes possible to quickly complete the scanning process, including the upright orientation determination, even when the size of the feature region to be registered is large. Note that in this embodiment, the resolution is adjusted according to the size of the feature region, but this is not limiting, and the same resolution may be used for all feature regions regardless of their size.
[0028] The determination information storage unit 21 may also store the size of the image in the predetermined format as the determination information. In the uprighting determination process, for example, if an uprighting determination is performed using the determination information when a document unrelated to the registered document is scanned, the processing time increases and there is a possibility of an erroneous determination. Therefore, in this embodiment, the size (width, height) of the image in the predetermined format (image for learning) is stored as the determination information so that uprighting determination is not performed on images related to documents that are not registered. This makes it possible to determine that uprighting determination is not performed on images to be determined that do not match or approximate the size of the stored image in the predetermined format.
[0029] The image receiving unit 22 receives input of learning images for learning (acquiring) information for determining the orientation of an image, and a determination target image (an image whose uprightness is to be determined) for determining the image orientation. The learning images and the determination target images are, for example, images related to documents such as forms, or other captured images (camera images, etc.). Note that, in this embodiment, the image receiving unit 22 receives input of these images from the image acquisition device 9, but is not limited to this, and may also acquire (accept input of) images stored in advance in the storage device 14.
[0030] The specification receiving unit 23 receives a specification of a characteristic region in a learning image by the user. For example, the specification receiving unit 23 receives the specification of a characteristic region by the user selecting a candidate of the characteristic region extracted by the candidate extracting unit 25. Note that the method in which the user selects (specifies) a characteristic region from the candidates (feature proposals) extracted by the candidate extracting unit 25 is hereinafter referred to as "auto."
[0031] Alternatively, the specification receiving unit 23 may receive the specification of a characteristic region by specifying a range within the learning image through a manual operation (such as a mouse operation) by the user. Note that the method of specifying a range related to a characteristic region by the user will be referred to as "manual" hereinafter.
[0032] Furthermore, the designation receiving unit 23 may receive designation of a feature region by the user selecting a feature region candidate extracted by the candidate extraction unit 25 in the proposal target region. Specifically, the designation receiving unit 23 receives from the user a range designation relating to a region in the training image from which a feature region candidate is to be extracted (proposal target region). Then, when the candidate extraction unit 25 extracts a feature region candidate in the proposal target region, the designation receiving unit 23 receives the feature region candidate by the user selecting the candidate. Note that hereinafter, the method in which the user selects (designates) a feature region from the candidates extracted in the designated proposal target region is referred to as "semi-automatic."
[0033] The feature extraction unit 24 extracts determination features (feature points, feature amounts, image data, etc.) related to a feature region designated in the learning image, and the position (x coordinate, y coordinate, etc.) of the feature region when the learning image is upright. The feature extraction unit 24 also extracts features (feature points, feature amounts, image data, etc.) from the determination target image (including a rotated image) at a position corresponding to the position of the feature region stored in the determination information storage unit 21. Specifically, the feature extraction unit 24 extracts an area (a region used for upright orientation determination (determination area)) related to a position corresponding to (corresponding to) the position of the feature region from the determination target image, and then extracts feature points, feature amounts, etc. from the determination area. For example, the feature extraction unit 24 extracts a determination area of the same size as the feature region from a position corresponding to (corresponding to) the position of the feature region in the determination target image.
[0034] Note that the extraction of feature points and feature amounts can be performed using known methods, such as SIFT (Scale-invariant feature transform), SURF (Speed-up robust feature), A-KAZE (Accelerated KAZE), etc. For each extracted feature point, the feature extraction unit 24 calculates the feature amount in a local region centered on that feature point.
[0035] The feature extraction unit 24 may extract feature points and feature quantities by setting the determination area to an expanded area (for example, an area expanded by 0.5 inches in all directions) obtained by adding a peripheral area to an area corresponding to a registered feature area (an area having the same position and size as the feature area). This makes it possible to extract feature points and feature quantities at positions corresponding to the positions of the feature areas in the determination target image, even if positional deviation occurs in the determination target image due to correction errors in the scanner device or scan image processing. The feature extraction unit 24 may also extract feature points and feature quantities related to the determination area after adjusting (matching) the resolution related to the determination area to the resolution related to the feature area.
[0036] The candidate extraction unit 25 extracts regions that are likely to be included in images having a predetermined style from the training images in an upright state, and sets these as candidate feature regions. For example, the candidate extraction unit 25 extracts multiple rectangular character regions from the training images having a predetermined style, and calculates a score for each rectangular character region taking into consideration the position within the image, the area of the rectangle, the aspect ratio, etc., and extracts the region with the highest score as a region (candidate feature region) that is likely to be included in images having the predetermined style.
[0037] For example, an area at the top of the image that has a large area and a small aspect ratio (which can be assumed to contain a character string with a large font size and a small number of characters) is likely to contain a company name, manuscript name, etc., and is therefore calculated to have a high score. In this way, the candidate extraction unit 25 calculates the score by performing a conversion process based on the position, area, aspect ratio, etc. of the rectangular area so that areas that are assumed to be more likely to contain a company name, manuscript name, etc., have a higher score. The candidate extraction unit 25 then extracts the area with the highest calculated score as a candidate characteristic area. Note that the candidate extraction unit 25 may extract multiple candidate characteristic areas by extracting multiple areas (e.g., the top three areas) in descending order of score. Alternatively, the candidate extraction unit 25 may extract candidate characteristic areas related to the same style based on multiple learning images related to the same style.
[0038] In addition, areas with large variations in gradation (unsmooth variations in gradation), large objects with connected pixels (such as areas with concentrated black pixels in binarization), and areas common to multiple learning images of the same predetermined format may be extracted as areas that are likely to be included in images of the predetermined format. For example, a large object with connected pixels is estimated to be a company logo, etc.
[0039] The rotation unit 26 rotates the image within a range of 0 degrees or more and less than 360 degrees. The rotation unit 26 rotates the image to be determined at one or more angles at which the outer edge shape of the image to be determined matches the outer edge shape of the image in the predetermined format when it is upright. For example, if the outer edge shape of the image to be determined matches the outer edge shape of the image in the predetermined format when it is upright at angles of 0 degrees, 90 degrees, 180 degrees, or 270 degrees (such as in the case of a square), the rotation unit 26 rotates the image to be determined at least at one of these angles. Furthermore, if the image in the predetermined format is a rectangular image (such as a rectangle) having long and short sides, the rotation unit 26 rotates the image to be determined at least at two angles (for example, 0 degrees and 180 degrees, or 90 degrees and 270 degrees) at which the relationship between the long side or short side and the vertical side or horizontal side matches the image in the predetermined format when it is upright. Furthermore, the rotation unit 26 rotates the learning image in an upright orientation in response to a user instruction to rotate in an upright orientation.
[0040] The upright orientation determination unit 27 determines whether or not a feature corresponding to the determination feature of the predetermined format image exists in a position (determination area) in the determination target image corresponding to the position of the feature area in the predetermined format image, for the determination target image at each rotation angle.The upright orientation determination unit 27 then determines whether or not the determination target image is upright by determining the orientation of the determination target image (the rotation angle with respect to the predetermined format image) based on the determination result for the determination target image at each rotation angle.The processing performed by the upright orientation determination unit 27 will be described in detail below.Note that the predetermined format images registered in the determination information storage unit 21 will be referred to as registered images hereinafter.
[0041] The upright orientation determination unit 27 determines whether or not a feature point exists in a determination area within the determination target image for each rotation angle (e.g., 0, 90, 180, and 270 degrees). If a feature point is extracted in the determination area, the upright orientation determination unit 27 performs a matching process (feature point (feature amount) matching) between the registered image and the determination target image based on the feature point and feature amount. The matching process between two images can be performed using a known method, such as a brute-force method or a fast approximate neighbor search method (FLANN, Fast Library for Approximate Nearest Neighbors). In this embodiment, if no feature point is extracted in the determination area, the matching process is not performed. However, this is not limited to this, and the matching process may be performed even when no feature point is extracted, in the same way as when a feature point is extracted. However, in this case, since no feature point is detected, the number of matching feature points is determined to be zero.
[0042] For example, the upright orientation determination unit 27 calculates the distance in feature point space between the feature amount related to a feature point in the registered image (feature region) and the feature amount related to a feature point in the determination target image (determination region), and determines the point with the smallest calculated distance or the point below a threshold as the matching (corresponding) feature point. For example, if the feature amount is a SIFT feature amount, the distance between two points in a 128-dimensional space is calculated. The upright orientation determination unit 27 calculates the number of matching feature points between the two images (between the feature region and the determination region) through this feature point matching.
[0043] The uprightness determination unit 27 also compares feature amounts related to feature points in the registered image (feature region) with feature amounts related to feature points in the determination target image (determination region) to determine the degree of similarity between the two points. The uprightness determination unit 27 calculates the similarity of the feature amounts between the two points by, for example, performing a conversion process based on the distance between the feature amounts of the matched feature points so that the closer the feature amount distance, the higher the similarity. Note that, for example, a distance such as Euclidean distance or Hamming distance is used as the distance between the feature amounts. The uprightness determination unit 27 calculates the similarity of the feature amounts between the feature region (registered image) and the determination target image (the similarity of the feature amounts between the two images) using the similarity of the feature amounts calculated for each matched feature point. For example, the similarity of the feature amounts calculated for each matched feature point may be calculated as a representative value such as the average, median, or mode, or a total value, etc., of the similarity of the feature amounts calculated for each matched feature point.
[0044] The upright orientation determination unit 27 determines whether the target image is upright or not based on the results of the matching process (the number of matched feature points and the similarity of the feature amounts). For example, the upright orientation determination unit 27 performs the upright orientation determination by comparing the number of matched feature points and the similarity of the feature amounts with a predetermined threshold. Specifically, for the target image at each rotation angle, the unit 27 compares the number of matched feature points and the similarity of the feature amounts with a predetermined threshold, and determines the rotation angle of the target image at which both exceed the threshold as the orientation of the target image.
[0045] For example, when a matching process is performed between a determination target image with a rotation angle of 0 degrees and a registered image, if the number of matched feature points and the similarity of the feature amounts exceed respective predetermined thresholds, the orientation of the determination target image is determined to be upright (the angle difference with the registered image is 0 degrees). Also, for example, when a matching process is performed between a determination target image rotated 90 degrees counterclockwise and a registered image, if the number of matched feature points and the similarity of the feature amounts exceed respective predetermined thresholds, the orientation of the determination target image (before rotation) is determined to be not upright (it is rotated 90 degrees clockwise with respect to the registered image). Note that the upright orientation determination unit 27 may set a plurality of the predetermined thresholds.
[0046] Furthermore, the uprightness determination unit 27 may determine the uprightness of the image to be determined at each rotation angle by determining a value indicating the likelihood of the image being upright at each rotation angle according to the number of matched feature points and the similarity of the feature amounts. In this case, the uprightness determination unit 27 determines the orientation of the image to be determined by comparing the values indicating the likelihood of the image being upright at each rotation angle. In this embodiment, the number of votes corresponding to the number of matched feature points is used as the value indicating the likelihood of the image being upright, but this is not limited to this as long as it indicates the likelihood of the image being upright. For example, the number of votes corresponding to the similarity of the feature amounts, or a score based on the number of matched feature points and / or the similarity of the feature amounts may be used.
[0047] The upright determination unit 27, for example, compares the number of votes for each rotation angle and determines the rotation angle with the highest number of votes as the angle of the image to be determined relative to the registered image, thereby determining whether the image to be determined is upright.
[0048] FIG. 5 is a diagram showing an example of the upright orientation determination process (with matching process) for a determination target image (rotation angle: 0 degrees) according to this embodiment. As shown in FIG. 5, in the case of a determination target image with a rotation angle of 0 degrees, as a result of the matching process between the determination region (determination target image) and the feature region (registered image), it is determined that six feature points match. In the example of FIG. 5, since the determination target image is upright, it is determined that there are six feature points matching with the registered feature region at multiple points. In this case, the upright orientation determination unit 27 determines the number of votes for a rotation angle of 0 degrees to be "6," which is the number of matched feature points.
[0049] 6 is a diagram showing an example of the upright orientation determination process (with matching process) for the determination target image (rotation angle: 90 degrees) according to this embodiment. As shown in FIG. 6, when the determination target image is rotated 90 degrees counterclockwise, the result of the matching process between the determination area (determination target image) and the feature area (registered image) determines that the feature points do not match. In this case, the upright orientation determination unit 27 determines the number of votes for the rotation angle of 90 degrees to be "0."
[0050] FIG. 7 is a diagram showing an example of the upright orientation determination process (without matching process) for the determination target image (rotation angle: 180 degrees) according to this embodiment. As shown in FIG. 7, when the determination target image is rotated 180 degrees, no image is formed in the determination area of the determination target image, and no feature points are detected from the determination area. For example, if the determination area is included in a white area, feature points may not be detected. In such a case, in this embodiment, the matching process is not performed for the 180 degree rotation angle, and the image is excluded from the comparison of the number of votes (the number of votes is not determined).
[0051] 8 is a diagram showing an example of the upright orientation determination process (with matching process) for the determination target image (rotation angle: 270 degrees) according to this embodiment. As shown in FIG. 8, when the determination target image is rotated counterclockwise by 270 degrees, the matching process between the determination region (determination target image) and the feature region (registered image) determines that one feature point matches. In this case, the upright orientation determination unit 27 determines the number of votes for the rotation angle of 270 degrees to be "1," which is the number of matched feature points.
[0052] 5, 6, and 8, and determines (adopts) the rotation angle of 0 degrees, which has the highest number of votes, as the rotation angle of the image to be determined relative to the registered image. This enables the upright determination unit 27 to determine that the image to be determined is upright.
[0053] In this embodiment, the number of matched feature points and the similarity of feature amounts are each compared with a threshold value, but the present invention is not limited to this, and the uprightness determination may be performed by comparing only one of them with a predetermined threshold value. Also, the number of votes may be corrected using the result of the uprightness determination using another method.
[0054] Furthermore, the upright determination unit 27 may determine whether or not a feature corresponding to the determination feature exists in accordance with the resolution adjusted according to the size of the feature region. Specifically, the upright determination unit 27 may determine whether or not a feature corresponding to the determination feature stored at the adjusted resolution exists in a determination region whose resolution has been adjusted according to the size of the feature region of the registered image. This makes it possible to quickly complete the scanning process even when the size of the feature region to be registered is large, as described above.
[0055] Furthermore, the upright orientation determination unit 27 may perform a determination on, among the input determination target images, those that match or are similar in size to the predetermined format image stored in the determination information storage unit 21. As a result, as described above, it becomes possible to quickly complete the scan process without performing an upright orientation determination on images for which determination information is not registered.
[0056] Furthermore, the upright orientation determination unit 27 may be configured to detect an erroneous determination. Because feature point matching by the upright orientation determination unit 27 is performed point by point, there is a possibility that points that do not relate to the same determination feature may match. As a result, a determination target image having a different style from the registered image may be erroneously determined to have a feature corresponding to the determination feature of the registered image, or a determination target image having the same style as the registered image may be erroneously determined not to have a feature corresponding to the determination feature of the registered image.
[0057] FIG. 9 is a diagram showing an example of relative positions of matching points related to a correct determination according to this embodiment. As shown in FIG. 9, because the determination target image is upright, the matching process determines that the registered image and the determination target image match at many feature points. As such, when the matching process is performed correctly (no erroneous determination), it can be seen that the relative positions of the matched feature points (the positions of the feature points in the determination target image relative to the feature points in the registered image) generally match for all feature points. The relative positions are the positions of the feature points in the determination target image relative to the positions (X-axis, Y-axis coordinates, etc.) of the feature points in the registered image when the determination target image and the registered image are arranged side by side, as shown in FIGS. 9 and 10. For example, the relative positions are represented by the distance and slope (or vector) of the line segments connecting the matched feature points between the registered image and the determination target image, as shown in FIGS. 9 and 10. In the example of Figure 9, it can be seen that the distance and slope of the line segments connecting the matched feature points between the registered image and the image to be determined are roughly the same for all combinations of matched feature points (five pairs in the example of Figure 9).
[0058] FIG. 10 is a diagram showing an example of the relative positions of matching points in an erroneous determination according to this embodiment. As shown in FIG. 10, even though the image to be determined is not upright, there are cases where feature points match with the registered image due to an erroneous determination. However, in the case of an erroneous determination, as shown in FIG. 10, it can be seen that the relative positions of the matching feature points do not match for all feature points. In other words, in the example of FIG. 10, it can be seen that the distance and slope of the line segment connecting the matching feature points between the registered image and the image to be determined do not match for all combinations of the matching feature points (two pairs in the example of FIG. 10).
[0059] Therefore, the uprightness determination unit 27 detects erroneous determination by determining the relative positions of matched points (feature points in the registered image and feature points in the image to be determined) through a matching process (feature point matching). For example, the uprightness determination unit 27 first removes outlying points (noise) from the matched feature points, and then calculates the distance (difference) between the matched points and the slope of the line connecting the matched points for each feature point. Then, based on the distance and slope for each feature point, the unit calculates the variances of the distance and slope, respectively, and obtains the reliability of the determination from these variances. If the reliability is equal to or less than a threshold, the unit determines that the determination is erroneous.
[0060] The uprightness determination unit 27 removes, as outliers (noise), feature points with fewer adjacent feature points than a predetermined number or feature points with a distance between matching points exceeding a predetermined value, for example. Furthermore, the less likely the determination is to be erroneous, the closer the relative positions of the feature points are to match and the closer the variances of distance and tilt are to 0. Therefore, the uprightness determination unit 27 calculates the reliability by, for example, performing a conversion process on each of the variances of distance and tilt such that the closer the variances are to 0, the higher the reliability. For example, the reliability is calculated by a method of finding the complement of a normalized variance or a method of associating a rank with each range of variance values. Alternatively, instead of the reliability, the variances of distance and tilt may be compared with a threshold to determine whether or not the determination is erroneous.
[0061] Furthermore, the uprightness determination unit 27 may dynamically change the order of feature regions (registered images) used for the uprightness determination. When the determination information storage unit 21 stores determination features related to multiple images (feature regions) in a predetermined format, the uprightness determination process for each feature region is performed sequentially, which may prevent the scan process, including the uprightness determination, from being completed quickly. Therefore, by dynamically changing the order of feature regions (registered images) used for the uprightness determination, the scan process, including the uprightness determination, can be completed quickly. Specifically, the uprightness determination unit 27 determines the order of the images in the predetermined format used for the uprightness determination of the image to be determined, based on the uprightness determination result of the image to be determined that was upright before the image to be determined.
[0062] For example, at a scanning site, documents of the same format (for example, the same form) are often scanned in large quantities, and in this case, by first performing the uprighting determination process with a registered image (characteristic region) related to that document, the uprighting determination process with other registered images (characteristic regions) becomes unnecessary. Therefore, when there is a registered image (characteristic region) that matches the image to be determined (determined to be an image related to the same format), the uprighting determination unit 27 first performs the uprighting determination process with the matching registered image for the next image to be determined.
[0063] 11 is a diagram showing an example of changing the order of feature regions used in the upright orientation determination process according to this embodiment. As shown in FIG. 11, upright orientation determination unit 27 performs the upright orientation determination process on a determination target image in the order of feature regions (registered images) A, B, C, and D. However, if the second determination target image matches feature region C, the feature region that will be first subjected to the upright orientation determination process on the third determination target image is designated as feature region C. Similarly, if the third determination target image matches feature region B, the feature region that will be first subjected to the upright orientation determination process on the fourth determination target image is designated as feature region B. Similarly, in the fifth and sixth determination target images, since the fourth and fifth determination target images match feature region B, the feature region that will be first subjected to the upright orientation determination process on the fifth and sixth determination target images is designated as feature region B.
[0064] In the example of FIG. 11, the third to sixth target images are images of the same format (for example, images of the same form), and by dynamically changing the order of the feature regions, the fourth to sixth target images can be determined as upright by performing the upright determination process on only feature region B, which enables the scanning process to be completed quickly.
[0065] Orientation correction unit 28 corrects the orientation of the determination target image based on the determination result by upright orientation determination unit 27. For example, when it is determined that the determination target image is an image rotated clockwise by R degrees (e.g., 90 degrees) with respect to the registered image, orientation correction unit 28 corrects the orientation by rotating the determination target image by R degrees (e.g., 90 degrees) in the opposite direction (counterclockwise) to the rotation direction of the determination target image with respect to the registered image so that the determination target image is upright.
[0066] The display unit 29 executes various display processes via the output device 17 in the information processing device 1. For example, the display unit 29 generates a registration screen or the like on which the user registers determination information related to an image in a predetermined format (image for learning), and displays (outputs) the generated screen via the output device 17 such as a display.
[0067] <Processing flow> Next, the flow of processing executed by the information processing device according to this embodiment will be described using a flowchart. Note that the specific content and processing order of the processing shown in the flowchart described below are an example for implementing the present disclosure. The specific content and processing order may be selected as appropriate depending on the embodiment of the present disclosure.
[0068] 12 is a flowchart showing an outline of the flow of the determination information registration process according to this embodiment. The determination information registration process according to this embodiment is executed when a user presses a new registration button or the like on a registration screen in the information processing device 1 and issues an instruction to scan a form for which new determination information is to be registered. Note that in this embodiment, an image related to the form is acquired by scanning the form, but the captured image is not limited to a form, and may be an image related to another document, a photograph, or the like.
[0069] In step S101, input of an image in a predetermined format (learning image) is accepted. When the image acquisition device 9 performs a scan process on the form for which judgment information is to be registered, the image accepting unit 22 acquires an image (learning image) related to the scanned form from the image acquisition device 9. Note that the image accepting unit 22 may acquire a pre-stored learning image from the storage device 14 in response to an instruction from the user to acquire a learning image. Thereafter, the process proceeds to step S102.
[0070] In step S102, the learning image is rotated to an upright orientation in response to a user instruction. When rotation unit 26 receives an instruction from the user to rotate the learning image to an upright orientation, rotation unit 26 rotates the learning image whose input was received in step S101 to an upright orientation. The image rotation instruction from the user may be received by an instruction input receiving unit (not shown).
[0071] 13 is a schematic diagram showing an example of a registration screen for rotating a learning image according to this embodiment. As shown in FIG. 13, for example, if a user visually determines that an acquired learning image can be upright by rotating it 180 degrees and presses the "Rotate 180 degrees" button, the rotation unit 26 rotates the learning image 180 degrees. Similarly, if the user presses the "Rotate Left 90 Degrees" button or the "Rotate Right 90 Degrees" button, the rotation unit 26 rotates the learning image 90 degrees counterclockwise (leftward) and 90 degrees clockwise (rightward), respectively. Thereafter, the process proceeds to step S103.
[0072] In step S103, candidate feature regions are extracted. The candidate extraction unit 25 extracts, from the upright learning image, regions that are likely to be included in images having a predetermined style as candidate feature regions (automatic processing). The candidate extraction unit 25 extracts one or more candidate feature regions. Thereafter, the process proceeds to step S104.
[0073] In step S104, it is determined whether the extracted feature region candidate has been selected by the user. The designation receiving unit 23 determines whether the feature region candidate extracted in step S103 has been selected by the user on a registration screen or the like.
[0074] FIG. 14 is a schematic diagram showing an example of a registration screen for determining a characteristic region according to this embodiment. As shown in FIG. 14, the designation receiving unit 23 determines that a candidate for a characteristic region extracted in step S103 has been selected, for example, when the user presses one of the buttons "Candidate 1" to "Candidate 3." If a candidate for a characteristic region has been selected, the process proceeds to step S109. On the other hand, if a candidate for a characteristic region has not been selected, the process proceeds to step S105. Note that when a candidate for a characteristic region has been selected by the user, the candidate for a characteristic region may be displayed in "image of characteristic region" on the registration screen by the display unit 29, for example, as shown in FIG. 14.
[0075] In step S105, a designation of manual or semi-automatic is accepted as the extraction method for feature region candidates. For example, as shown in FIG. 14, the designation accepting unit 23 accepts a designation that the candidate extraction method is manual when the user presses the "Custom" button. In this case, the processing proceeds to step S106. Also, when the user presses the "Candidate (range designation)" button, the designation accepting unit 23 accepts a designation that the candidate extraction method is semi-automatic. In this case, the processing proceeds to step S107.
[0076] In step S106, a range designation for a characteristic region by the user is accepted. The designation accepting unit 23 accepts the range designation for a characteristic region by a manual operation by the user, such as a mouse drag operation. As in step S104, the designated range (region) may be displayed in the "image of characteristic region" on the registration screen shown in Fig. 14. Then, the process proceeds to step S109.
[0077] In step S107, a proposal target region is designated by the user. The designation receiving unit 23 receives a range designation for the proposal target region (a region from which a candidate feature region is to be extracted) by a manual operation by the user, such as a mouse drag operation. Then, the process proceeds to step S108.
[0078] In step S108, feature region candidates are extracted from the proposal target region. As in step S103, within the proposal target region designated in step S107, the candidate extraction unit 25 extracts regions that are likely to be commonly included in images having a predetermined style as feature region candidates. Note that, as in the case of automatic processing, multiple feature region candidates may be extracted from the proposal target region, or only one may be extracted. Also, as in step S104, the extracted feature region candidates are displayed in "Feature Region Image" on the registration screen, as shown in FIG. 14, for example. Thereafter, the process proceeds to step S109.
[0079] In step S109, a characteristic region to be used for orientation correction is determined from the regions extracted automatically, manually, or semi-automatically. The designation receiving unit 23 receives input for determining the characteristic region, for example, by selecting one of "Candidate 1" to "Candidate 3," "Candidate (Specified Range)," and "Custom" on the registration screen shown in FIG. 14 and then pressing the "Register" button. Before determining the characteristic region, the user may press a "Test Scan" button or the like to check whether another document of the same format can be correctly rotated (corrected) to an upright orientation using the selected characteristic region candidates ("Candidate 1" to "Candidate 3," "Candidate (Specified Range)," and "Custom"). Then, the process proceeds to step S110.
[0080] In step S110, determination information related to the confirmed characteristic region is registered (stored). The feature extraction unit 24 extracts determination features (feature points, feature amounts, etc.) related to the characteristic region confirmed in step S109 and the position of the characteristic region when the learning image is upright. For example, the feature extraction unit 24 extracts (calculates) the feature points and feature amounts related to the characteristic region using a method such as SIFT, SURF, or A-KAZE. The determination information storage unit 21 then stores the determination features, including the feature points, feature amounts, and image data related to the characteristic region, in association with the position of the characteristic region when the learning image is upright.
[0081] At this time, the determination information storage unit 21 may store the determination features at a resolution adjusted according to the size of the characteristic region. The determination information storage unit 21 may also store the determination features by treating the region (extended region) obtained by adding the surrounding region to the region confirmed (specified) by the user in step S109 as the characteristic region. Furthermore, the determination information storage unit 21 may store the size of the learning image in association with the determination features and the positions of the characteristic region. After that, the processing shown in this flowchart ends.
[0082] Note that, although the feature region candidates are extracted (selected) by the branching process shown in steps S103 to S108 above, the present invention is not limited to this, and any branching process can be executed as long as feature region candidates are extracted by one or more of the following methods: automatic, manual, and semi-automatic, and a feature region is selected from among them. Therefore, the process of automatically extracting feature region candidates (step S103) may be performed only when a user instruction to automatically extract feature regions (automatic process) is received. For example, if a user instruction to manually select a feature region is received after step S102, the process of step S103 need not be performed.
[0083] 15 and 16 are flowcharts showing an outline of the flow of the uprighting determination process according to this embodiment. The uprighting determination process according to this embodiment is executed in response to, for example, a user issuing a scan instruction for a form for which uprighting determination is desired in the information processing device 1. Note that in this embodiment, an image related to the form is acquired by scanning the form, but this is not limited to this as long as it is a captured image, and images related to other documents, photographs, etc. may also be used.
[0084] In step S201, input of a determination target image is accepted. When a scan process is performed on a form to be subjected to upright orientation judgment in image acquisition device 9, image acceptance unit 22 acquires an image of the scanned form (determination target image) from image acquisition device 9. Note that image acceptance unit 22 may acquire a pre-stored determination target image from storage device 14 in response to an instruction from a user to acquire the determination target image. Thereafter, the process proceeds to step S202. Note that in step S201, upright orientation judgment unit 27 may not execute the subsequent processes (steps S202 to S219) if the acquired determination target image does not match or approximate the size of the predetermined format image stored in determination information storage unit 21.
[0085] In step S202, the target image is rotated. The rotation unit 26 rotates the target image to any angle at which the outer edge shape of the target image matches the outer edge shape of the image in an upright state in a predetermined format. Each time the target image is rotated to any angle in step S202, the processing (repeated processing) of steps S203 to S213 is executed, and this repetitive processing is executed until it is determined in step S214 that rotation has been performed for all angles at which the outer edge shapes match. For example, as shown in FIGS. 5 to 8, if the registered image and the target image are square and have matching outer edge shapes, in step S202, the target image is rotated to any angle at which the outer edge shapes match (0 degrees, 90 degrees, 180 degrees, 270 degrees), and extraction of a determination area and matching processing are executed in steps S203 to S213. Thereafter, the processing proceeds to step S203. However, in this embodiment, if the orientation of the image to be determined is determined in step S208, the repeated processing is stopped.
[0086] In step S203, an area to be used for determining whether the image is upright is extracted. The feature extraction unit 24 extracts an area (determination area) corresponding to the feature area at a position corresponding to the position of the feature area in the determination target image for each rotation angle. For example, as shown in Fig. 5, a determination area (dotted frame) having the same size as the feature area is extracted from the determination target image at a position corresponding to (corresponding to) the position of the feature area.
[0087] In step S203, the feature extraction unit 24 may adjust (match) the resolution of the image related to the extracted determination region to a resolution adjusted according to the size of the feature region related to the registered image. For example, if the size of the feature region is large and the determination features are stored at a resolution (e.g., 100 dpi) lower than a predetermined value (e.g., 300 dpi), the resolution of the image related to the extracted determination region is also changed to 100 dpi. In addition, the feature extraction unit 24 may set the determination region to an expanded region obtained by adding a peripheral region to the region corresponding to the feature region (e.g., an expanded region by 0.5 inches in all directions). Then, the process proceeds to step S204.
[0088] In step S204, feature points related to the determination region are extracted. The feature extraction unit 24 extracts feature points in the determination region extracted in step S203. For example, the feature extraction unit 24 extracts feature points related to the determination region using a method such as SIFT, SURF, or A-KAZE. Then, the process proceeds to step S205.
[0089] In step S205, it is determined whether or not feature points have been extracted in the determination area. In step S204, the upright determination unit 27 determines whether or not feature points have been extracted in the determination area by the feature extraction unit 24. If it is determined that feature points have not been extracted, the upright determination unit 27 does not perform the matching process described below, and the process proceeds to step S214. On the other hand, if it is determined that feature points have been extracted, the process proceeds to step S206.
[0090] In step S206, a matching process is performed between the registered image and the image to be determined based on the extracted feature points and feature amounts. The uprightness determination unit 27 calculates the number of matched feature points between the two images based on the feature points and feature amounts related to the registered image and the image to be determined. The uprightness determination unit 27 also compares the feature amounts related to the feature points in the registered image (feature region) with the feature amounts related to the feature points in the image to be determined (determination region) to calculate (determine) the similarity of the feature amounts between the two images. The uprightness determination unit 27 performs this matching process using a brute force method, a fast approximate neighbor search method, or the like. Thereafter, the process proceeds to step S207.
[0091] In the following steps S207 to S219, the matching degree (similarity degree) between the registered image and the image to be determined is divided into multiple stages (four stages from stage 1 to stage 4) according to the number of matched feature points and the similarity of the feature amounts, thereby determining whether the image to be determined is upright and correcting the orientation of the image to be determined. The division into stages based on the matching degree is not limited to four stages, and other multiple stages may be used.
[0092] In step S207, it is determined whether the number X of matched feature points and the feature similarity Y are equal to or greater than a predetermined threshold X1 and a predetermined threshold Y1, respectively. As a result of the matching process in step S206, the upright determination unit 27 determines whether the number X of matched feature points is equal to or greater than a predetermined threshold X1 and the feature similarity Y is equal to or greater than a predetermined threshold Y1. If X is equal to or greater than the predetermined threshold X1 and Y is equal to or greater than the predetermined threshold Y1 (stage 1), the process proceeds to step S208. On the other hand, if X is not equal to or greater than the predetermined threshold X1 and Y is not equal to or greater than the predetermined threshold Y1, the process proceeds to step S209.
[0093] In step S208, the orientation of the determination target image is determined. The upright orientation determination unit 27 determines the orientation (rotation angle) of the determination target image when it is determined that the number of matched feature points is particularly large and the similarity of the feature amounts is particularly high as the orientation (rotation angle) of the determination target image relative to the registered image. For example, as a result of performing a matching process between the determination target image with a rotation angle of 0 degrees and the registered image, if the number of matched feature points and the similarity of the feature amounts each exceed predetermined thresholds (X1, Y1), the determination target image is determined to be upright (the angle difference with the registered image is 0 degrees). Then, the process proceeds to step S215.
[0094] In step S209, it is determined whether the number X of matched feature points and the feature similarity Y are equal to or greater than a predetermined threshold X2 and a predetermined threshold Y2, respectively. As a result of the matching process in step S206, the upright determination unit 27 determines whether the number X of matched feature points is equal to or greater than a predetermined threshold X2 and the feature similarity Y is equal to or greater than a predetermined threshold Y2. If X is equal to or greater than the predetermined threshold X2 and Y is equal to or greater than the predetermined threshold Y2 (stage 2), the process proceeds to step S210. On the other hand, if X is not equal to or greater than the predetermined threshold X2 and Y is not equal to or greater than the predetermined threshold Y2, the process proceeds to step S211.
[0095] In step S210, the number of votes based on the feature points and / or feature amounts is determined for the image orientation (rotation angle). The upright orientation determination unit 27 determines, for example, the number of matched feature points as the number of votes for the rotation angle in step S202. For example, as shown in Fig. 5, when the number of matched feature points between the determination target image and the registered image with a rotation angle of 0 degrees is six, the upright orientation determination unit 27 determines the number of votes for the rotation angle of 0 degrees to be six. Thereafter, the process proceeds to step S214.
[0096] In step S211, it is determined whether the number X of matched feature points and the feature similarity Y are equal to or greater than a predetermined threshold X3 and a predetermined threshold Y3, respectively. As a result of the matching process in step S206, the upright orientation determination unit 27 determines whether the number X of matched feature points is equal to or greater than a predetermined threshold X3 and the feature similarity Y is equal to or greater than a predetermined threshold Y3. If X is equal to or greater than the predetermined threshold X3 and Y is equal to or greater than the predetermined threshold Y3 (stage 3), the process proceeds to step S212. On the other hand, if X is not equal to or greater than the predetermined threshold X3 and Y is not equal to or greater than the predetermined threshold Y3 (stage 4), the process proceeds to step S214. Note that in stage 4, it is determined that the image to be determined does not have a feature corresponding to the determination feature of the registered image, and the number of votes for the image orientation (rotation angle) in this case is not determined.
[0097] In step S212, it is determined whether the determination is correct (whether it is an erroneous determination or not). The upright determination unit 27 detects an erroneous determination based on the relative positions of the feature points matched by the matching process in step S206 (the feature points in the registered image and the feature points in the image to be determined). The upright determination unit 27 calculates a reliability based on, for example, the variance value of the distances between the matched feature points and the variance value of the slopes of the lines connecting the matched points, and determines whether it is an erroneous determination by comparing the reliability with a predetermined threshold. If it is determined not to be an erroneous determination, the process proceeds to step S213. On the other hand, if it is determined to be an erroneous determination, the process proceeds to step S214. Note that if it is determined to be an erroneous determination, the number of votes is not determined for the orientation (rotation angle) of the image in that case so that the image to be determined is not corrected to an incorrect orientation due to the erroneous determination.
[0098] In step S213, the number of votes based on the feature points and / or feature amounts is determined for the image orientation (rotation angle). Note that the process in step S213 is the same as the process in step S210, so a description thereof will be omitted. Thereafter, the process proceeds to step S214.
[0099] In step S214, it is determined whether the rotation process has been completed for all angles at which the outer edge shape of the image to be determined matches the outer edge shape of the image in the predetermined format in an upright state. The upright orientation determination unit 27 determines whether the rotation process has been completed for all angles at which the outer edge shapes match, and if the rotation process has not been completed for all angles, the process returns to step S202, and after the image to be determined is rotated, the processes of steps S203 to S213 are executed again. On the other hand, if it is determined that the rotation process has been completed for all angles, the process proceeds to step S215.
[0100] In step S215, it is determined whether the orientation of the image to be determined has been determined. The upright orientation determination unit 27 determines whether the orientation of the image to be determined has been determined in step S208. If the orientation has been determined, the process proceeds to step S216. On the other hand, if the orientation has not been determined, the process proceeds to step S217.
[0101] In step S216, the orientation of the determination target image is corrected. For example, if the upright orientation determination unit 27 determines that the determination target image is an image rotated 90 degrees clockwise with respect to the registered image, the orientation correction unit 28 corrects the orientation of the determination target image to an upright orientation by rotating the determination target image 90 degrees counterclockwise. Note that if it is determined in step S208 that the orientation of the determination target image is upright (the angular difference with the registered image is 0 degrees), the correction process in step S216 is not necessary. Thereafter, the process shown in this flowchart ends.
[0102] In step S217, it is determined whether or not there is an orientation for which the number of votes has been determined. The upright orientation determination unit 27 determines whether or not there is an orientation for which the number of votes has been determined in steps S210 and S213. For example, in the examples of FIGS. 5 to 8, the number of votes is determined as 6 for a rotation angle of 0 degrees, 0 for a rotation angle of 90 degrees, and 1 for a rotation angle of 270 degrees, and thus the upright orientation determination unit 27 determines that there is an orientation for which the number of votes has been determined. If there is an orientation for which the number of votes has been determined, the process proceeds to step S218. On the other hand, if there is no orientation for which the number of votes has been determined, the process shown in this flowchart ends. Note that if there is no orientation for which the number of votes has been determined, that is, if the image to be determined does not have features corresponding to the determination features of the registered image and an upright orientation cannot be determined, conventional upright orientation correction processing such as upright orientation correction using OCR may be performed.
[0103] In step S218, the orientation of the image to be determined is determined based on the number of votes. The upright orientation determination unit 27 determines the orientation (rotation angle) with the highest number of votes as the orientation of the image to be determined. For example, in the examples of FIGS. 5 to 8, the number of votes is highest when the rotation angle is 0 degrees, so the upright orientation determination unit 27 determines that the rotation angle of the image to be determined with respect to the registered image is 0 degrees, that is, the image to be determined is upright. Then, the process proceeds to step S219.
[0104] In step S219, the orientation of the determination target image is corrected. Note that the processing of step S219 is the same as the processing of step S216, and therefore description thereof will be omitted. Thereafter, the processing shown in this flowchart ends. Note that if characteristic regions relating to a plurality of images in a predetermined format are stored in the determination information storage unit 21, the upright orientation determination processing shown in FIGS. 15 and 16 is executed sequentially for each characteristic region until the image orientation of the determination target image is determined (until it matches any of the characteristic regions).
[0105] According to the system shown in this embodiment, it is possible to determine the orientation of an image using information related to characteristic areas within an image having a predetermined format, and therefore it is possible to improve the accuracy of determining the orientation of an image even in cases where it is difficult to determine the orientation (the accuracy is reduced) using conventional known recognition methods.
[0106] For example, when using known character recognition methods (such as OCR), it can be difficult to determine the orientation of documents with characteristics that result in low accuracy (such as overlapping characters and background, uppercase letters, bold letters, dotted characters, handwritten characters, etc.) or difficult-to-recognize features (such as few characters, a mixture of vertical and horizontal writing, or only other features such as a person's face). However, the system described in this embodiment can determine the orientation by using information related to characteristic regions within an image with a predetermined format, thereby improving the accuracy of orientation determination even in such cases. For the same reason, it is also possible to accurately determine the orientation of images that do not contain objects for which recognition methods have been established.
[0107] Furthermore, when extremely high accuracy (e.g., 100%) is required, such as in an operational environment where QC (Quality Control) processes are to be reduced, known character recognition methods (such as OCR) often use generic character recognition processes, making it difficult to perform highly accurate character recognition due to the presence of similar characters (e.g., characters that look similar when inverted). However, the system described in this embodiment is capable of performing upright orientation determination using information related to a specific feature area in an image having a specific format, allowing users to easily customize the upright orientation determination process at each scanning site. Therefore, the system described in this embodiment allows for highly accurate orientation determination by customizing the upright orientation determination process, even in cases where extremely high accuracy is required.
[0108] Furthermore, for documents in the same format as a registered image that has been registered in the determination information storage unit 21, the uprightness determination and orientation correction will be performed automatically thereafter, so visual uprightness determination and manual orientation correction during the QC process will remain unnecessary, making it possible to improve the efficiency of scanning operations and operator productivity.
[0109] For example, in the past, when digitizing a large number of documents, such as forms, documents were loaded onto a document loading section of an ADF (Auto Document Feeder) scanner and scanned all at once. However, due to labor and other constraints, it was difficult to perform sufficient preprocessing (such as sorting and correcting the orientation of documents) at all sites. Therefore, documents loaded onto the document loading section may be scanned with misaligned orientations. In this case, document images with misaligned orientations are output. During the QC process, operators must check the image orientation of each document and manually rotate images with misaligned orientations. This frequent occurrence leads to reduced operator productivity. Even in such cases, the system described in this embodiment can automatically determine the orientation of documents using information related to characteristic regions within images with a predetermined format, thereby improving the efficiency of scanning operations and operator productivity.
[0110] [Second embodiment] Next, a second embodiment will be described. In the second embodiment, items that overlap with the contents described in the first embodiment will be assigned the same reference numerals and descriptions thereof will be omitted.
[0111] In the first embodiment described above, feature region candidates are extracted from a single training image having a predetermined format using one or more of the following methods: automatic, manual, and semi-automatic. A user then selects a feature region from among these candidate candidates. However, the number of training images used to determine the feature region is not limited to one, and the setting of the feature region is not limited to manual setting by the user as shown in the first embodiment. In this embodiment, a feature region used to determine whether an image having the predetermined format is upright (determines whether the image is facing the direction desired by the user) is automatically extracted using a plurality of training images having the predetermined format, and the feature region is automatically registered (set).
[0112] In this embodiment, as in the first embodiment, after the characteristic regions are registered, an upright orientation determination process is performed on the determination target image based on the determination information related to the characteristic regions. The flow of the upright orientation determination process in this embodiment is roughly the same as that described in the first embodiment with reference to Figures 15 and 16, so a description thereof will be omitted.
[0113] The configuration of the information processing system 2 according to this embodiment is generally similar to that described in the first embodiment with reference to Fig. 1, and therefore will not be described again. However, the information processing system 2 according to this embodiment may further include a server (not shown) that can be connected to the information processing device 1 via a network.
[0114] 17 is a diagram showing an outline of the functional configuration of the information processing device according to this embodiment. The information processing device 1 functions as an information processing device including a determination information storage unit (determination information database) 21, an image receiving unit 22, a feature extraction unit 24, a rotation unit 26, an upright position determination unit 27, an orientation correction unit 28, a display unit 29, an instruction input receiving unit 30, a common region extraction unit 31, a reliability determination unit 32, and a region determination unit 33, by a program recorded in a storage device 14 being read into a RAM 13 and executed by a CPU 11, which controls each piece of hardware provided in the information processing device 1.
[0115] In this embodiment and other embodiments described later, each function of the information processing device 1 is executed by a CPU 11, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors. Furthermore, each functional unit of the information processing device 1 is not limited to being implemented in a device (one device) consisting of a single housing, but may be implemented remotely and / or distributedly (for example, on the cloud).
[0116] Furthermore, the determination information storage unit (determination information database) 21, image receiving unit 22, feature extraction unit 24, rotation unit 26, upright position determination unit 27, orientation correction unit 28, and display unit 29 according to this embodiment are generally similar to those described in the first embodiment with reference to Fig. 2, and therefore description thereof will be omitted. However, since the image receiving unit 22 and feature extraction unit 24 differ in some respects from the description in the first embodiment, the differences from the first embodiment will be described below.
[0117] The image receiving unit 22 receives input of a plurality of (two or more) learning images having a predetermined format (acquires a plurality of learning images) in order to learn (acquire) information for determining the orientation of an image (determining whether the image is upright). For example, when the image acquisition device 9 scans a plurality of documents (e.g., 2 to 5 documents) relating to the same type of document (format), the image receiving unit 22 acquires the scanned images of the plurality of scanned documents as learning images.
[0118] The feature extraction unit 24 extracts features (feature points and feature amounts) for each acquired training image. The feature extraction unit 24 extracts features that are robust to errors (illumination, scaling, rotation, etc.) caused by devices or scans from the entire region of the training image (scanned image). In this embodiment, feature extraction is performed for training images in an upright state. Feature points and feature amounts can be extracted using known methods, such as feature extraction methods A-KAZE, SIFT, SURF, and ORB (Oriented FAST and Rotated BRIEF). In this embodiment, features are extracted for training images that are grayscale images (including grayscaled color images), but the present invention is not limited to this.
[0119] The instruction input receiving unit 30 receives from the user an instruction to rotate to an upright orientation any training image that is not upright among all training images whose input has been received by the image receiving unit 22. As will be described later, feature extraction processing and feature point matching are performed on training images that are in an upright state, so the user issues an instruction to rotate the training image that is not upright to an upright state.
[0120] The common area extraction unit 31 extracts an area (common area) that is common to a plurality of upright learning images. , one or more candidate feature regions (partial regions) to be used for determining whether the image is upright are extracted (generated). The common region extraction unit 31 performs a matching process between upright training images based on the features (feature points and feature amounts) of each training image, thereby determining (extracting) a region common to multiple training images. More specifically, the common region extraction unit 31 determines a region where feature points matched between training images by the matching process are located (a region including matching points) as a region common to the training images (candidate feature region (candidate partial region)). In this embodiment, candidate feature regions are determined based on a region common to each pair of two training images (hereinafter referred to as a "candidate feature region for each pair").
[0121] The common region extraction unit 31 determines one or more pairs of two training images from the acquired training images and performs feature point matching for each pair of training images. For example, the common region extraction unit 31 determines (generates) the training image pairs by dividing the acquired training images into two sets of images. However, if there is an odd number of training images, for example, any one image may be used twice in two sets so that all training images are used in the matching process. For example, if there are five training images, training image pairs are generated such as a pair of the first and second images, a pair of the third and fourth images, a pair of the fifth and first images, and feature point matching between two images is performed for each pair. Note that the method for generating training image pairs is not limited to the example described in this embodiment, and the combination of training images and the number of pairs generated are arbitrary.
[0122] FIG. 18 is a diagram showing an example of a matching result according to this embodiment. FIG. 18 shows the result of feature point matching performed between a first training image (training image 1) and a second training image (training image 2) (a pair of training image 1 and training image 2). In FIG. 18, training image 1 and training image 2 are arranged side by side in the horizontal direction, and two feature points (a feature point in training image 1 and a feature point in training image 2 (shown as a circle)) connected by a line segment (dashed line) are shown as feature points that have matched with each other through feature point matching. In the example of FIG. 18, as a result of feature point matching performed between training image 1 and training image 2, feature points match at multiple locations.
[0123] Note that various known techniques can be used for matching between two images, such as the k-Nearest Neighbor (KNN) method, a brute force search method, or a fast approximate neighbor search method. In this embodiment, if the coordinate positions of matched feature points differ significantly between the two images, it is determined that matching has not been performed correctly (the feature points are not related to the same shape), and the feature points are excluded from the matched feature points before generating feature region candidates (pair-specific feature region candidates). However, since it is expected that the positions of scanned images (learning images) may differ from one document to another, if the deviation in coordinate positions is within an allowable range (a predetermined range, such as 0.5 inches), the coordinates are treated as being the same.
[0124] Based on the results of matching processing between the learning images, the common region extraction unit 31 generates (extracts) one or more regions that include (surround) feature points common to the learning images as candidate feature regions to be used for upright orientation determination. In this embodiment, the common region extraction unit 31 first detects, for each pair, regions where matched feature points (common feature points) are concentrated by clustering, and classifies the matched feature points into multiple clusters (groups). Then, for each classified cluster, the common region extraction unit 31 generates regions that surround feature points (matched feature points) belonging to the cluster as candidate feature regions for each pair (candidate partial regions for each pair).
[0125] FIG. 19 is a diagram illustrating an example of a method for generating paired feature region candidates according to this embodiment. FIG. 19 illustrates an example in which feature points matched between a pair of training image 1 and training image 2 are classified into multiple clusters by clustering, thereby generating multiple paired feature region candidates for the pair of training image 1 and training image 2. As shown in FIG. 19, feature points matched between two images are classified into multiple clusters, and for each cluster, circumscribing rectangles (bold-line frames in the figure) of all feature points belonging to the cluster are generated as paired feature region candidates. Note that in this embodiment, the paired feature region candidates are regions surrounded by circumscribing rectangles of the feature points belonging to the cluster, but this is not limited thereto; regions having shapes other than rectangles may be used as long as they surround the feature points belonging to the cluster. Furthermore, the paired feature region candidates may be regions obtained by expanding or reducing the circumscribing rectangles of the feature points belonging to the cluster.
[0126] Note that various known methods can be used for clustering the matched feature points, such as non-hierarchical cluster analysis, such as the K-means method (k-point average method). For example, by using the K-means method, feature points within a predetermined region (e.g., within a region of ±1 inch from the center) centered on the feature point (best point) with the highest matching degree (0 to 100%) in feature point matching are divided into two clusters: one including the best point and one not including the best point. Then, the circumscribed rectangle of the feature points in the cluster including the best point is set as a candidate paired feature region, and the feature points in the cluster including the best point are excluded. The best point is again determined from the matched feature points, and clustering is performed. By repeating this process, the matched feature points can be divided into many (multiple) clusters, and multiple candidate paired feature regions can be generated.
[0127] Clustering is performed based on the similarity of the positions (coordinate positions, etc.) of matched feature points. The number of clusters can be set arbitrarily. While FIG. 19 shows an example in which paired feature region candidates for a pair of training image 1 and training image 2 are generated using matched feature points in training image 2, the present invention is not limited to this, and paired feature region candidates may also be generated using matched feature points in training image 1. Furthermore, the present invention is not limited to clustering, and various known methods may be used as long as they are capable of grouping matched feature points based on similarity.
[0128] The common region extraction unit 31 determines feature region candidates based on the paired feature region candidates generated for each pair. Specifically, the common region extraction unit 31 detects paired feature region candidates that have the same (matching) position (coordinates) and size in all pairs of learning images, and determines feature region candidates based on these paired feature region candidates that are common to all pairs. Note that, taking into account differences in device model, individual device differences, errors between scans, etc., the position and size may be determined to be common even when the position and size of paired feature region candidates are roughly the same (e.g., the difference is less than a predetermined value).
[0129] FIG. 20 is a diagram illustrating an example of a method for generating feature region candidates according to this embodiment. FIG. 20 shows pairwise feature region candidates (thick-line frames (including dashed-line frames) in the figure) generated for each pair of the first and second training images and the third and fourth training images when there are four training images. As shown in FIG. 20 , the pairwise feature region candidates associated with the dashed-line frames are not regions that share a common position and size across all pairs (two pairs in the case of FIG. 20 ). Such regions are unlikely to be common regions across images having a predetermined style, and are therefore ineligible as feature regions used for determining upright orientation. Therefore, they are excluded from the candidate feature region pool. On the other hand, as shown in FIG. 20 , the pairwise feature region candidates associated with the thick-line frames indicated by solid lines are regions that share a common position and size across the two pairs, and are therefore estimated to have the potential to be common regions across images having a predetermined style. In this way, the common region extraction unit 31 determines feature region candidates based on pairwise feature region candidates that share a common position and size across all pairs of training images.
[0130] The common region extraction unit 31 determines feature region candidates by detecting overlapping regions (rectangles) between paired feature region candidates determined to match in position and size for all pairs of learning images. As an example, a paired feature region candidate based on the first and second learning images, indicated by a rectangle with coordinates (0,0) to (2,2), and a paired feature region candidate based on the third and fourth learning images, indicated by a rectangle with coordinates (1,1) to (3,3), both of which are determined to match in position and size, are shown. In this case, the overlapping rectangles of the paired feature region candidates are rectangles (1,1) to (2,2), and the common region extraction unit 31 determines (generates) this overlapping rectangle (1,1) to (2,2) as the feature region candidate. Note that if the overlapping region (rectangle) is small, such as smaller than a predetermined size, it may be expanded around the overlapping region, and the expanded region may be determined as the feature region candidate. Furthermore, if the expanded region overlaps with an overlapping region of other paired feature region candidates, the region combined with the overlapping region may be set as a feature region candidate.
[0131] FIG. 21 is a diagram showing an example of feature region candidates according to this embodiment. FIG. 21 illustrates an example in which feature region candidates are generated (determined) as a result of detecting overlaps (overlapping regions) between the paired feature region candidates based on training images 1 and 2 and the paired feature region candidates based on training images 3 and 4 shown in FIG. 20. In FIG. 21, feature region candidates are indicated by bold frames. For example, feature region candidate 83 shown in FIG. 21 is a region determined based on the overlapping region detected between paired feature region candidate 81 based on training images 1 and 2 and paired feature region candidate 82 based on training images 3 and 4 in FIG. 20. For example, the common region extraction unit 31 detects overlaps between the paired feature candidates generated for each pair by superimposing the two images shown in FIG. 20 (image 1 related to paired feature region candidates based on training images 1 and 2 and image 2 related to paired feature region candidates based on training images 3 and 4).
[0132] The feature region candidate 83 may be, for example, the overlapping region itself between the paired feature region candidate 81 and paired feature region candidate 82, or may be a region obtained by expanding or contracting this overlapping region. In this way, when N sets of learning images for performing matching processing (two sets in the above example) are determined, the feature region candidate is determined based on (the overlapping regions of) the N (two in the above example) paired feature region candidates that are determined to match each other in position and size.
[0133] When there are two training images, only one set (pair) of training images is determined, and therefore the feature region candidates are set as the paired feature region candidates generated from these two training images (feature region candidates = paired feature region candidates).
[0134] The reliability determination unit 32 determines, for each of one or more feature region candidates, a common region reliability that indicates the degree to which the feature region candidate is suitable as a feature region used for uprightness determination (degree to which it is suitable for orientation determination). The reliability determination unit 32 includes a commonality determination unit 32A, an incompatibility determination unit 32B, an exclusion degree determination unit 32C, and a reliability calculation unit 32D.
[0135] The commonality determination unit 32A determines a commonality that indicates the likelihood that a characteristic region candidate is a region that has a predetermined style and is shared at approximately the same position among multiple different images. The commonality determination unit 32A determines (calculates) the commonality of the characteristic region candidate based on one or more attributes (evaluation indexes) related to the characteristic region candidate. More specifically, the commonality determination unit 32A calculates scores c1 to cm for each evaluation index (evaluation index 1 to evaluation index m, where m is the number of evaluation indexes) related to the commonality, and calculates the commonality C based on each score. In this embodiment, the calculated scores are multiplied together to determine the commonality (C = c1 × ... × cm).
[0136] Examples of the attributes (evaluation indices) of the feature region candidates include the degree of matching (the number (density) of matching feature points, degree of matching, etc.) between the learning images in the region related to the feature region candidate (the region of the feature region candidate, or the region of the set of feature region candidates that formed the feature region candidate), the size (area) and position of the region related to the feature region candidate, and the possibility of changing the character strings in the region related to the feature region candidate. The commonality determination unit 32A determines the commonality based on at least one attribute (score for the evaluation index) out of these attributes (evaluation indices).
[0137] <Commonality evaluation index 1: Matching degree of matched feature points> The commonality determiner 32A determines the score c1 for evaluation index 1 by using the degree of matching of the matched feature points related to the feature region candidate as evaluation index 1 for commonality. In this embodiment, the score c1 for evaluation index 1 is determined by using the degree of matching of the matched feature points (see the circles in FIG. 19 ) within the region of the paired feature region candidate (hereinafter referred to as the "paired feature region candidate corresponding to the feature region candidate") used (the basis for the feature region candidate) when determining the feature region candidate, as the degree of matching of the matched feature points related to the feature region candidate. The matched feature points within the region of the paired feature region candidate refer to the feature points (the feature points used to determine the circumscribing rectangle) matched by the feature point matching performed by the common region extractor 31 when generating the paired feature region candidate.
[0138] It is estimated that the higher the matching degree in feature point matching between training images is for an area that includes feature points, the more likely it is that the shapes of objects in that area will match between the training images. Therefore, the score c1 of evaluation index 1 is determined so that the higher the matching degree (0 to 100%) for each matched feature point related to the feature region candidate, the higher the commonality of the feature region candidate.
[0139] The commonality determination unit 32A first calculates a score c1(n) (n = 1 to N, N is the number of sets) relating to the degree of matching of feature points located within the group-based characteristic region candidate for each of all group-based characteristic region candidates corresponding to the characteristic region candidate. In the example of Figures 20 and 21, the group-based characteristic region candidates corresponding to characteristic region candidate 83 are group-based characteristic region candidates 81 and 82, and in order to determine score c1 for characteristic region candidate 83, score c1(1) relating to the degree of matching of feature points in group-based characteristic region candidate 81 and score c1(2) relating to the degree of matching of feature points in group-based characteristic region candidate 82 are calculated.
[0140] For example, for each group-specific characteristic region candidate, the score c1(n) relating to the degree of matching of the matched feature points is calculated as follows: c1(n) = sum of the degrees of matching of all feature points located within the group-specific characteristic region candidate divided by the number of feature points located within that region. In this way, c1(n) may be a representative value (for example, an average value) of the degrees of matching of all feature points located within the group-specific characteristic region candidate. The commonality determiner 32A then calculates, as the score c1 for the characteristic region candidate, a representative value (for example, an average value) of the scores (c1(1) to c1(N)) relating to the degree of matching for all group-specific characteristic region candidates corresponding to the characteristic region candidate.
[0141] <Commonality evaluation index 2: Size (area) of the region related to the feature region candidate> The commonality determiner 32A determines the score c2 for evaluation index 2 by using the size (area) of the region related to the candidate characteristic region as the evaluation index 2 related to the commonality. More specifically, the score c2 for evaluation index 2 is determined by using the area of the region of the paired candidate characteristic region corresponding to the candidate characteristic region as the area of the region related to the candidate characteristic region. It is desirable to select a region related to an object that is common to a predetermined style, such as a title or logo, as the feature region used for the upright orientation determination. Titles, logos, etc. are often relatively large objects (having an area equal to or greater than a certain level). For this reason, the score c2 for evaluation index 2 is determined so that the larger the area of the candidate characteristic region, the higher the commonality of the candidate characteristic region. This score determination method also makes it possible to reduce the commonality of candidate characteristic regions related to small areas (such as parts of text) that are likely to coincidentally match between images (possibility of erroneous matching) through feature point matching.
[0142] The commonality determination unit 32A first calculates a score c2(n) (n = 1 to N, N is the number of sets) relating to the area of each of all group-based characteristic region candidates corresponding to the characteristic region candidates. In the example of Figures 20 and 21, to determine the score c2 for the characteristic region candidate 83, a score c2(1) relating to the area of the region of group-based characteristic region candidate 81 and a score c2(2) relating to the area of the region of group-based characteristic region candidate 82 are calculated.
[0143] For example, for each group of characteristic region candidates, the score c2(n) relating to the area of the group of characteristic region candidates is calculated as follows: c2(n) = area of group of characteristic region candidate × offset magnification. For example, when the area of the group of characteristic region candidates is 100 pixels × 100 pixels, the offset magnification is set to 0.0001 so that the score c2(n) = 1 (100%). Then, the commonality determination unit 32A calculates the representative value (for example, the average value) of the scores (c2(1) to c2(N)) for all group of characteristic region candidates corresponding to the characteristic region candidate as the score c2 for the characteristic region candidate.
[0144] <Commonality evaluation index 3: Density of matched feature points> The commonality determiner 32A determines the score c3 for evaluation index 3 by using the density (number) of matched feature points for the feature region candidate as the evaluation index 3 for commonality. In this embodiment, the score c3 for evaluation index 3 is determined by using the density of matched feature points (see circles in FIG. 19 ) within the region of the paired feature region candidate corresponding to the feature region candidate as the density of matched feature points for the feature region candidate. Note that, similar to evaluation index 1, the matched feature points within the region of the paired feature region candidate refer to the feature points matched by the feature point matching performed by the common region extractor 31 when generating the paired feature region candidate.
[0145] It is estimated that the more feature points a region contains that match between the training images, the more likely it is that the shape of the object in that region matches between the training images. Therefore, the score c3 of evaluation index 3 is determined so that the higher the density of matching feature points (the number of feature points per unit area) for the feature region candidate, the higher the commonality of the feature region candidate.
[0146] The commonality determination unit 32A first calculates a score c3(n) (n = 1 to N, N is the number of sets) relating to the density of feature points located within the group-based characteristic region candidate for each of all group-based characteristic region candidates corresponding to the characteristic region candidate. In the example of Figures 20 and 21, to determine the score c3 for the characteristic region candidate 83, the commonality determination unit 32A calculates a score c3(1) relating to the density of feature points in the group-based characteristic region candidate 81 and a score c3(2) relating to the density of feature points in the group-based characteristic region candidate 82.
[0147] For example, for each group-based characteristic region candidate, the score c3(n) relating to the density of feature points is calculated as follows: c3(n) = number of feature points located in the group-based characteristic region candidate region ÷ area of the region × offset magnification. For example, if 10 feature points are contained within an area of 100 pixels × 100 pixels, the offset magnification is set to 1000 so that c3(n) = 1 (100%). The commonality determiner 32A then calculates, as the score c3 for the characteristic region candidate, a representative value (for example, the average value) of the scores (c3(1) to c3(N)) for all group-based characteristic region candidates corresponding to the characteristic region candidate.
[0148] <Commonality evaluation index 4: Location of regions related to feature region candidates> The commonality determiner 32A determines the score c4 for the evaluation index 4 by using the position of the region related to the candidate characteristic region as the evaluation index 4 related to the commonality. More specifically, the score c4 for the evaluation index 4 is determined by using the position in the image of the paired candidate characteristic region corresponding to the candidate characteristic region as the position of the region related to the candidate characteristic region. As described above, it is desirable to select a region related to a title, logo, etc. as the characteristic region used for determining uprightness, and titles, logos, etc. are often located relatively close to the top or bottom of the document. Therefore, the score c4 for the evaluation index 4 is determined so that the closer the position of the region related to the candidate characteristic region is to the top or bottom of the document (learning image (predetermined format)), the higher the commonality of the candidate characteristic region.
[0149] The commonality determination unit 32A first calculates a score c4(n) (n = 1 to N, N is the number of sets) relating to the position of each of all paired characteristic region candidates corresponding to the characteristic region candidate. In the example of Figures 20 and 21, to determine the score c4 for the characteristic region candidate 83, the commonality determination unit 32A calculates a score c4(1) relating to the position (position within the image) of the region of the paired characteristic region candidate 81 and a score c4(2) relating to the position of the region of the paired characteristic region candidate 82.
[0150] For example, for each group of characteristic region candidates, the score c4(n) relating to the position of the group of characteristic region candidates is calculated as follows: c4(n) = abs(coordinate position (height direction) of the center (center of gravity) of the group of characteristic region candidate (region) ÷ document height - 0.5) × offset magnification. The center coordinate position (height direction) of the group of characteristic region candidate is, for example, the coordinate position in the height direction (up and down direction) where the bottom edge of the document (learning image) is set to 0. For example, the offset magnification is set to 2.0 so that the score c4(n) = 1 (100%) when the center of the region of the group of characteristic region candidate is at the top or bottom edge of the document. The commonality determination unit 32A then calculates the score c4 for the characteristic region candidate as a representative value (for example, the average value) of the scores (c4(1) to c4(N)) for all group of characteristic region candidates corresponding to the characteristic region candidate. The central coordinate position (position within the learning image) of the paired feature region candidate may be a position within any learning image among the acquired multiple learning images.
[0151] <Commonality evaluation index 5: Possibility of changing character strings in regions related to feature region candidates> The commonality determination unit 32A determines the score c5 for the evaluation index 5 by using the possibility of changing character strings in a region related to a candidate characteristic region as the commonality evaluation index 5. It is conceivable that there may be a bias in the training images, such as when all training images are images related to forms containing the same date or the same client. In this case, in a region related to character strings whose content, such as dates, may be changed, multiple training images may be determined to locally match, and the region may be extracted as a candidate characteristic region. However, a candidate characteristic region related to character strings whose content (numbers, letters, etc.) may be changed is estimated to be unlikely to be a common region between images having a predetermined format.
[0152] Therefore, the score c5 of the evaluation index 5 is determined so that the more a region related to a candidate characteristic region includes character strings whose content is likely to be changed, the lower the commonality of the candidate characteristic region. Examples of character strings whose content is likely to be changed include dates, names, place names, numbers, etc. In this embodiment, a character string means a combination (sequence) of one or more letters, numbers, or symbols.
[0153] The commonality determination unit 32A extracts (recognizes) character strings within the candidate characteristic region by, for example, performing OCR (Optical Character Recognition) on the candidate characteristic region. The commonality determination unit 32A then determines whether the extracted (recognized) character strings include character strings whose contents are likely to be changed. If a character string with a high changeability is included, the area of the region related to the character string (for example, the area of the circumscribed rectangle of the character string with a high changeability) is calculated. A score c5 related to the changeability of character strings within the candidate characteristic region is then calculated, for example, by c5 = (area of the candidate characteristic region - area of the region related to the character string with a high changeability) ÷ area of the candidate characteristic region.
[0154] As described above, for evaluation indexes 1 to 4, an example has been shown in which scores for each group of characteristic region candidates are first calculated, and then the representative values of these scores are calculated to determine the scores c1 to c4 for the characteristic region candidates, respectively. However, instead of calculating scores c1 to c4 based on the scores for each group of characteristic region candidates, it is also possible to multiply the scores c1(n) to c4(n) for each group of characteristic region candidates, and then calculate the commonality based on the representative value of the multiplication results (c1(n) × ... × c4(n)).
[0155] Furthermore, for evaluation indexes 2 and 4, instead of using the area or position of the region of the group-specific characteristic region candidate, costs c2 and c4 may be calculated using the area or position of the region of the characteristic region candidate itself, as in evaluation index 5. Conversely, for evaluation index 5, instead of using the proportion of character strings that are highly likely to be changed in the region of the characteristic region candidate, cost c5 may be calculated using the proportion of character strings that are highly likely to be changed in the region of the group-specific characteristic region candidate corresponding to the characteristic region candidate, as in evaluation indexes 1 to 4. In this case, as described above, each group-specific characteristic region candidate may be multiplied by scores c1(n) to c5(n), and the commonality may be calculated based on a representative value of the multiplication results (c1(n) × ... × c5(n)).
[0156] Note that when there are two training images, only one set (pair) of training images is determined, and therefore the score for the feature region candidate (e.g., c1) is the same as the score for the pair of feature region candidates generated from the two training images (e.g., c1(1)) (score for feature region candidate = score for pair of feature region candidates).
[0157] Furthermore, for evaluation indexes 1 and 3, instead of using the number of feature points (matched feature points) within the region of each set of feature region candidates corresponding to the determined feature region candidates or the matching degree, it is also possible to use the number of feature points that exist within the region of the determined feature region candidate, among the matched feature points in each set, or the matching degree.
[0158] The incompatibility determination unit 32B determines whether the feature region candidate extracted (generated) by the common region extraction unit 31 corresponds to an incompatibility region, which is a region that is not suitable for determining whether the image is upright (a region that has a negative effect on determining the orientation of the image). Incompatibility regions include regions that may not be detectable, regions where the image orientation may be mistakenly recognized as a different direction, and regions related to objects with rotationally symmetric shapes (regions where the orientation cannot be uniquely identified). In this embodiment, it is determined whether the feature region candidate corresponds to each of the three types of incompatibility regions described above, but the present invention is not limited to this, and it is also possible to determine whether only one or two of the three types of incompatibility regions correspond to each other.
[0159] <Non-conforming area 1 (area that may not be detected)> The incompatibility determination unit 32B determines whether or not the region of the feature region candidate extracted by the common region extraction unit 31 matches a region in all upright learning images that is located at the same (or approximately the same) position as the feature region candidate. More specifically, the incompatibility determination unit 32B determines whether or not it is possible to correctly determine the orientations of all learning images when the feature region candidate is used as the feature region. For example, the upright orientation determination unit 27 performs the upright orientation determination process shown in FIGS. 15 and 16 by determining the region of the feature region candidate in the learning images extracted by the common region extraction unit 31 as the "feature region of the registered image" and all upright learning images as the "images to be determined."
[0160] The incompatibility determination unit 32B determines whether there is at least one learning image among all the learning images that was not correctly determined to be in an upright state in the upright orientation determination process. If there is a learning image whose orientation was not correctly determined, the feature region candidate is presumed to be a region that cannot be detected from the image (document) (a region that may (may) not be detectable), and therefore the incompatibility determination unit 32B determines (decides) the feature region candidate as an incompatibility region.
[0161] In this embodiment, an upright orientation determination process is performed using the feature region candidate and the learning images to determine whether or not the feature region candidate matches a region in each learning image that is in an upright state and that is at a position corresponding to (substantially the same position as) the position of the feature region candidate. However, the method for determining whether or not the two match is not limited to the example given in this embodiment, and may be performed using a method that uses so-called template matching (pattern matching).
[0162] <Incompatible area 2 (area that may cause the image orientation to be mistaken for a different orientation)> The incompatibility determination unit 32B determines whether a region of a feature region candidate extracted by the common region extraction unit 31 matches a region in a training image rotated from an upright state and located at the same (approximately the same) position as the feature region candidate. More specifically, the incompatibility determination unit 32B determines whether, when the feature region candidate is used as a feature region, the orientation of a training image that is not upright is erroneously determined to be upright. For example, the uprightness determination unit 27 determines the feature region candidate region in the training image extracted by the common region extraction unit 31 as the "feature region of the registered image," and all training images that are rotated (for example, rotated 90 degrees, 180 degrees, or 270 degrees from the upright state) as the "images to be determined," and performs the uprightness determination process shown in FIGS. 15 and 16 .
[0163] The incompatibility determination unit 32B determines whether there is at least one or more learning images that have been erroneously determined to be in an upright state in the upright orientation determination process among all the rotated learning images. If there is an erroneously determined learning image, the characteristic region candidate is presumed to be an area where the image to be determined is likely to be erroneously recognized as being in a different orientation, and the incompatibility determination unit 32B determines the characteristic region candidate as an incompatibility region. For example, if there are point-symmetric icons (objects) in the upper left and lower right of a document, the two icons will match when the image is rotated 180 degrees, and therefore the characteristic region candidate related to the icon is determined to be an area where the image orientation will be erroneously recognized as being in a different orientation.
[0164] In this embodiment, an upright orientation determination process is performed using the feature region candidate and the learning images to determine whether or not the feature region candidate matches a region in each of the rotated learning images at a position corresponding to (substantially the same position as) the position of the feature region candidate. However, the method for determining whether or not the two match is not limited to the example given in this embodiment, and template matching (pattern matching) may also be used.
[0165] <Incompatible region 3 (rotationally symmetric region)> The incompatibility determination unit 32B determines whether the feature region candidate extracted by the common region extraction unit 31 is a rotationally symmetric region (a region related to an object (such as a character, symbol, or graphic) having a rotationally symmetric shape). More specifically, the incompatibility determination unit 32B performs pattern matching (template matching) between the feature region candidate in an upright state and the feature region candidate in a rotated state (for example, a state rotated 90 degrees, 180 degrees, or 270 degrees from the upright state). If there is a rotated feature region candidate that matches an upright feature region candidate, that is, if the feature region candidate is a rotationally symmetric region, it is presumed that the feature region candidate is a region that would be erroneously detected as an object of the same shape in a different location, and therefore the incompatibility determination unit 32B determines the feature region candidate as an incompatibility region.
[0166] Areas of rotational symmetry (areas whose orientation cannot be uniquely identified) include, for example, areas of point symmetry (the shape matches when rotated 180 degrees) (letters (I, 8, N), symbols (%, map symbols for power stations, etc.), figures (British flag), etc.), and areas of four-fold rotational symmetry (the shape matches at 0 degrees, 90 degrees, 180 degrees, and 270 degrees) (letters (O, X), symbols (+, map symbols for police stations), figures (Japanese flag), etc.). Note that, strictly speaking, barcodes are not rotationally symmetric objects (areas), but since it is expected that the same features will be obtained even if the position or thickness of the striped lines of a barcode changes (even if they are different), they may be treated as rotationally symmetric.
[0167] The exclusion degree determiner 32C determines an exclusion degree indicating the degree to which a feature region candidate is excluded from the feature region candidates used for upright orientation determination, based on the determination result by the incompatibility determiner 32B. In this embodiment, the exclusion degree determiner 32C determines, as the exclusion degree, an exclusion magnification E, which is an index set so as to lower the common region reliability of a feature region candidate when the feature region candidate is an unsuitable region. The exclusion magnification is a parameter that can be set (determined) between 0 and 1, for example. For example, if a feature region candidate is determined to be an unsuitable region as described above, the exclusion magnification is set to 0 (times), and if the feature region candidate is determined not to be an unsuitable region, the exclusion magnification is set to 1 (times). Note that in this embodiment, if a feature region candidate is determined to be an unsuitable region, the exclusion magnification is set to 0 (i.e., the common region reliability is 0) to exclude the feature region candidate from the feature region candidates. However, the value of the exclusion magnification is not limited to this and may be set to a value other than 0 or 1 (e.g., 0.5) depending on the degree to which the feature region candidate is an unsuitable region. In addition, the exclusion degree is not limited to a parameter (exclusion multiplier) multiplied by the commonality, as long as it can increase or decrease the common area reliability, but may be, for example, a parameter added to or subtracted from the commonality.
[0168] The reliability calculation unit 32D calculates the common region reliability based on the commonality and the determination result by the incompatibility determination unit 32B. For example, the common region reliability is calculated based on the commonality and the exclusion degree. In this embodiment, the reliability calculation unit 32D calculates the common region reliability R by: common region reliability R = commonality C × exclusion magnification E. However, the method of calculating the common region reliability is not limited to the example given in this embodiment, and the common region reliability may be calculated by a method that does not calculate the exclusion degree. For example, if the incompatibility determination unit 32B determines that the feature region candidate corresponds to an incompatibility region, the common region reliability R may be calculated as R = 0. On the other hand, if the feature region candidate is determined not to correspond to an incompatibility region, the common region reliability may be calculated as R = commonality.
[0169] The region determination unit 33 determines one or more feature regions to be used for upright orientation determination from one or more feature region candidates based on the common region reliability. For example, the region determination unit 33 determines the one feature region candidate with the highest common region reliability as the feature region. Alternatively, the region determination unit 33 may determine, as feature regions, multiple feature region candidates selected in descending order of common region reliability. Note that when multiple feature regions are determined, the orientation of the target image may be determined if it is determined in the upright orientation determination process that the registered image and the target image match (are similar) in all of these feature regions.
[0170] 22 is a flowchart showing an outline of the flow of the determination information registration process according to this embodiment. The determination information registration process according to this embodiment is executed when a user presses a new registration button or the like on a registration screen (not shown) in the information processing device 1 and issues an instruction to scan a form for which new determination information is to be registered. Note that in this embodiment, an image related to the form is acquired by scanning the form, but the captured image is not limited to a form, and may be an image related to other documents, a photograph, or the like.
[0171] In step S301, input of a plurality of learning images having a predetermined format is accepted. When the image acquisition device 9 performs a scan process on a plurality of originals relating to a form for which judgment information is to be registered (for which erection correction is to be performed), the image accepting unit 22 acquires images (a plurality of learning images) relating to the scanned plurality of originals from the image acquisition device 9 (accepts the input of learning images). Note that instead of acquiring learning images from the image acquisition device 9, the image accepting unit 22 may acquire learning images stored in advance in the storage device 14.
[0172] In this embodiment, in step S301, all images scanned by the image acquisition device 9 are acquired as learning images, and the processes from step S302 onward are executed, but this is not limited to this. For example, a portion (however, multiple) of all scanned images may be selected by the user as learning images, and the processes from step S302 onward may be executed for these selected portions (learning images). Thereafter, the process proceeds to step S302.
[0173] In step S302, the learning images are rotated to an upright orientation in response to a user instruction. The instruction input receiving unit 30 receives from the user an instruction to rotate to an upright orientation any learning images that are not upright among all the learning images whose input was received in step S301. Then, the rotation unit 26 rotates the learning images related to the rotation instruction to an upright orientation. Note that if all the learning images whose input was received in step S301 are in an upright state, the process of step S302 does not need to be executed. Then, the process proceeds to step S303. In steps S303 to S307, a process of automatically determining a feature region is performed.
[0174] In step S303, features of each training image are extracted. The feature extraction unit 24 extracts features for each of all training images that were placed in an upright position in step S302. In this embodiment, the feature extraction unit 24 extracts feature points and feature amounts from the entire region of each training image. Then, the process proceeds to step S304.
[0175] In step S304, a matching process (feature point matching) is performed between the training images. The common region extraction unit 31 performs the matching process between the training images based on the features (feature points and feature amounts) of each training image extracted in step S303. In this embodiment, the training images are divided into sets (pairs) of two images, and feature point matching is performed for each set of training images. Then, the process proceeds to step S305.
[0176] In step S305, one or more feature region candidates are generated (extracted). The common region extraction unit 31 generates feature region candidates by extracting one or more regions (common regions) common to the learning images based on the results of the matching process in step S304. In this embodiment, the common region extraction unit 31 performs clustering on the feature points matched in the matching process in step S304 for each group, and classifies the matched feature points into one or more clusters. The common region extraction unit 31 then generates, for each classified cluster, a region (group-specific feature region candidate) that surrounds the feature points belonging to the cluster, and generates feature region candidates based on the group-specific feature region candidate generated for each group. Specifically, the common region extraction unit 31 determines feature region candidates based on group-specific feature region candidates that have a common position and size among all groups of learning images. The process then proceeds to step S306.
[0177] In step S306, common region reliability is determined. For each of one or more feature region candidates generated in step S305, the reliability determination unit 32 determines common region reliability R, which indicates the degree to which the feature region candidate is suitable as a feature region to be used in upright orientation determination. In this embodiment, the common region reliability R is determined (calculated) by multiplying the commonality C by the exclusion magnification E. Details of the common region reliability determination process will be described later using Figures 23 and 24. Then, the process proceeds to step S307.
[0178] In step S307, the characteristic region is automatically determined. The region determination unit 33 determines one or more characteristic regions to be used for upright orientation determination from one or more characteristic region candidates based on the common region reliability determined in step S306. The region determination unit 33 may, for example, determine the single characteristic region candidate with the highest common region reliability as the characteristic region, or may determine multiple characteristic region candidates with high common region reliability as the characteristic regions. Then, the process proceeds to step S308.
[0179] In step S308, determination information related to the characteristic region is registered (stored). The feature extraction unit 24 extracts determination features (feature points, feature amounts, etc.) related to the characteristic region determined in step S307 and the position of the characteristic region in the training image when the image is upright. The position of the characteristic region in the training image may be the position of the characteristic region in any of the acquired training images. For example, the feature extraction unit 24 extracts feature points and feature amounts related to the characteristic region using a method such as SIFT, SURF, or A-KAZE. The determination information storage unit 21 then stores the determination features, including the feature points, feature amounts, and image data related to the characteristic region, in association with the position of the characteristic region in the training image when the image is upright. After that, the processing shown in this flowchart ends.
[0180] 23 and 24 are flowcharts showing an outline of the flow of the common region reliability determination process according to this embodiment. The common region reliability determination process according to this embodiment is executed when a feature region candidate is generated in step S305 of the determination information registration process shown in Fig. 22. Note that the common region reliability determination process according to this embodiment (the processes of steps S401 to S416) is executed individually for each feature region candidate.
[0181] In step S401, initial values of the common area reliability, commonality, and exclusion magnification are set. The commonality determination unit 32A sets the initial value C0 of the commonality C to 1.0. The exclusion degree determination unit 32C sets the initial value E0 of the exclusion magnification E to 1.0. In this embodiment, the common area reliability is calculated by common area reliability R = commonality C × exclusion magnification E, and the reliability calculation unit 32D sets the initial value R0 of the common area reliability R to 1.0 by R0 = C0 × E0. Then, the process proceeds to step S402.
[0182] Hereinafter, in steps S402 to S415, whenever the commonality C (commonality-related scores c1 to c5) or the exclusion magnification E is calculated (changed), the reliability calculation unit 32D recalculates (changes) the common area reliability using the changed commonality or exclusion magnification and the calculation formula for the common area reliability (R=C×E). However, this is not limited to this, and the common area reliability may be calculated after the commonality and exclusion magnification are determined in steps S402 to S415.
[0183] In step S402, the commonality (score c1 related to evaluation index 1) is calculated based on the matching degree of the matched feature points related to the feature region candidate. The score c1 of evaluation index 1 is determined so that the higher the matching degree of each matched feature point related to the feature region candidate (the feature points matched in the matching process in step S304 of FIG. 22), the higher the commonality of the feature region candidate. The commonality determiner 32A calculates a score c1(n) related to the matching degree of the matched feature points located within the group feature region candidate for each group feature region candidate corresponding to the feature region candidate. The commonality determiner 32A then calculates the average value of the scores (c1(1) to c1(N)) related to the matching degree for all group feature region candidates corresponding to the feature region candidate as the score c1 for the feature region candidate. Furthermore, the commonality determiner 32A calculates (changes) the commonality by using the commonality C←C0*c1. After that, the process proceeds to step S403.
[0184] In step S403, the commonality (score c2 related to evaluation index 2) is calculated based on the size (area) of the region related to the characteristic region candidate. The score c2 of evaluation index 2 is determined so that the larger the area of the region related to the characteristic region candidate, the higher the commonality of the characteristic region candidate. The commonality determination unit 32A calculates a score c2(n) related to the area for each pair of characteristic region candidates corresponding to the characteristic region candidate. The commonality determination unit 32A then calculates the average value of the scores (c2(1) to c2(N)) for all pair of characteristic region candidates corresponding to the characteristic region candidate as the score c2 for the characteristic region candidate. Furthermore, the commonality determination unit 32A calculates (changes) the commonality by the commonality C←C*c2. After that, the process proceeds to step S404.
[0185] In step S404, the commonality (score c3 related to evaluation index 3) is calculated based on the density of matched feature points related to the feature region candidate. The score c3 of evaluation index 3 is determined so that the higher the density of matched feature points related to the feature region candidate (the number of feature points per unit area), the higher the commonality of the feature region candidate. For each paired feature region candidate corresponding to the feature region candidate, the commonality determiner 32A calculates a score c3(n) related to the density of matched feature points located within the paired feature region candidate. The commonality determiner 32A then calculates the average value of the scores (c3(1) to c3(N)) for all paired feature region candidates corresponding to the feature region candidate as the score c3 for the feature region candidate. Furthermore, the commonality determiner 32A calculates (changes) the commonality by using the commonality C←C*c3. After that, the process proceeds to step S405.
[0186] In step S405, the commonality (score c4 related to evaluation index 4) is calculated based on the position of the region related to the candidate characteristic region. The score c4 of evaluation index 4 is determined so that the closer the position of the region related to the candidate characteristic region is to the top or bottom edge of the document (learning image (predetermined format)), the higher the commonality of the candidate characteristic region. The commonality determination unit 32A calculates a score c4(n) related to the position of each pair of candidate characteristic region corresponding to the candidate characteristic region. The commonality determination unit 32A then calculates the average value of the scores (c4(1) to c4(N)) for all pair of candidate characteristic region corresponding to the candidate characteristic region as the score c4 for the candidate characteristic region. The commonality determination unit 32A then calculates (changes) the commonality by: commonality C←C*c4. After that, the process proceeds to step S406 ( FIG. 24 ). In steps S406, S409, and S412, it is determined whether the candidate characteristic region corresponds to an unqualified region.
[0187] In step S406, it is determined whether the feature region candidate is a region that fails to match with regions of all the learning images in an upright state (a region that may not be detectable). The incompatibility determination unit 32B determines whether or not the feature region candidate matches with all the learning images by determining (checking) whether it is possible to correctly determine the orientation of all the learning images (whether the upright orientation determination process is successful) when the feature region candidate is used as a feature region. If the feature region candidate matches with all the learning images (step S406: YES), the process proceeds to step S407. On the other hand, if the feature region candidate does not match with at least one of all the learning images (step S406: NO), the process proceeds to step S408.
[0188] In step S407, an exclusion magnification is determined. Because it has been determined that the feature region candidate does not fall under the category of an unsuitable region (a region that may not be detectable), the exclusion degree determination unit 32C multiplies the exclusion magnification by 1.0 (exclusion magnification E←E0*1.0). In other words, feature region candidates that do not fall under the category of an unsuitable region are regions suitable as feature regions, and therefore the exclusion magnification is set so as not to reduce the common region reliability. Then, processing proceeds to step S409.
[0189] In step S408, an exclusion magnification is determined. Because the feature region candidate is determined to be an unsuitable region (a region that may not be detectable), the exclusion degree determination unit 32C sets the exclusion magnification to 0 (exclusion magnification E←E0*0). In other words, since the feature region candidate that is an unsuitable region is a region that is not suitable as a feature region, the exclusion magnification is set to lower the common region reliability. Then, the process proceeds to step S409.
[0190] In step S409, it is determined whether the feature region candidate is a region that erroneously matches (mismatches at a different angle) with a region (other region) of the non-upright training image (a region that may cause the image orientation to be mistakenly recognized as a different direction). The mismatch determination unit 32B determines whether the feature region candidate will erroneously determine the orientation of the non-upright training image (fail the upright orientation determination process) when used as a feature region, thereby determining whether the feature region candidate matches with all non-upright training images. If the feature region candidate does not match with all non-upright training images (no mismatch (misjudgment)) (step S409: NO), the process proceeds to step S410. On the other hand, if the feature region candidate matches with at least one image of all non-upright training images (mismatch (misjudgment)) (step S409: YES), the process proceeds to step S411.
[0191] In step S410, an exclusion magnification is determined. Since it has been determined that the feature region candidate does not fall into an incompatible region (a region that may cause the image orientation to be mistakenly recognized as a different direction), the exclusion degree determination unit 32C multiplies the exclusion magnification by 1.0 (exclusion magnification E←E*1.0). Then, the process proceeds to step S412.
[0192] In step S411, an exclusion magnification is determined. The exclusion degree determination unit 32C sets the exclusion magnification to 0 (exclusion magnification E←E*0) because it has determined that the feature region candidate corresponds to an incompatible region (a region that may cause the image orientation to be mistakenly recognized as a different direction). Thereafter, the process proceeds to step S412.
[0193] In step S412, it is determined whether the feature region candidate is a rotationally symmetric region (non-matching region). The non-matching determination unit 32B determines whether the feature region candidate is a rotationally symmetric region by performing pattern matching between the feature region candidate in an upright state and the feature region candidate in a rotated state (for example, a state rotated 90 degrees, 180 degrees, or 270 degrees from the upright state). If the feature region candidate is not rotationally symmetric (does not match through pattern matching) (step S412: NO), the process proceeds to step S413. On the other hand, if the feature region candidate is rotationally symmetric (matches through pattern matching) (step S412: YES), the process proceeds to step S414.
[0194] In step S413, an exclusion magnification is determined. Because it has been determined that the feature region candidate does not correspond to an incompatible region (a rotationally symmetric region), the exclusion degree determination unit 32C multiplies the exclusion magnification by 1.0 (exclusion magnification E←E*1.0). Then, the process proceeds to step S415.
[0195] In step S414, an exclusion magnification is determined. Since the feature region candidate is determined to be an incompatible region (a rotationally symmetric region), the exclusion degree determination unit 32C sets the exclusion magnification to 0 (exclusion magnification E←E*0). Then, the process proceeds to step S415.
[0196] In step S415, the commonality (score c5 related to evaluation index 5) is calculated based on the possibility of changing the character strings in the feature region candidate. The score c5 of evaluation index 5 is determined so that the more character strings the feature region candidate contains that are more likely to be changed, the lower the commonality of the feature region candidate. The commonality determination unit 32A performs OCR on the feature region candidate to extract character strings within the region and determines whether the extracted (recognized) character strings include character strings that are likely to be changed. The commonality determination unit 32A then calculates the score c5 related to the possibility of changing the character strings in the feature region candidate by using the formula c5 = (area of feature region candidate - area of region containing character strings that are likely to be changed) ÷ area of feature region candidate. The commonality determination unit 32A then calculates (changes) the commonality by using the formula C←C*c5. Processing then proceeds to step S416.
[0197] In step S416, the common region reliability is determined. The reliability determination unit 32 determines the common region reliability calculated based on the commonality C and exclusion magnification E determined in steps S401 to S415 as the common region reliability for the characteristic region candidate. Thereafter, the processing shown in this flowchart ends. Then, the region determination unit 33 determines a characteristic region based on the common region reliability for each characteristic region candidate determined by the processing shown in this flowchart.
[0198] The processes of steps S401 to S405 (calculation processes of c1(n) to c4(n) and c1 to c4) may be performed for each set at the timing when the set-wise feature region candidates are generated in step S305 of FIG. 22. Steps S402 to S405 may be performed in any order. Similarly, steps S406 to S408, steps S409 to S411, and steps S412 to S414 may be performed in any order. Furthermore, the process of step S415 may be performed between steps S405 and S406.
[0199] Also, in this embodiment, an example has been described in which the processing proceeds to the next step even if the exclusion magnification is calculated to be 0 in steps S408, S411, and S414 of Fig. 24. However, the present invention is not limited to the branching shown in Fig. 24, and if the exclusion magnification is calculated to be 0, the processing shown in this flowchart may be terminated and the common area reliability may be determined to be 0. Also, in this embodiment, as shown in Fig. 23, an example has been shown in which the commonality calculation processing is performed first in the reliability calculation processing, but the present invention is not limited to this, and a determination as to whether or not the characteristic area corresponds to an incompatible area may be performed first.
[0200] According to the information processing system of this embodiment, feature region candidates are extracted based on multiple training images, and feature regions are determined based on the common region reliability of each feature region candidate. This makes it possible to automatically select (set) feature regions to be used when determining the orientation (uprightness) of an image. This reduces the user's effort of manually setting feature regions and improves convenience. Furthermore, because feature regions are automatically determined based on the common region reliability (automatically learning regions common to multiple training images), it becomes possible to extract regions suitable for orientation determination (orientation correction) that can more accurately (reliably) determine the orientation of an image compared to when a user manually selects feature regions.
[0201] For example, by automatically selecting a feature region based on the common region reliability, it is possible to prevent a region in which the orientation of an image cannot be determined from being selected as a feature region to be used for determining the upright orientation. Examples of regions in which the orientation of an image cannot be determined include (1) a region where the content changes (differs) depending on the form (original) even between forms of the same format, (2) a region in which the feature amount changes depending on the form (original) due to the inclusion of a background pattern (partial or complete) such as "Copy" or "Confidential" and therefore the originals of the same format are not determined to be of the same format, and (3) a region in which a shape that is point-symmetric with a shape (a shape that is rotationally symmetric) at a predetermined position in a form of a different format is located at a point-symmetric position with the predetermined position, resulting in an erroneous determination of the orientation of the form of a different format.
[0202] <Modification> Modifications of this embodiment will be described below. The functional configurations of the information processing system 2 (information processing device 1, server 3) according to modifications 1 to 8 are partially different from those described in this embodiment with reference to Fig. 17, and will be described below with reference to Figs. 25 and 26. Note that items (functional units) that overlap with those described in the first embodiment and this embodiment will be assigned the same reference numerals and descriptions thereof will be omitted.
[0203] 25 is a diagram showing an outline of the functional configuration of an information processing device according to Modifications 1 to 7 of this embodiment. A program recorded in storage device 14 is read into RAM 13 and executed by CPU 11, which controls each piece of hardware provided in information processing device 1, so that information processing device 1 functions as an information processing device including a determination information storage unit (determination information database) 21, an image receiving unit 22, a feature extraction unit 24, a rotation unit 26, an upright orientation determination unit 27, an orientation correction unit 28, a display unit 29, an instruction input receiving unit 30, a common region extraction unit 31, a reliability determination unit 32, and a region determination unit 33, as well as a matching unit 34, a learning image orientation correction unit 35, an orientation determination unit 36, a dimensional compression unit 37, a style determination unit 38, and an update unit 39.
[0204] <Variation 1> In this embodiment, since feature extraction processing and matching processing are performed between upright training images, the orientation of training images that are not upright is corrected in advance to an upright orientation by receiving a rotation instruction from the user for each image. However, the method for correcting training images to an upright orientation is not limited to this. For example, based on one (or more) upright training images among multiple training images, the other training images may be automatically corrected to an upright orientation.
[0205] In the first modification, an example is shown in which, based on one (or more) upright training images, other training images are automatically corrected to an upright orientation. The matching unit 34 performs pattern matching (template matching) between one (or more) upright training images among the multiple training images and other training images other than the upright training image, thereby determining the orientation (uprightness or non-uprightness) of the other training images. Specifically, pattern matching is performed between the upright training image and other training images rotated at an angle whose outer edge shape matches the outer edge shape of the upright training image in its upright state, and the orientation is determined based on the rotation angle at which the best match is achieved. Then, the training image orientation correcting unit 35 corrects training images determined by the matching unit 34 to be not upright to an upright orientation.
[0206] For example, when a user instructs to rotate a non-erect learning image to an upright orientation, rotation unit 26 corrects the non-erect learning image to an upright orientation. Then, matching unit 34 performs pattern matching between the learning image corrected to an upright orientation by rotation unit 26 and other learning images.
[0207] Furthermore, for example, rotation unit 26 may correct a learning image that is not upright to an upright orientation based on the determination result of orientation determination unit 36, which automatically determines the orientation of the learning image by performing OCR processing on the learning image and determining whether the characters are upright. That is, rotation unit 26 corrects a learning image that is determined by orientation determination unit 36 to be not upright (in a different orientation) to an upright orientation. Then, matching unit 34 performs pattern matching between the learning image that has been corrected to an upright orientation by rotation unit 26 and the other learning images.
[0208] Alternatively, for example, by accepting a user's designation of a training image that is upright among a plurality of training images, pattern matching may be performed between the designated training image and the other training images. Then, the training image orientation corrector 35 corrects the other training images to the orientation determined by pattern matching.
[0209] As a result, all learning images are automatically corrected to an upright orientation, which reduces the effort required to align all scanned documents to the same orientation (upright orientation).
[0210] <Variation 2> In this embodiment, features (feature points and feature amounts) for each learning image are extracted by the feature extraction unit 24. In this modified example, an example is shown in which the dimensions of the extracted feature amounts are compressed in order to speed up the process of determining the feature region.
[0211] The dimension compression unit 37 compresses the dimensions of the features for each learning image extracted by the feature extraction unit 24. Various known methods may be used to compress the dimensions of the features, and for example, PCA (Principal Component Analysis) is used to compress the dimensions of the features (data). In this modification, the common region candidate extraction unit 31 performs matching processing using the features whose dimensions have been compressed by the dimension compression unit 37. This makes it possible to speed up the processing up to determining the feature regions.
[0212] <Variation 3> In this embodiment, barcodes are treated as rotationally symmetric regions (non-conforming regions) and are excluded from the target of feature regions (the exclusion magnification is set to 0). However, in this modified example, when barcodes of the same type and approximately the same size are present in similar positions in multiple learning images having a predetermined format, the barcodes are determined to be feature regions.
[0213] In this modification, the region determination unit 33 determines whether a one-dimensional code (barcode) or a two-dimensional code of the same type and approximately the same size is present in a similar position among a plurality of training images having a predetermined format. If a code with a common position, type, and size is present, the region determination unit 33 determines the region associated with the code as a characteristic region. If the region determination unit 33 determines that a code with a common position, type, and size is present among a plurality of training images, the region determination unit 33 may determine the region associated with the code as a characteristic region regardless of the value of the common region reliability for the region associated with the code. Furthermore, if the region determination unit 33 determines that a common code exists, the region may be determined as a characteristic region without calculating the common region reliability. For one-dimensional codes (barcodes), examples of the code type include JAN codes, EAN codes, UPC codes, etc.
[0214] <Variation 4> In this modified example, as in modified example 3, if there is a seal impression of the same type and approximately the same size in a similar position in multiple learning images having a predetermined format, the seal impression is determined to be a feature area.
[0215] In this modification, the region determination unit 33 determines whether there are seal impressions of the same type and approximately the same size at similar positions among multiple learning images having a predetermined format. If there are seal impressions (stamps) with a common position, type, and size among multiple learning images, the region determination unit 33 determines the region related to the seal impression as a characteristic region. If the region determination unit 33 determines that there are seal impressions with a common position, type, and size among multiple learning images, it may determine the region related to the seal impression as a characteristic region regardless of the value of the common region reliability for the region related to the seal impression. Furthermore, if it determines that there is a common seal impression, it may determine the region as a characteristic region without calculating the common region reliability. Examples of the type of seal impression include a seal impression of a company (specific company), a seal impression of a specific individual's registered seal, etc.
[0216] <Variation 5> In this embodiment, features are extracted from training images that are grayscale images. However, the training images used for feature extraction are not limited to grayscale images and may be color images. In this modified example, feature extraction is performed on training images that include color images for each component constituting a color space (such as an RGB (red, green, blue) color space or a CMYK (cyan, magenta, yellow, black) color space). In grayscaled images (e.g., when a color document is scanned in black and white), the weights applied to each color component during grayscaling are different, so the feature amounts vary depending on the color of the original document. Therefore, even if documents (images) of the same type but with different colors (e.g., a blue document and a red document) are similar, it may be impossible to extract common areas (features may not match) between the grayscaled images.
[0217] In this modification, the feature extraction unit 25 extracts feature points and feature quantities for each component of the color space (e.g., color components (red, green, blue)) for each of multiple training images (including color images). Furthermore, the common area extraction unit 31 performs matching processing between multiple training images for the same component and different components based on the feature points and feature quantities extracted for each component of the color space. This makes it possible to accurately detect (extract) common areas displayed in different colors between training images. For example, even in the case of documents of the same type but with different colors, such as the first and second sheets of a shipping slip (two slips with the same format but different printing colors (customer copy (printed in red) and store copy (printed in blue))), it is possible to appropriately extract common areas between these documents.
[0218] <Variation 6> As described above with reference to FIG. 22 , the determination information registration process (characteristic region determination process) according to this embodiment is triggered by the user pressing a new registration button or the like on the registration screen to acquire an image as a learning image. However, the registration process of the characteristic region (determination information) is not limited to being triggered by a manual operation by the user to intentionally start the characteristic region extraction process, and the registration process may be performed automatically without the user's knowledge. This modification shows an example in which the characteristic region (determination information) is automatically registered without the user's knowledge, that is, without the need for a manual operation by the user to start the registration process.
[0219] In this modified example, for example, when a user rotates a scanned image in a desired direction, the scanned image (or image features) is stored in the storage device 14. Then, when the user rotates another scanned image in a desired direction, it is determined whether the scanned image and the saved scanned image are images of the same type of form (same format). If these images are of the same format, the determination information registration process ( FIG. 22 ) of this embodiment is performed on these images, and if a feature region is found, it is automatically registered. When a user rotates a scanned image, it is generally assumed that the user intends to rotate the image in an upright orientation. Therefore, as described above, if there are multiple images rotated by the user, these images are used as learning images, and the process of automatically extracting and registering feature regions is initiated.
[0220] In this modification, the instruction input receiving unit 30 receives an instruction to rotate an image from a user, and the rotation unit 26 rotates the image according to the rotation instruction in the instructed direction. If there are multiple images rotated by the rotation unit 26, the style determination unit 38 determines whether the images have the same style (whether they are images related to the same type of document). The image receiving unit 22 then receives input of the multiple images determined to have the same style as multiple learning images having a predetermined style. The style determination unit 38 determines whether the images have the same style by, for example, pattern matching (template matching). This allows the registration process of the characteristic regions to be started automatically, reducing the manual operation required by the user to register the characteristic regions (determination information), and improving convenience.
[0221] <Variation 7> In this embodiment, an example of automatically extracting and registering a characteristic region is shown, but in this modified example, an example of updating an already registered characteristic region is shown. For example, by scanning an additional document of the same type (document having the same format) as the document (predetermined format) related to the already registered characteristic region, the determination information registration process is executed again to update the characteristic region. Note that instead of scanning an additional document, it is also possible to obtain a learning image of the same format that has not yet been used to register the characteristic region.
[0222] In this modification, an updating unit 39 that updates characteristic regions (determination information) stored in a storage device (determination information database) updates already registered characteristic regions. Specifically, first, the image accepting unit 22 accepts input of other learning images having the predetermined format other than the plurality of learning images having the predetermined format used when determining the characteristic regions. Then, the region determining unit 33 determines one or more characteristic regions to be used for upright orientation determination again based on the plurality of learning images and the other learning images having the predetermined format. Then, the updating unit 39 updates the characteristic regions (determination information) stored in the determination information storage unit 21 based on the re-determined characteristic regions.
[0223] This makes it possible to correct (change) the feature regions, and register regions that are more suitable for determining upright orientation, even in cases where there are only two learning images at first and an incorrect feature region common to only those two learning images is registered. In this way, by updating the feature regions as needed, it becomes possible to improve the accuracy of determining image orientation (upright orientation determination).
[0224] <Variation 8> 26 is a diagram showing an outline of the functional configuration of an information processing system (information processing device, server) according to Modification 8 of this embodiment. The information processing system 2 includes an information processing device 1 and a server 3. A program recorded in a storage device 14 is read into a RAM 13 and executed by a CPU 11, which controls each piece of hardware included in the information processing device 1, so that the information processing device 1 functions as an information processing device including a determination information storage unit (determination information database) 21A, an image receiving unit 22A, a feature extraction unit 24A, a rotation unit 26A, an upright position determination unit 27A, an orientation correction unit 28, a display unit 29, and an instruction input receiving unit 30.
[0225] The server 3 functions as an information processing device including a determination information storage unit (determination information database) 21B, an image receiving unit 22B, a feature extraction unit 24B, a rotation unit 26B, an upright determination unit 27B, a common area extraction unit 31, a reliability determination unit 32, an area determination unit 33, a style determination unit 38, and an area information notification unit 40. Items (functional units) that overlap with those described in the first embodiment and this embodiment are denoted by the same reference numerals and will not be described again. Furthermore, since the determination information storage unit (determination information database) 21A, image receiving unit 22A, feature extraction unit 24A, rotation unit 26A, upright determination unit 27A, determination information storage unit (determination information database) 21B, image receiving unit 22B, feature extraction unit 24B, rotation unit 26B, and upright determination unit 27B have the same functions as the determination information storage unit (determination information database) 21, image receiving unit 22, feature extraction unit 24, rotation unit 26, and upright determination unit 27 described in the first embodiment and this embodiment, respectively, their descriptions will be omitted.
[0226] In this embodiment, the characteristic area (information for determination) is determined and registered in the information processing device 1. However, the characteristic area (information for determination) may also be determined and registered in a server 3 that can be connected to the information processing device 1 via a network. In this modified example, an example is shown in which the characteristic area (information for determination) is determined and registered in the server 3.
[0227] In this modification, the server 3 includes functional units related to the determination information registration process, and performs a characteristic region determination process (determination information registration process) based on training images acquired from the information processing device 1. Furthermore, the server 3 includes a region information notification unit 40 that notifies the information processing device 1 of information (such as determination features) related to the characteristic region determined by the region determination unit 33 in the characteristic region determination process. The server 3 may determine characteristic regions using training images acquired from multiple information processing devices 1. In this case, the server 3 notifies the multiple information processing devices 1 of information related to the determined characteristic regions. The notified information processing devices 1 then reflect (store) the notified information related to the characteristic regions in the determination information database 21A. The server 3 (style determination unit 38) determines whether the training images acquired from the information processing device 1 have the same style, and determines characteristic regions based on two or more training images determined to have the same style by the style determination unit 38.
[0228] This makes it possible to collectively manage the feature areas (information for determination) for various forms (various predetermined formats) on the server 3. Furthermore, it becomes possible to use the same feature areas (information for determination) between information processing devices, thereby improving convenience.
[0229] Furthermore, similar to the sixth modification, the server 3 may also automatically register characteristic regions (determination information) without the user's knowledge. For example, when a user rotates a scanned image in any direction on the information processing device 1, the scanned image (or image features) is transmitted (transferred) to the server 3. The server 3 then stores the scanned image as a rotation-failed image or the like in the server 3. The server 3 (format determination unit 38) determines whether the images stored in the server 3, including scanned images from other users (other information processing devices 1, etc.), have the same format. If these images have the same format, the server 3 performs the determination information registration process ( FIG. 22 ) of this embodiment on these images. If a characteristic region is found, the server 3 automatically registers the characteristic region. The region information notification unit 40 then notifies (transfers) information related to the characteristic region to each user (information processing device 1), and the notified information related to the characteristic region is reflected (stored) in the determination information database 21A in the notified information processing device 1. <Other variations>
[0230] The feature extraction unit 24 may enlarge or reduce the learning image (scanned image) to extract a number of feature points appropriate for the accuracy and performance of the matching process by the common area extraction unit 31. For example, if the number of extracted feature points is too large, the number of extracted feature points may be reduced by reducing the image.
[0231] In this embodiment, the feature extraction unit 24 extracts features from the entire region of each acquired training image, but is not limited to this. For example, for the first and second training images, features are extracted from the entire region of the image to detect common regions, and for the third training image, feature extraction is performed only near the common regions detected in the first and second training images. Similarly, for the fourth training image, feature extraction is performed only near the common regions detected in the first, second, and third training images. This makes it possible to reduce the time required for feature extraction and matching processes when there are a large number of training images.
[0232] Similarly, as a method for narrowing down (reducing) the target areas for feature extraction, for example, for the first and second learning images, features are extracted from the entire image area to detect common areas and calculate the common area reliability for the common areas. Then, for the third and subsequent learning images, feature extraction may be performed only near areas with high common area reliability among the common areas detected in the first and second images. This, as with the above, makes it possible to reduce the time required for feature extraction processing and matching processing.
[0233] In this embodiment, the region determination unit 33 determines one or more feature regions to be used for upright orientation determination from one or more feature region candidates based on the common region reliability. However, thresholds for determining (judging) a region as a feature region, such as the common region reliability and the number of common regions (feature regions), may be stored. In this case, the threshold for determining a region as a feature region may be changeable by the user. The display unit 29 may also display a preview of the results of changing the threshold.
[0234] Furthermore, as described above, when registering the characteristic region (information for determination), the user may select a pre-scanned image as the learning image, rather than acquiring the learning image by scanning the document. However, since the overall image processing during scanning affects and changes the feature amount of the image, a process may be executed to convert the group of scanned images selected by the user (all selected images) so that they have a common image quality (resolution, color, etc.).
[0235] Furthermore, in this embodiment, after the features of the learning image are extracted, a matching process, a feature region candidate extraction process, a common region reliability determination process, etc. are executed, and then the feature region is determined and the information for determination is registered. However, while these processes are in progress, the remaining time of the processes may be displayed to the user. For example, the remaining time of the processes may be predicted from the number of feature points extracted by the feature extraction process and displayed to the user. Similarly, during the automatic registration process of the feature region (information for determination), the progress of the process may be displayed to the user using animation or the like.
[0236] In addition, in the present embodiment, an example is shown in which characteristic regions (determination information) for an image (one type of document) having a predetermined format are registered, but the present invention is not limited to this. For example, documents of multiple formats (multiple types of documents) may be scanned together, and the user may sort (classify) the scanned images by document type (format), so that automatic registration processing of characteristic regions for multiple types of documents may be performed simultaneously in parallel.
[0237] [Third embodiment] Next, a third embodiment will be described. In the third embodiment, items that overlap with the contents described in the first and second embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted.
[0238] In this embodiment, an embodiment will be described that combines the information processing system described in the first embodiment with the information processing system described in the second embodiment. Specifically, an embodiment will be described in which the method for determining a characteristic region is changed depending on the number of acquired learning images (scanned documents). More specifically, when multiple learning images are acquired, the process of automatically determining a characteristic region described in the second embodiment is executed. On the other hand, when only one learning image is acquired, the process of determining a characteristic region is executed by the user selecting a characteristic region candidate extracted by one of the "automatic," "semi-automatic," or "manual" methods described in the first embodiment.
[0239] The configuration of the information processing system 2 according to this embodiment is roughly the same as that described in the first embodiment with reference to Fig. 1, and therefore will not be described again. However, the information processing system 2 according to this embodiment may further include a server (not shown) that can be connected to the information processing device 1 via a network. The functional configuration of the information processing device 1 according to this embodiment will be described below.
[0240] FIG. 27 is a diagram illustrating an outline of the functional configuration of an information processing device according to this embodiment. A program stored in a storage device 14 is loaded into RAM 13 and executed by a CPU 11, which controls the various hardware components of the information processing device 1. The information processing device 1 functions as an information processing device including a determination information storage unit (determination information database) 21, an image receiving unit 22, a designation receiving unit 23, a feature extraction unit 24, a candidate extraction unit 25, a rotation unit 26, an upright orientation determination unit 27, an orientation correction unit 28, a display unit 29, an instruction input receiving unit 30, a common region extraction unit 31, a reliability determination unit 32, and a region determination unit 33. The reliability determination unit 32 includes a commonality determination unit 32A, an incompatibility determination unit 32B, an exclusion degree determination unit 32C, and a reliability calculation unit 32D. The functional units of the information processing device 1 according to this embodiment are generally similar to those described in the first embodiment with reference to FIG. 2 and the second embodiment with reference to FIG. 17, and therefore, description thereof will be omitted.
[0241] 28 is a flowchart showing an outline of the flow of the determination information registration process according to this embodiment. The determination information registration process according to this embodiment is executed when the user presses a new registration button or the like on a registration screen (not shown) in the information processing device 1 and issues a scan instruction for a form for which new determination information is to be registered.
[0242] In step S501, input of one or more learning images having a predetermined format is accepted. When the image acquisition device 9 performs a scan process on one or more documents related to the form for which judgment information is to be registered (for which erection correction is to be performed), the image accepting unit 22 acquires images related to the scanned one or more documents from the image acquisition device 9. Note that instead of acquiring learning images from the image acquisition device 9, the image accepting unit 22 may acquire learning images stored in advance in the storage device 14. Thereafter, the process proceeds to step S502.
[0243] In step S502, the learning image is rotated to an upright position in response to a user instruction. If the learning image whose input was accepted in step S501 is not upright, the instruction input accepting unit 30 accepts an instruction from the user to rotate the learning image to an upright position. Then, the rotation unit 26 rotates the learning image related to the rotation instruction to an upright position. Note that if the learning image whose input was accepted in step S501 is upright, the process of step S502 does not need to be executed. Thereafter, the process proceeds to step S503.
[0244] In step S503, the number of learning images that have been received as input is determined. CPU 11 determines whether the number of learning images acquired in step S501 is one or more. If the number of learning images is more than one (step S503: YES), the process proceeds to step S504. On the other hand, if the number of learning images is one (step S503: NO), the process proceeds to step S505.
[0245] In step S504, a characteristic region determination process is performed. Note that the process in step S504 is the same as the process in steps S303 to S307 in Fig. 22, and therefore a detailed description thereof will be omitted. Thereafter, the process proceeds to step S511.
[0246] In step S505, a designation of automatic, manual, or semi-automatic is accepted as the extraction method for feature region candidates. The designation accepting unit 23 accepts the designation of the extraction method for feature region candidates when the user selects one of "automatic," "manual," or "semi-automatic" on a registration screen, for example. If "manual" is selected, the process proceeds to step S506. If "semi-automatic" is selected, the process proceeds to step S507. If "automatic" is selected, the process proceeds to step S509.
[0247] In step S506, a range designation for a characteristic region by the user is accepted. Note that the process in step S506 is similar to the process in step S106 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the process proceeds to step S510.
[0248] In step S507, the user's designation of the proposal target area is accepted. Note that the process in step S507 is similar to the process in step S107 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the process proceeds to step S508.
[0249] In step S508, feature region candidates are extracted from the proposal target region. Note that the process in step S508 is similar to the process in step S108 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the process proceeds to step S510.
[0250] In step S509, candidates for characteristic regions are extracted. Note that the process in step S509 is the same as the process in step S103 in Fig. 12, so a detailed description will be omitted. Thereafter, the process proceeds to step S510.
[0251] In step S510, a feature region to be used for orientation correction (erecting whether the image is upright) is determined from the region extracted automatically, manually, or semi-automatically. Note that the process in step S510 is similar to the process in step S109 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the process proceeds to step S511.
[0252] In step S511, determination information related to the determined (confirmed) characteristic region is registered (stored). Note that the process in step S511 is similar to the process in step S110 in Fig. 12 and step S308 in Fig. 22, and therefore a detailed description thereof will be omitted. Thereafter, the process shown in this flowchart ends.
[0253] In this embodiment, after the characteristic regions are registered, an upright orientation determination process is performed on the determination target image based on the determination information related to the characteristic regions, as in the first and second embodiments. The flow of the upright orientation determination process in this embodiment is roughly the same as that described in the first embodiment with reference to Figures 15 and 16, and therefore a description thereof will be omitted.
[0254] [Fourth embodiment] Next, a fourth embodiment will be described. In the fourth embodiment, items that overlap with the contents described in the first, second, and third embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted.
[0255] In this embodiment, an implementation that combines the information processing system described in the first embodiment with the information processing system described in the second embodiment will be described. Specifically, an implementation will be described in which the method for determining a feature region is changed depending on the method for extracting feature region candidates selected by the user. More specifically, if the user selects automatic extraction (determination) of feature regions, the process for automatically determining feature regions described in the second embodiment is executed. On the other hand, if the user selects "semi-automatic" or "manual" extraction of feature regions, the process for determining feature regions is executed by the user selecting feature region candidates extracted by the "semi-automatic" or "manual" method described in the first embodiment.
[0256] The configuration of the information processing system 2 according to this embodiment is generally similar to that described in the first embodiment with reference to Fig. 1, and therefore a description thereof will be omitted. However, the information processing system 2 according to this embodiment may further include a server (not shown) that can be connected to the information processing device 1 via a network. Furthermore, the functional configuration of the information processing device 1 according to this embodiment is generally similar to that described in the third embodiment with reference to Fig. 27, and therefore a description thereof will be omitted.
[0257] 29 is a flowchart showing an outline of the flow of the determination information registration process according to this embodiment. The determination information registration process according to this embodiment is executed in the information processing device 1 when the user selects a method for extracting feature region candidates on a registration screen (not shown), for example.
[0258] In step S601, a designation of automatic determination, manual, or semi-automatic is accepted as the extraction (determination) method for feature region candidates. The designation accepting unit 23 accepts a designation of the extraction method for feature region candidates, for example, when the user selects one of "automatic," "manual," or "semi-automatic" on a registration screen or the like. If the designation of "automatic (automatic determination)" is accepted, the process proceeds to step S602. If the designation of "manual" is accepted, the process proceeds to step S605. If the designation of "semi-automatic" is accepted, the process proceeds to step S608.
[0259] In steps S602 to S604, input of a plurality of learning images having a predetermined format is accepted, and the images that are not upright are rotated to an upright orientation, and then a feature region determination process is performed. Note that the processes in steps S602 to S604 are the same as the processes in steps S301 to S307 in Figure 22, and therefore detailed description will be omitted. Thereafter, the process proceeds to step S613.
[0260] In steps S605 and S606, input of one learning image having a predetermined format is accepted, and the image that is not upright is rotated so that it is upright. Note that the processing in steps S605 and S606 is the same as the processing in steps S101 and S102 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the processing proceeds to step S607.
[0261] In step S607, a range designation for the characteristic region by the user is accepted. Note that the process in step S607 is similar to the process in step S106 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the process proceeds to step S612.
[0262] In steps S608 and S609, input of one learning image having a predetermined format is accepted, and the image that is not upright is rotated so that it is upright. Note that the processing in steps S608 and S609 is the same as the processing in steps S101 and S102 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the processing proceeds to step S610.
[0263] In steps S610 and S611, a proposal target region specified by the user is accepted, and candidate feature regions are extracted from the proposal target region. Note that the processing in steps S610 and S611 is similar to the processing in steps S107 and S108 in Fig. 12, and therefore detailed description thereof will be omitted. Thereafter, the processing proceeds to step S612.
[0264] In step S612, a feature region to be used for orientation correction is determined from the regions extracted manually or semi-automatically. Note that the process in step S612 is similar to the process in step S109 in Fig. 12, and therefore a detailed description thereof will be omitted. Thereafter, the process proceeds to step S613.
[0265] In step S613, determination information related to the determined (confirmed) characteristic region is registered (stored). Note that the process in step S613 is similar to the process in step S110 in Fig. 12 and step S308 in Fig. 22, and therefore a detailed description thereof will be omitted. Thereafter, the process shown in this flowchart ends.
[0266] In this embodiment, after the characteristic region is registered, an upright orientation determination process is performed on the determination target image based on the determination information related to the characteristic region, as in the first to third embodiments. The flow of the upright orientation determination process in this embodiment is generally similar to that described in the first embodiment with reference to Figures 15 and 16, and therefore will not be described again. [Explanation of symbols]
[0267] 1. Information processing equipment 21 Judgment information storage unit 22 Image Reception Department 23 Designated Reception Department 24 Feature Extraction Unit 25 Candidate Extraction Unit 26 Rotating part 27 Upright judgment section 28 Orientation correction unit 29 Display section 30 Instruction input reception unit 31 Common area extraction part 32 Reliability determination unit 33 Area determination part 9. Image acquisition device
Claims
1. a determination information storage means for storing determination features relating to a predetermined partial area in a predetermined format image, which is an image having a predetermined format, and the position of the partial area when the predetermined format image is in an upright position; an upright position determination means for determining whether or not a feature corresponding to the determination feature exists at the position of the input image to be determined, thereby determining whether or not the image to be determined is upright, the determination features include feature points and feature amounts, the upright determination means performs matching between the partial region in the image in the predetermined format and a region related to the position in the image to be determined using the determination features, thereby determining whether or not a feature corresponding to the determination features exists at the position in the image to be determined. Information processing device.
2. a determination information storage means for storing determination features relating to a predetermined partial area in a predetermined format image, which is an image having a predetermined format, and the position of the partial area when the predetermined format image is in an upright position; an uprightness determination means for determining whether or not a feature corresponding to the determination feature exists at the position of the input image to be determined, thereby determining whether or not the image to be determined is upright; image receiving means for receiving input of learning images having the predetermined format; a candidate extraction means for extracting, from the learning image in an upright state, an area that is likely to be included in images having the predetermined style, and setting the area as a candidate for the partial area; a designation receiving means for receiving a designation of the partial region in the learning image, the designation receiving means receives designation of the partial region by a user who has referred to the extracted candidates; the determination information storage means stores, as the determination feature and the position of the partial area when the learning image is upright, the determination feature and the position of the partial area accepted by the designation accepting means. Information processing device.
3. image receiving means for receiving input of learning images having the predetermined format; and a designation receiving means for receiving a designation of the partial region in the learning image, the determination information storage means stores, as the determination feature and the position of the partial area when the learning image is upright, the determination feature and the position of the partial area accepted by the designation accepting means. The information processing device according to claim 1 .
4. further comprising a rotation means for rotating the image within a range of 0 degrees or more and less than 360 degrees; the erection determination means performs a determination on the determination target image in a state rotated by the rotation means.
4. The information processing device according to claim 1.
5. the rotation means rotates the image to be determined at least one angle among one or a plurality of angles at which an outer edge shape of the image to be determined coincides with an outer edge shape of the image in an upright state of the predetermined format; the erection determination means performs a determination on the determination target image in a state rotated to any one of the one or more angles; The information processing device according to claim 4 .
6. When the image in the predetermined format is a rectangular image having long and short sides, the rotation means rotates the image to be determined to at least one of two angles at which the relationship between the long side or short side and the vertical side or horizontal side coincides with that of the image in the predetermined format when it is upright, the erection determination means performs a determination on the determination target image in a state rotated to one of the two angles; The information processing device according to claim 5 .
7. the determination information storage means stores the determination features at a resolution adjusted according to the size of the partial region; the erection determination means determines whether or not a feature corresponding to the determination feature exists in accordance with the resolution. The information processing device according to claim 1 .
8. the determination information storage means stores the determination features by defining an area obtained by adding a peripheral area to an area designated by a user as the partial area; The information processing device according to claim 1 .
9. When the peripheral area exceeds the edge of the image in the predetermined format, the determination information storage means adds a margin area as the peripheral area for the portion exceeding the edge, and stores the determination feature. The information processing device according to claim 8 .
10. the determination information storage means further stores a size of the image in the predetermined format; the upright determination means determines whether or not a determination target image that matches or is close to the size of the input determination target image is upright; The information processing device according to claim 1 .
11. a candidate extraction means for extracting an area that is likely to be included in images having the predetermined style from the learning image in an upright state, and setting the area as a candidate for the partial area; the designation receiving means receives a designation by a user who has referred to the extracted candidates; The information processing device according to claim 3 .
12. the predetermined format image is an image of a document having a predetermined format; The information processing device according to claim 1 .
13. the upright orientation determination means determines whether or not a feature corresponding to the determination feature exists in an area obtained by adding a peripheral area to an area related to the position of the determination target image. The information processing device according to claim 1 .
14. the upright orientation determination means detects an erroneous determination based on the relative positions of the feature points in the matched image in the predetermined format and the feature points in the determination target image as a result of the matching. The information processing device according to claim 1 .
15. the determination information storage means stores the determination features and the positions of the partial regions in a plurality of images in a predetermined format; the upright determination means determines the order of the predetermined format images used in determining the upright orientation of the determination target image based on the upright orientation determination result of the determination target image that was determined to be upright before the determination target image; The information processing device according to claim 1 .
16. The computer a determination information storage step for storing determination features relating to a predetermined partial region in a predetermined format image, which is an image having a predetermined format, and a position of the partial region when the predetermined format image is in an upright state; an uprightness determination step of determining whether or not a feature corresponding to the determination feature exists at the position of the input determination target image, thereby determining whether or not the determination target image is upright; the determination features include feature points and feature amounts, In the upright orientation determination step, a matching is performed between the partial region in the image in the predetermined format and a region related to the position in the image to be determined, using the determination features, thereby determining whether or not a feature corresponding to the determination features exists at the position in the image to be determined. Image orientation determination method.
17. Computer, a determination information storage means for storing determination features relating to a predetermined partial area in a predetermined format image, which is an image having a predetermined format, and the position of the partial area when the predetermined format image is in an upright position; the input image to be determined is made to function as an upright position determination means for determining whether or not a feature corresponding to the determination feature exists at the position of the input image to be determined, and the determination features include feature points and feature amounts, the upright determination means performs matching between the partial region in the image in the predetermined format and a region related to the position in the image to be determined using the determination features, thereby determining whether or not a feature corresponding to the determination features exists at the position in the image to be determined. program.
18. The computer a determination information storage step for storing determination features relating to a predetermined partial region in a predetermined format image, which is an image having a predetermined format, and a position of the partial region when the predetermined format image is in an upright state; an uprightness determination step of determining whether or not a feature corresponding to the determination feature exists at the position of the input determination target image, thereby determining whether or not the determination target image is upright; an image receiving step of receiving an input of a learning image having the predetermined format; a candidate extraction step of extracting a region that is likely to be included in images having the predetermined style from the learning image in an upright state, and setting the region as a candidate for the partial region; a designation receiving step of receiving designation of the partial region in the learning image; In the designation receiving step, designation of the partial region by a user who has referred to the extracted candidates is received; In the determination information storage step, the determination feature related to the partial area accepted in the designation accepting step and the position of the partial area in an upright state of the learning image are stored as the determination feature and the position. Image orientation determination method.
19. Computer, a determination information storage means for storing determination features relating to a predetermined partial area in a predetermined format image, which is an image having a predetermined format, and the position of the partial area when the predetermined format image is in an upright position; an uprightness determination means for determining whether or not a feature corresponding to the determination feature exists at the position of the input image to be determined, thereby determining whether or not the image to be determined is upright; image receiving means for receiving input of learning images having the predetermined format; a candidate extraction means for extracting, from the learning image in an upright state, an area that is likely to be included in images having the predetermined style, and setting the area as a candidate for the partial area; a designation receiving means for receiving a designation of the partial region in the learning image; the designation receiving means receives designation of the partial region by a user who has referred to the extracted candidates; the determination information storage means stores, as the determination feature and the position of the partial area when the learning image is upright, the determination feature and the position of the partial area accepted by the designation accepting means. program.
20. 1. An information processing system for determining a partial region to be used for determining whether an input image to be determined is upright, based on determination features related to the partial region in an image having a predetermined format and a position of the partial region when the image having the predetermined format is upright, comprising: image receiving means for receiving input of a plurality of learning images having the predetermined format; a common region extraction means for extracting one or more regions common to the plurality of learning images in an upright state as partial region candidates that are candidates for partial regions to be used for the upright orientation determination; a reliability determining means for determining a common region reliability for each of the one or more partial region candidates, the reliability indicating the degree to which the partial region candidate is suitable as a partial region used for the erection determination; a region determining means for determining one or more partial regions to be used for the upright determination from the one or more partial region candidates based on the common region reliability; An information processing system comprising:
21. The reliability determination means a commonality determining means for determining a commonality indicating a degree of likelihood that the partial region candidate is a region that is shared at approximately the same position among a plurality of images that have the predetermined style and are different from one another; an incompatibility determination means for determining whether the partial region candidate corresponds to an incompatibility region that is an area that is not suitable for the erection determination; a reliability calculation means for calculating a common area reliability based on the commonality and the determination result by the non-conformity determination means, 21. The information processing system according to claim 20.
22. the reliability determination means further includes an exclusion degree determination means for determining an exclusion degree indicating a degree to which the partial region candidate is to be excluded from the partial region candidates used for the upright orientation determination, in accordance with a determination result by the incompatibility determination means; the reliability calculation means calculates the common area reliability based on the commonality and the exclusion degree.
22. The information processing system according to claim 21.
23. the exclusion degree is an exclusion magnification, which is an index set so that the common region reliability of the partial region candidate decreases when the partial region candidate is the unsuitable region; the reliability calculation means calculates the common area reliability by multiplying the commonality by the exclusion magnification; 23. The information processing system according to claim 22.
24. The incompatible region is at least one of a region that may not be detected, a region that may cause the image orientation to be erroneously recognized as a different direction, and a region related to an object having a rotationally symmetric shape.
24. An information processing system according to any one of claims 21 to 23.
25. The region that may not be detected is a region that does not match at least one of the plurality of learning images in an upright state.
25. The information processing system according to claim 24.
26. The region that causes the image orientation to be erroneously recognized is a region that erroneously matches with at least one of the plurality of learning images that is not upright.
25. The information processing system according to claim 24.
27. the commonality determination means determines the commonality of the partial region candidates based on the attributes of the partial region candidates; 27. An information processing system according to any one of claims 21 to 26.
28. The attribute is at least one of information regarding the degree of matching between the plurality of learning images in the region related to the partial region candidate, the size and position of the region related to the partial region candidate, and the possibility of changing a character string in the region related to the partial region candidate.
28. The information processing system according to claim 27.
29. the degree of matching between the plurality of learning images in the region related to the partial region candidate includes at least one of information on the number of matching feature points between the plurality of learning images in the region related to the partial region candidate and the matching degree of the feature points; 29. The information processing system according to claim 28.
30. the common region extraction means extracts the partial region candidate based on a region common to two learning images selected from the plurality of learning images in the upright state; The region related to the partial region candidate is the region of the partial region candidate or a region common to the two learning images that served as the basis for the partial region candidate.
30. An information processing system according to claim 28 or 29.
31. further comprising a feature extraction means for extracting feature points and feature amounts for each of the plurality of learning images; the common region extraction means performs a matching process between the plurality of learning images based on the extracted feature points and feature amounts, and extracts a region including one or more feature points matched by the matching process as the partial region candidate; 31. The information processing system according to any one of claims 20 to 30.
32. the common region extraction means performs clustering on the feature points matched by the matching process to detect a region where the matched feature points are concentrated, and extracts the partial region candidate based on the concentrated region.
32. The information processing system according to claim 31.
33. The common area extraction means determining one or more pairs of two learning images from the plurality of learning images, and performing the matching process for each pair; extracting a region including one or more feature points matched by the matching process for each pair as a pair-wise partial region candidate, and extracting the partial region candidate based on the pair-wise partial region candidate for the determined one or more pairs; 33. An information processing system according to claim 31 or 32.
34. a matching means for performing pattern matching between at least one upright training image among the plurality of training images and other training images other than the at least one upright training image to determine the orientation of the other training images; and an orientation correction means for correcting the learning image determined by the comparison means to be not upright to an upright orientation, the common region extraction means extracts the partial region candidates using the plurality of upright learning images; 34. An information processing system according to any one of claims 20 to 33.
35. an instruction input receiving means for receiving an instruction to rotate an image from a user; a rotation means for rotating the image according to the rotation instruction; a format determining means for determining whether the plurality of rotated images have the same format as each other when the plurality of images are rotated by the rotating means; the image receiving means receives input of a plurality of images determined to have the same format as a plurality of learning images having the predetermined format; 35. An information processing system according to any one of claims 20 to 34.
36. further comprising an area information notification means for notifying information about the partial area determined by the area determination means; 36. An information processing system according to any one of claims 20 to 35.
37. further comprising a dimension compression means for compressing the dimensions of the extracted feature quantities, the common region extraction means performs the matching process using the feature values whose dimensions have been compressed; 34. An information processing system according to any one of claims 31 to 33.
38. the region determination means determines a region relating to a one-dimensional code or a two-dimensional code of the same type and approximately the same size at a similar position among a plurality of learning images having the predetermined format as the partial region; 38. An information processing system according to any one of claims 20 to 37.
39. the region determining means determines a region relating to a seal impression of the same type and of approximately the same size at a similar position among a plurality of learning images having the predetermined format as the partial region; 39. An information processing system according to any one of claims 20 to 38.
40. the feature extraction means extracts the feature points and the feature amounts for each of the plurality of learning images in each component of a color space; the common area extraction means performs a matching process between the plurality of learning images for matching between the same components and between different components based on the feature points and feature amounts extracted for each component of the color space.
38. An information processing system according to any one of claims 31 to 33 and 37.
41. The image forming apparatus further includes a determination information storage means for storing information about the partial area determined by the area determination means and used for determining whether the image is upright, the image receiving means receives input of other learning images having the predetermined format other than the plurality of learning images having the predetermined format; the region determining means determines one or more partial regions to be used for determining whether the image is upright again, based on the plurality of learning images and the other learning images having the predetermined format; an updating unit configured to update the information stored in the determination information storage unit based on the partial region determined based on the plurality of learning images and the other learning images; 41. The information processing system according to any one of claims 20 to 40.
42. a computer that determines a partial area to be used for determining whether an input image to be determined is upright based on determination features related to a partial area in an image having a predetermined format and the position of the partial area when the image having the predetermined format is upright, an image receiving step of receiving input of a plurality of learning images having the predetermined format; a common region extraction step of extracting one or more regions common to the plurality of learning images in an upright state as partial region candidates that are candidates for partial regions to be used for the upright orientation determination; a reliability determination step of determining a common region reliability for each of one or more of the partial region candidates, the reliability indicating the degree to which the partial region candidate is suitable as a partial region used for the erection determination; a region determining step of determining one or more partial regions to be used for the upright orientation determination from the one or more partial region candidates based on the common region reliability; Area determination method.
43. a computer that determines a partial area to be used for determining whether an input image to be determined is upright, based on determination features related to a partial area in an image having a predetermined format and the position of the partial area when the image having the predetermined format is upright; image receiving means for receiving input of a plurality of learning images having the predetermined format; a common region extraction means for extracting one or more regions common to the plurality of learning images in an upright state as partial region candidates that are candidates for partial regions to be used for the upright orientation determination; a reliability determining means for determining a common region reliability for each of the one or more partial region candidates, the reliability indicating the degree to which the partial region candidate is suitable as a partial region used for the erection determination; and determining, based on the common area reliability, one or more partial areas to be used for the upright orientation determination from the one or more partial area candidates. program.
Citation Information
Patent Citations
Card arrangement direction identification method, device and image processing device
CN108564081A
Image recognition device
JP2000032247A
Apparatus for generating form dictionary, apparatus for identifying form, method of generating form dictionary, and program
JP2010262578A
Device and method for reading driver's license
JP2014071784A
Image processor, image processing method, and program
JP2014071866A