Image processing apparatus, image processing system, and program

The image processing apparatus addresses the challenge of accurately extracting characters from images by using contour extraction and non-character pixel removal techniques, resulting in improved OCR accuracy and character visibility.

JP2025091589APending Publication Date: 2025-06-19RICOH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023206905
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Conventional image processing techniques struggle with accurately extracting characters from images where characters resemble straight lines or are overlapped with figures other than rectangles, leading to reduced character extraction accuracy.

Method used

An image processing apparatus that includes a contour extraction unit to extract contours of foregrounds in both the original image and a character-extracted image, a non-character pixel identification unit to identify non-character pixels by comparing the contours, and a non-character pixel removal unit to generate an image excluding these identified pixels.

Benefits of technology

The proposed solution significantly improves the accuracy of character extraction from images, enhancing OCR accuracy and character visibility by effectively removing non-character pixels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025091589000001_ABST
    Figure 2025091589000001_ABST
Patent Text Reader

Abstract

To improve the accuracy of extracting characters from an image.SOLUTION: An image processing apparatus includes: a contour extraction unit which extracts a contour of foreground of a first image, and extracts a contour of foreground of a second image formed by extracting a set of pixels from the foreground of the first image, the set of pixels being estimated to constitute characters; a non-character pixel identifying unit which identifies pixels that do not constitute characters in the set of pixels, on the basis of the comparison between the contour extracted for the first image and the contour extracted for the second image; and a non-character pixel exclusion unit which generates a third image having, as foreground, a set of pixels obtained by excluding the pixels identified not to constitute characters in the set of pixels.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing system, and a program.

Background Art

[0002] By generating an image that extracts only the characters entered in a document from an image of the document including ruled lines such as an order form, it is possible to improve OCR accuracy and the visibility of the characters. However, when the ruled line and the characters overlap in the image from which the characters are extracted, a part of the ruled line may remain in the extracted image without being completely erased.

[0003] On the other hand, as a technique for extracting handwritten characters, a straight line is extracted using a conventional image processing technique, intersections between the straight line and the characters are extracted, and an entry field (rectangle) is extracted, and only the characters are extracted from an image in which the characters and the straight line overlap within the entry field. Furthermore, a technique has been proposed to improve OCR accuracy by registering the image, the OCR misrecognition result, and the correct result when OCR misrecognizes from the result of OCR of the extracted characters (Patent Document 1).

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the conventional technique, for an image in which characters similar to a straight line are continuous or an image including a figure other than a rectangle such as an underline, a part other than the characters is erroneously extracted, resulting in a problem that the accuracy of character extraction is lowered.

[0005] The present invention has been made in view of the above points, and an object thereof is to improve the accuracy of character extraction from an image.

Means for Solving the Problems

[0006] Therefore, in order to solve the above problems, an image processing apparatus includes: a contour extraction unit that extracts a contour of a foreground of a first image and extracts a contour of a foreground of a second image from which a set of pixels estimated to constitute characters has been extracted from the foreground of the first image; a non-character pixel identification unit that identifies, based on a comparison between the contour extracted for the first image and the contour extracted for the second image, pixels in the set of pixels that do not constitute characters; and a non-character pixel removal unit that generates a third image having, as a foreground, a set of pixels excluding the pixels identified as not constituting characters from the set of pixels.

Effect of the Invention

[0007] The extraction accuracy of characters from an image can be improved.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a diagram showing a hardware configuration example of an image processing apparatus 10 in a first embodiment. The image processing apparatus 10 in FIG. 1 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, and an interface device 105, etc., which are mutually connected by a bus B.

[0010] A program for realizing the processing in the image processing apparatus 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the installation of the program does not necessarily have to be performed from the recording medium 101, and it may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program and also stores necessary files, data, etc.

[0011] When an activation instruction for the program is given, the memory device 103 reads out and stores the program from the auxiliary storage device 102. The processor 104 is a CPU or a GPU (Graphics Processing Unit), or both a CPU and a GPU, and executes the functions related to the image processing apparatus 10 according to the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to a network.

[0012] FIG. 2 is a diagram showing a functional configuration example of the image processing apparatus 10 in the first embodiment. In FIG. 2, the image processing apparatus 10 includes an input image acquisition unit 11, an input image correction unit 12, a character extraction unit 13, a character extraction image correction unit 14, an image processing unit 15, a communication unit 16, an output unit 17, and the like. Each of these units is realized by a process executed by the CPU 104 for one or more programs installed in the image processing apparatus 10. The image processing apparatus 10 also uses an input image storage unit 121, a corrected input image storage unit 122, a character extraction storage unit 123, a character extraction image storage unit 124, a corrected character extraction image storage unit 125, an output image storage unit 126, and the like. Each of these storage units can be realized using, for example, an auxiliary storage device 102, a memory device 103, or a storage device that can be connected to the image processing apparatus 10 via a network.

[0013] The input image acquisition unit 11 acquires an image (hereinafter referred to as an "input image") input (scanned or photographed, etc.) by a scanner, a camera, or the like. The input image acquisition unit 11 stores the acquired input image in the input image storage unit 121.

[0014] The input image correction unit 12 executes image correction processing on the input image stored in the input image storage unit 121, and stores the image (hereinafter referred to as a "corrected input image") generated by the image correction processing in the corrected input image storage unit 122. The image correction processing is, for example, image processing such as γ correction, skew correction, and noise removal when the input image is a scanned image, and image processing such as distortion correction when the input image is a camera image. The position of each pixel in the corrected input image is the same as the position of the pixel corresponding to the pixel in the input image.

[0015] The character extraction unit 13 extracts a set of pixels (hereinafter referred to as "character pixels") estimated to constitute handwritten characters from the corrected input image stored in the corrected input image storage unit 122, and generates an image (hereinafter referred to as "character extraction image") with the character pixels as the foreground and the other pixels as the background. The character extraction unit 13 stores the generated character extraction image in the character extraction image storage unit 124. Note that the character extraction image is a binary image in which the pixel value of the foreground pixel is 1 and the pixel value of the background pixel is 0. The position of each character pixel is the same as the position of the pixel corresponding to the character pixel in the corrected input image. In other words, the character extraction image is the same image as the image in which only the character pixels are the foreground and the rest are the background in the corrected input image.

[0016] The character extraction unit 13 is, for example, a character extraction AI (character extraction model), and generates a character extraction image based on the learned model parameters using data including the characters to be extracted in advance. For the character extraction AI, for example, an algorithm for labeling each pixel in the image can be used. For example, Semantic Segmentation using CNN can be cited as an example of the character extraction AI. By performing character extraction (hereinafter referred to as "character extraction") using an AI (machine learning model), it is possible to handle differences in handwritten characters due to writing utensils, writers, etc., and various documents with different printed characters, ruled lines, drawings, etc. In addition, it becomes possible to easily extract handwritten characters from free-format documents that cannot be fully handled by conventional rule-based processing. Note that the character extraction unit 13 uses the character extraction storage unit 123 as a temporary memory for operations for character extraction (for example, during the operation of CNN).

[0017] The character extraction image correction unit 14 performs correction processing (hereinafter referred to as "character extraction correction processing") to exclude pixels that are highly likely not to constitute characters but are extracted as character pixels (for example, pixels constituting grid lines) from the character extraction image stored in the character extraction image storage unit 124. The character extraction image correction unit 14 stores the image obtained as a result of the character extraction correction processing (hereinafter referred to as "corrected character extraction image") in the corrected character extraction image storage unit 125.

[0018] The image processing unit 15 performs processing according to the usage purpose on the corrected character extraction image stored in the corrected character extraction image storage unit 125, and stores the data obtained as a processing result in the output image storage unit 126.

[0019] For example, the image processing unit 15 may generate an image such as a PDF with text embedded in the character pixel part, where the text is obtained by performing OCR on the corrected character extraction image.

[0020] Alternatively, the image processing unit 15 may generate an image such as a PDF with text by performing OCR for handwritten characters on the handwritten character part of the corrected input image stored in the corrected input image storage unit 122 and performing OCR for typeset characters on the typeset part. At this time, the image processing unit 15 identifies the handwritten character part (character pixels) based on the corrected character extraction image stored in the corrected character extraction image storage unit 125.

[0021] Alternatively, the image processing unit 15 may generate an image in which only the handwritten character part of the corrected input image stored in the corrected input image storage unit 122 is colored. At this time, the image processing unit 15 identifies the handwritten character part (character pixels) based on the corrected character extraction image stored in the corrected character extraction image storage unit 125.

[0022] Alternatively, the image processing unit 15 may generate an image obtained by deleting only the handwritten characters in the corrected input image stored in the corrected input image storage unit 122. At this time, the image processing unit 15 identifies the handwritten character portions (character pixels) based on the corrected character extraction image stored in the corrected character extraction image storage unit 125.

[0023] The communication unit 16 transmits, as necessary, the data (processing result by the image processing unit 15) stored in the output image storage unit 126 to, for example, a PC or the like connected via a network.

[0024] The output unit 17 controls, for example, the printing out of the image stored in the output image storage unit 126 when a printer is implemented in the image processing apparatus 10.

[0025] Since the image processing apparatus 10 has the functional configuration as shown in FIG. 2, for example, the image processing apparatus 10 can be applied to the following situations.

[0026] FIG. 3 is a first diagram for explaining an application example of the image processing apparatus 10. In FIG. 3, a case will be described where a scanned image generated by scanning a document in which characters are handwritten on a paper document with a scanner such as a multifunction device is the input image.

[0027] On the paper document, characters, ruled lines, drawings, etc. are printed in advance, and characters are handwritten thereon. The handwritten characters include various characters due to differences in the color and thickness of the writing instrument used to write the characters, the degree of rubbing, and the size and shape of the characters by the writer. On the other hand, the characters, ruled lines, drawings, etc. printed in the document in advance also have various colors, fonts, sizes, thicknesses, etc.

[0028] The input image acquisition unit 11 acquires, as an input image, a scanned image generated by scanning such a paper document with a scanner such as a multifunction printer. The input image correction unit 12 generates a corrected input image by performing processes such as gamma correction, color correction, skew correction, and noise removal on the input image, which is a scanned image, as necessary. The character extraction unit 13, which is a character extraction AI that has been learned in advance to extract handwritten characters, generates a character extraction image by extracting pixels that constitute the handwritten characters from the corrected input image.

[0029] For example, when a ruled line overlaps with a handwritten character in the input image (corrected input image), a part of the ruled line that overlaps with the handwritten character (for example, a part of the ruled line included in the circle c1 in the character extraction image of FIG. 3) may remain in the character extraction image. Therefore, the character extraction image correction unit 14 generates a corrected character extraction image in which such a part of the ruled line is deleted by performing a character extraction correction process on the character extraction image. Thereby, it is possible to improve the OCR accuracy and the visibility of the handwritten characters.

[0030] Here, the characters to be extracted are handwritten characters, but not limited to handwritten characters. By changing the data during the learning of the character extraction AI, it is also possible to extract characters such as printed characters and numbers.

[0031] FIG. 4 is a second diagram for explaining an application example of the image processing apparatus 10. In FIG. 4, a case where the input image is an image obtained by photographing a label attached to a cargo or the like with a camera will be described.

[0032] The cargo has a label with printed characters and handwritten characters. The input image acquisition unit 11 acquires, as an input image, an image obtained by photographing such a label with a camera. The input image correction unit 12 generates a corrected input image by performing processes such as distortion correction and color correction on the input image as necessary. The character extraction unit 13, which is a character extraction AI that has been learned in advance to extract handwritten characters, generates a character extraction image by extracting pixels that constitute the handwritten characters from the corrected input image.

[0033] For example, when a ruled line overlaps with a handwritten character in the input image, there may be a case where a part of the ruled line that overlaps with the handwritten character (for example, a part of the ruled line included in the circle c2 in the character extraction image of FIG. 4) remains. Therefore, the character extraction image correction unit 14 generates a corrected character extraction image in which such a part of the ruled line is deleted by executing character extraction correction processing on the character extraction image. Thereby, it is possible to improve the OCR accuracy and the visibility of the handwritten characters.

[0034] Here, the characters to be extracted are handwritten characters. However, not limited to handwritten characters, by changing the data at the time of character extraction AI learning, it is also possible to extract characters such as printed characters, symbols, and numbers written in a specific area within the label.

[0035] Next, the details of the character extraction image correction unit 14 will be described.

[0036] FIG. 5 is a diagram showing a configuration example of the character extraction image correction unit 14. In FIG. 5, the character extraction image correction unit 14 includes a binarization unit 141, a contour extraction unit 142, a non-character pixel specifying unit 143, and a non-character pixel removing unit 144. The processes executed by these units will be described with reference to FIGS. 6 and 7.

[0037] FIGS. 6 and 7 are diagrams for explaining the character extraction correction process in the first embodiment.

[0038] The binarization unit 141 binarizes the corrected input image. By binarization, the set of pixels constituting the corrected input image is divided into a foreground and a background. The foreground includes pixels with a pixel value of 1, such as pixels that become handwritten characters, ruled lines, and printed characters (hereinafter referred to as "foreground pixels"). The background includes pixels with a pixel value of 0.

[0039] In FIG. 6, an image g1 with the label "<input image>" is a part of the corrected input image stored in the corrected input image storage unit 122. An image g2 with the label "<character extraction image>" is an image of a portion (range) corresponding to the image g1 in the character extraction image stored in the character extraction image storage unit 124. Hereinafter, the image g2 is referred to as "target character extraction image g2".

[0040] For the purpose of the following description, an image g3 is obtained by superimposing the range of a rectangular region R1 in an image (hereinafter referred to as "target input image g1") obtained by binarizing the image g1 by the binarization unit 141 and the range of a rectangular region R2 in the target character extraction image g2. The rectangular region R2 is a rectangular region having the same coordinates as the rectangular region R1. Note that the image g3 is an image illustrated for convenience of explanation and is not an actually generated image.

[0041] Among the regions of the image g3, pixels whose boundaries are shown are foreground pixels, and pixels constituting regions other than the foreground pixels are background pixels. Also, among the foreground pixels, pixels with hatching are pixels that overlap with the character pixels extracted by the character extraction unit 13. In other words, among the foreground pixels, pixels without hatching are pixels that are not character pixels (hereinafter also referred to as "non-character images").

[0042] (1) to (7) of FIG. 7 are diagrams showing in order how the rectangular region R3 in the image g3 changes by the processing executed by the character extraction image correction unit 14 for the rectangular region R3 in the image g3.

[0043] The contour extraction unit 142 of the character extraction image correction unit 14 extracts the foreground contour of the target character extraction image g2 and the foreground contour of the target input image g1. More specifically, the contour extraction unit 142 focuses on a certain pixel (hereinafter, the pixel being focused on is referred to as the "focus pixel"), and extracts the next pixel that constitutes the contour from among the pixels adjacent to the focus pixel (for example, extracts one pixel whose pixel value difference from the focus pixel is within a threshold value (for example, 0)). The contour extraction unit 142 repeats the extraction of the contour with the next pixel as the focus pixel.

[0044] The non-character pixel identification unit 143 of the character extraction image correction unit 14 identifies, based on a comparison between the contour extracted for the target input image g1 and the contour extracted for the target character extraction image g2, the pixels that do not constitute characters (that is, the pixels that are part of the grid lines) among the set of pixels that constitute the foreground in the target character extraction image g2.

[0045] In Fig. 7(1), a process is shown in which the contour extraction unit 142 performs contour extraction of the foreground image by tracing the focus pixel one pixel at a time in the directions of arrows a11 to a17 with the foreground pixel marked with a black circle (●) as the starting point of the focus pixel. Each time the contour extraction unit 142 extracts the next pixel that becomes the contour from the focus pixel, it determines whether the next pixel is the same in the target input image g1 and the target character extraction image g2. Here, assume that the focus pixel has moved to the pixel marked with "*" in Fig. 7(1) (that is, assume that the contour extraction has reached the focus pixel *). At this time, for the target character extraction image g2, the contour extraction unit 142 extracts the next pixel (contour) in the direction of arrow g18 (rightward direction), but for the target input image g1, it extracts the next pixel (contour) in the direction of arrow g19, so the next pixel is different between the target character extraction image g2 and the target input image g1.

[0046] In this case, when the non-character pixel specifying unit 143 satisfies the condition (hereinafter referred to as the "straight-line search condition") that the number of pixels constituting the foreground in the direction of the pixel to be extracted next to the target pixel in the target input image g1 (the direction of arrow a19) is equal to or greater than a predetermined value (a threshold value set in advance, which is 4 indicated by both arrows in the example of FIG. 7) (hereinafter referred to as the "first condition"), and the number of pixels constituting the foreground in the target character extraction image in the direction opposite to the said direction (i.e., the direction of arrow a18) is equal to or greater than a predetermined value (a threshold value set in advance, which is 2 in the example of FIG. 7) (hereinafter referred to as the "second condition"), all the pixels that form a straight line by tracing back the pixels extracted as the contour from the target pixel (the target pixel and each contour pixel located on the root side of each of arrows a13 to a17 that have been traced until reaching the target pixel) are specified as pixels that do not constitute a character (non-character pixels). That is, the non-character pixel specifying unit 143 determines that these contour pixels are pixels that have been erroneously extracted as character pixels by the character extraction unit 13 despite being grid lines.

[0047] The non-character pixel removing unit 144 generates an image with the foreground being the set of pixels excluding the pixels (non-character pixels) specified by the non-character pixel specifying unit 143 from the set of pixels constituting the foreground of the target character extraction image g2. In other words, the non-character pixel removing unit 144 removes the pixels (non-character pixels) specified by the non-character pixel specifying unit 143 from the foreground.

[0048] In FIG. 7(2), a cross (×) is given to the pixels removed by the non-character pixel removing unit 144. Subsequently, the contour extraction unit 142 performs contour extraction on the foreground pixels again. This state is (2). In the state of (2), the target pixel moves to the foreground image with a "*" given at the tip of arrow a11.

[0049] In (3), the contour extraction unit 142 continues to perform contour extraction on the image g3 that does not include the pixels removed in (2). At this time, for the contour extraction, the pixels removed in (2) (the pixels surrounded by the dashed line in (3)) are not the objects of contour extraction in either the target input image g1 or the target character extraction image g2. The contour extraction unit 142 extracts the contours of the target input image g1 and the target character extraction image g2 in the same manner as in (1) (arrows a31 to a35). When the directions of contour extraction (the next pixel of the pixel of interest) are different between the target character extraction image g2 and the target input image g1 (arrows a36 and a37), the non-character pixel identification unit 143 determines whether the straight-line search condition is satisfied. When the condition is satisfied, the non-character pixel identification unit 143 identifies the pixel of interest at that time and each contour pixel located on the root side of each of arrows a33 to a35 as non-character pixels. The non-character pixel removal unit 144 removes these contour pixels from the foreground of the target character extraction image g2 as shown in (4).

[0050] By repeating the above, the rectangular region R3 changes as shown in (5) to (7). As a result, the pixels of the ruled line that the character extraction unit 13 has erroneously extracted as character pixels can be deleted from the target character extraction image g2.

[0051] Note that for contour extraction, a known contour tracing algorithm such as Freeman's chain code is used, for example.

[0052] Hereinafter, the processing procedure executed by the image processing apparatus 10 will be described. FIG. 8 is a flowchart for explaining an example of the processing procedure executed by the image processing apparatus 10 in the first embodiment.

[0053] In step S101, the input image acquisition unit 11 acquires an input image and stores the input image in the input image storage unit 121.

[0054] Subsequently, the input image correction unit 12 generates a corrected input image by performing image correction processing on the input image stored in the input image storage unit 121, and stores the corrected input image in the corrected input image storage unit 122 (S102).

[0055] Subsequently, the character extraction unit 13 generates a character extraction image (for example, an image including the target character extraction image g2 in FIG. 6) that is a binary image from the corrected input image stored in the corrected input image storage unit 122, and stores the character extraction image (hereinafter referred to as the "target character extraction image") in the character extraction image storage unit 124 (S103).

[0056] Steps S104 and subsequent steps are character extraction correction processing by the character extraction image correction unit 14.

[0057] In step S104, the binarization unit 141 acquires the corrected input image from the corrected input image storage unit 122 and binarizes the target input image (S104). By binarization, a binary image (for example, an image including the target input image g1 described in FIG. 6) is generated in which the pixel values of the characters and ruled lines (foreground image) in the corrected input image are 1 and the pixel values of the background part (background pixels) are 0. Hereinafter, the binary image is referred to as the "target input image".

[0058] Subsequently, the contour extraction unit 142 performs a raster scan on the target character extraction image for only one scan line in raster order (S105).

[0059] FIG. 9 is a diagram for explaining the raster scan and contour extraction. In FIG. 9, an example in which the raster scan proceeds in the order of (1), (3), and (5) is shown. In this example, the order from the upper left to the lower right of the image is the raster order. What is scanned in one raster scan is a column of pixels for one row until character pixels are detected in the main scanning direction or until the right end is reached.

[0060] Subsequently, the contour extraction unit 142 determines whether the raster scan has reached the end of the target character extraction image in raster order (S106). Here, the end corresponds to the pixel at the lower right vertex in the example of FIG. 9, and the state of (3) is the state where the end has been reached.

[0061] When the end has not been reached (No in S106), the contour extraction unit 142 determines whether a character pixel (a pixel with a pixel value of 1) has been detected by raster scanning (whether it has reached the left end) (S107). For example, in the first arrow and the second arrow in FIG. 9(1), the left end has been reached. In this case (No in S107), the contour extraction unit 142 returns to step S105 and continues raster scanning.

[0062] On the other hand, when a character pixel (a pixel with a pixel value of 1) is detected by raster scanning as in the third arrow in FIG. 9(1) (Yes in S107), the contour extraction unit 142 sets the character pixel as a target pixel for performing the contour extraction described in FIG. 7 (S108). When step S108 is executed for the first time following step S107, the contour extraction unit 142 stores the coordinate information of the target pixel in the memory device 103 or the auxiliary storage device 102 as contour information.

[0063] FIG. 10 is a diagram showing a first example of contour information. On the upper side of FIG. 10, extraction sequence numbers (1) to (8) are assigned to each pixel extracted in the contour extraction described in FIG. 6(1). On the lower side of FIG. 10, the contour information (1) generated in the contour extraction is shown.

[0064] First, pixel (1) is the target pixel. Therefore, the first row (the X coordinate and Y coordinate of pixel (1)) in the contour information (1) is stored.

[0065] Subsequently, the contour extraction unit 142 performs contour extraction with the target pixel as the starting point (for example, in the state of FIG. 6(1), the pixel marked with a black circle (●)) for each of the target character extraction image and the target input image (S109). In one contour extraction, any one of the pixels adjacent to the target pixel is extracted as the next pixel constituting the contour. In FIG. 6(1), the pixel at the tip of arrow a11 (pixel (2) in FIG. 10) is extracted as the next pixel constituting the contour.

[0066] Subsequently, the contour extraction unit 142 determines whether the directions of contour extraction are different between the target character extraction image and the target input image (S110). Specifically, the contour extraction unit 142 determines whether the positions (coordinates) of the next pixels extracted for each of the target character extraction image and the target input image in step S109 are different.

[0067] For example, when the pixel marked with a black circle (●) in FIG. 6(1) is the pixel of interest, the directions of contour extraction are the same between the target character extraction image and the target input image. In this case) (No in S110), the contour extraction unit 142 repeats steps S108 and subsequent steps. However, in step S108 executed when the result in step S110 is No, the contour extraction unit 142 sets the next pixel extracted in step S109 as the pixel of interest, and adds the coordinate information of the pixel of interest to the contour information.

[0068] On the other hand, when the pixel marked with "*" in FIG. 6(1) (pixel (8) in FIG. 10) is the pixel of interest, as described in FIG. 6, the directions of contour extraction are different between the target character extraction image and the target input image. Therefore, in this case (Yes in S110), the non-character pixel identification unit 143 executes a straight line search process (S111). Specifically, the non-character pixel identification unit 143 determines whether the straight line search conditions are satisfied. The straight line search conditions are that the first condition that the number of pixels constituting the foreground in the direction to the pixel extracted next to the pixel of interest in the target input image is equal to or greater than a predetermined value, and the second condition that the number of pixels constituting the foreground in the target character extraction image in the direction opposite to the said direction (that is, the direction to the pixel extracted next to the pixel of interest in the target character extraction image) is equal to or greater than a predetermined value are both satisfied.

[0069] When the straight line search conditions are not satisfied (No in S112), the contour extraction unit 142 repeats steps S108 and subsequent steps. However, in step S108 executed when the result in step S112 is No, the contour extraction unit 142 sets the next pixel extracted in step S109 as the pixel of interest, and adds the coordinate information of the pixel of interest to the contour information.

[0070] When the straight line search condition is satisfied (Yes in S112), the non-character pixel specifying unit 143 specifies, in the target character extraction image, all the pixels from the current pixel of interest up to the pixel forming a straight line when tracing back as non-character pixels (S113). Subsequently, the non-character pixel removal unit 144 removes the pixels specified as non-character pixels from the foreground (outline) of the target character extraction image (S114). Specifically, the non-character pixel removal unit 144 sets the pixel value of the corresponding pixel in the target character extraction image to 0. The non-character pixel removal unit 144 also deletes the coordinate information of the corresponding pixel from the outline information. As a result, the rectangular region R3 and the outline information (1) in FIG. 10 become the state of the rectangular region R3 and the outline information (2) in FIG. 11.

[0071] Subsequently, the outline extraction unit 142 determines whether or not the outline extraction has reached the end (S115). That the outline extraction has reached the end means that the next pixel extracted in step S109 matches the character pixel detected by the raster scan in step S105 (that is, the pixel that became the start point of the outline extraction) (for example, as shown by the arrow in (2) of FIG. 9, the outline extraction has returned to the start point).

[0072] When the outline extraction has not reached the end (No in S115), the outline extraction unit 142 repeats steps S108 and subsequent steps. However, in step S108 executed when the result in step S115 is No, the outline extraction unit 142 sets the pixel corresponding to the last coordinate information in the outline information (pixel (2) in FIG. 11) as the pixel of interest. At this time, the outline information is not updated.

[0073] When the contour extraction reaches the end (Yes in S115), the contour extraction unit 142 re-executes the steps after S105. In this case, the raster scan resumes from the pixel next to the pixel reached in the previous raster scan. For example, if the raster scan has progressed up to Fig. 9(1), the raster scan resumes as shown in Fig. 9(3). When the character image is reached again, contour extraction is performed as shown in Fig. 9(4). Thereafter, the raster scan is executed as shown in Fig. 9(5). Note that when resuming the raster scan, a known method is used, such as excluding the pixels that have been contour-extracted or the pixels surrounded by the contour from the detection target, to determine the pixel at the resumption location of the raster scan.

[0074] When the raster scan reaches the end of the target character extraction image as shown in Fig. 9(5) (Yes in S106), the non-character pixel removal unit 144 corrects the target character extraction image at that time and stores it in the corrected character extraction image storage unit 125 as the corrected character extraction image (S116).

[0075] As described above, according to the first embodiment, it is possible to identify the pixels that are erroneously detected as pixels constituting the character, and generate an image (corrected character extraction image) from which the pixels have been removed. Therefore, the extraction accuracy of characters from the image can be improved. As a result, for example, the OCR accuracy and the visibility of handwritten characters can be improved.

[0076] Next, a second embodiment will be described. In the second embodiment, the points different from the first embodiment will be described. Therefore, for the points not particularly mentioned, they may be the same as those in the first embodiment.

[0077] In the second embodiment, even when there is a deviation in the contour of the ruled line portion between the corrected input image and the corrected character extraction image, it is considered.

[0078] Fig. 12 is a diagram for explaining the character extraction correction process in the second embodiment. In Fig. 12, the parts corresponding to Fig. 7 are denoted by the same reference numerals.

[0079] In (0) of FIG. 12, an example is shown in which, in the binarized image of image g1 (hereinafter referred to as the “target input image”), the width of the foreground pixels corresponding to the grid lines increases by 4 pixels on the left side of the marked pixel with “*”. In this case, at the marked pixel *, the direction of contour extraction in the target input image is the direction of arrow a20. Since the direction of arrow a20 is different from the direction of arrow a19, the character extraction image correction unit 14 can detect that the direction of contour extraction is different between the target input image and image g2 (hereinafter referred to as the “target character extraction image”). However, since the direction of arrow a20 does not satisfy the condition that “the number of pixels constituting the foreground is equal to or more than a predetermined value”, the grid line portion included in the target character extraction image cannot be deleted.

[0080] Therefore, in the second embodiment, the straight line search condition is changed to the condition that, as shown in FIG. 12(1), the number of pixels constituting the foreground is equal to or more than a predetermined value in the direction from the marked pixel to the next pixel extracted in the target character extraction image, and the number of pixels constituting the foreground is equal to or more than a predetermined value in the target input image in the direction opposite to the said direction. As shown in FIG. 12(2), the character pixels to be deleted from the contour are the same as those in the first embodiment. The same processing is performed from FIG. 12(3) onwards, and finally the same result as FIG. 7 is obtained.

[0081] As described above, according to the second embodiment, even when the shape of the grid line is different between the input image and the character extraction image, the non-character extraction image can be removed.

[0082] Incidentally, as shown in FIG. 2, a configuration in which the image processing apparatus includes all functions is suitable, for example, when it is not desired to upload data to the cloud from a security perspective, or when it is desired to complete all processing only with an edge device. However, each of the above embodiments is not limited to the arrangement form as shown in FIG. 2 for each function.

[0083] FIG. 13 is a diagram showing a first modification of the arrangement form of the functional configuration. In FIG. 13, the same parts as those in FIG. 2 are denoted by the same reference numerals, and the description thereof is omitted as appropriate.

[0084] FIG. 13 shows a configuration assuming that after performing handwritten character extraction and character extraction correction processing on an edge device such as a multifunction device, a corrected character extraction image is used on the cloud server 20. The image processing device 10 on the edge side can confirm the result with the output unit 17 (printer or monitor), and the cloud server 20 can create a database by inputting it to OCR for text conversion.

[0085] Specifically, the image processing device 10 is connected to the cloud server 20 via a network such as the Internet. The cloud server 20 has the image processing unit 15 and the output image storage unit 126 that the image processing device 10 had in FIG. 2. That is, in FIG. 2, the image processing unit 15 is realized by the processing executed by the program installed in the cloud server 20 on the processor of the cloud server 20. The output image storage unit 126 can be realized using cloud storage or the like.

[0086] The communication unit 16 of the image processing device 10 transmits the corrected character extraction image stored in the corrected character extraction image storage unit 125 to the cloud server 20. The communication unit 16 may also transmit the corrected input image stored in the corrected input image storage unit 122 to the cloud server 20.

[0087] The image processing unit 15 of the cloud server 20 executes the processing described in FIG. 2.

[0088] The output unit 17, for example, when a monitor or a printer is mounted on the image processing device 10, displays or prints out the data stored in the output image storage unit 126.

[0089] FIG. 14 is a diagram showing a second modification of the arrangement form of the functional configuration. In FIG. 14, the same parts as those in FIG. 2 are denoted by the same reference numerals, and the description thereof is omitted as appropriate.

[0090] FIG. 14 shows a configuration assuming that a camera image is acquired by an edge device including a camera installed in, for example, a factory or a warehouse, and handwriting character extraction and character extraction image correction processing are executed on the cloud. In the case of FIG. 14, the device can be miniaturized by minimizing the edge-side implementation.

[0091] Specifically, the image processing apparatus 10 is connected to a cloud server 20 via a network such as the Internet. The cloud server 20 has an input image correction unit 12, a character extraction unit 13, a character extraction image correction unit 14, an image processing unit 15, a communication unit 16, an output unit 17, a corrected input image storage unit 122, a character extraction image storage unit 124, a corrected character extraction image storage unit 125, an output image storage unit 126, etc., which the image processing apparatus 10 had in FIG. 2.

[0092] The communication unit 16 of the image processing apparatus 10 transmits the input image stored in the input image storage unit 121 by the input image acquisition unit 11 to the cloud server 20. The cloud server 20 executes the processing procedure described in FIG. 8 based on the input image.

[0093] Note that each function of the above-described embodiments can be realized by one or a plurality of processing circuits. Here, the "processing circuit" in this specification refers to a processor programmed to execute each function by software like a processor implemented by an electronic circuit, an ASIC (Application Specific Integrated Circuit) designed to execute each function described above, a DSP (digital signal processor), an FPGA (field programmable gate array), or a device such as a conventional circuit module.

[0094] Note that in each of the above embodiments, the target input image is an example of the first image. The target character extraction image is an example of the second image. The corrected character extraction image is an example of the third image.

[0095] As described above in detail regarding the embodiments of the present invention, the present invention is not limited to such specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.

[0096] Aspects of the present invention are, for example, as follows. <1> A contour extraction unit that extracts the contour of the foreground of the first image and extracts the contour of the foreground of the second image from which a set of pixels estimated to constitute characters has been extracted from the foreground of the first image; A non-character pixel identification unit that identifies non-character pixels among the set of pixels based on a comparison between the contour extracted for the first image and the contour extracted for the second image; A non-character pixel removal unit that generates a third image having, as the foreground, a set of pixels excluding the pixels identified as not constituting characters among the set of pixels; An image processing apparatus characterized by comprising: <2> In the process where the contour extraction unit extracts the contour of the foreground of the second image pixel by pixel, when the pixel extracted next to a certain pixel is different from the pixel extracted next to the corresponding pixel of the certain pixel in the first image, all the pixels forming a straight line traced back from the certain pixel among the pixels extracted as the contour for the second image are identified as non-character pixels. The image processing apparatus according to <1>, characterized by the above. <3> The non-character pixel identification unit, when the number of pixels constituting the foreground in the direction from the pixel corresponding to the certain pixel in the first image to the next extracted pixel is equal to or greater than a predetermined value, and the number of pixels constituting the foreground in the second image in the direction opposite to the said direction is equal to or greater than a predetermined value, identifies all the pixels forming a straight line traced back from the certain pixel among the pixels extracted as the contour as non-character pixels. The image processing apparatus according to <2>, characterized by the above. <4> When the number of pixels constituting the foreground in the direction from a certain pixel to the next pixel extracted in the second image is equal to or greater than a predetermined value, and the number of pixels constituting the foreground of the first image in the direction opposite to the said direction is equal to or greater than a predetermined value, all pixels that form a straight line by tracing back the pixels extracted as the contour from the said certain pixel are specified as pixels that do not constitute characters. The image processing apparatus according to <2>, characterized in that. <5> The second image is an image from which a set of pixels estimated to constitute handwritten characters has been extracted from the foreground of the first image. The image processing apparatus according to any one of <1> to <4>, characterized in that. <6> A contour extraction unit that extracts the contour of the foreground of the first image and extracts the contour of the foreground of the second image from which a set of pixels estimated to constitute characters has been extracted from the foreground of the first image; A non-character pixel specifying unit that specifies, based on a comparison between the contour extracted for the first image and the contour extracted for the second image, pixels that do not constitute characters among the set of pixels; A non-character pixel removal unit that generates a third image having, as the foreground, a set of pixels excluding the pixels specified as not constituting characters among the set of pixels; An image processing system, characterized by comprising. <7> A contour extraction procedure that extracts the contour of the foreground of the first image and extracts the contour of the foreground of the second image from which a set of pixels estimated to constitute characters has been extracted from the foreground of the first image; A non-character pixel specifying procedure that specifies, based on a comparison between the contour extracted for the first image and the contour extracted for the second image, pixels that do not constitute characters among the set of pixels; A non-character pixel removal procedure that generates a third image having, as the foreground, a set of pixels excluding the pixels specified as not constituting characters among the set of pixels; A program, characterized by causing a computer to execute.

Explanation of Signs

[0097] 10 Image processing device 11 Input image acquisition unit 12 Input image correction unit 13 Character extraction unit 14 Character extraction image correction unit 15 Image processing unit 16 Communication unit 17 Output unit 20 Cloud server 100 Drive device 101 Recording medium 102 Auxiliary storage device 103 Memory device 104 Processor 105 Interface device 121 Input image memory unit 122 Corrected input image memory unit 123 Memory unit for character extraction 124 Character extraction image memory unit 125 Corrected character extraction image memory unit 126 Output image memory unit 141 Binarization unit 142 Contour extraction unit 143 Non-character pixel identification unit 144 Non-character pixel removal unit B bus

Prior art documents

Patent documents

[0098]

Patent Document 1

Claims

1. A contour extraction unit that extracts the contour of the foreground of the first image and extracts the contour of the foreground of the second image from which a set of pixels estimated to constitute characters has been extracted from the foreground of the first image; A non-character pixel identification unit that identifies non-character pixels among the set of pixels based on a comparison between the contour extracted for the first image and the contour extracted for the second image; A non-character pixel removal unit that generates a third image having, as the foreground, a set of pixels excluding the pixels identified as not constituting characters among the set of pixels; An image processing apparatus, characterized by comprising the above.

2. In the process where the contour extraction unit extracts the contour of the foreground of the second image one pixel at a time, when the pixel extracted next to a certain pixel is different from the pixel extracted next to the pixel corresponding to the certain pixel in the first image, all the pixels forming a straight line traced back from the certain pixel among the pixels extracted as the contour for the second image are identified as non-character pixels. The image processing apparatus according to claim 1, characterized by the above.

3. When the number of pixels constituting the foreground in the direction from the pixel corresponding to the certain pixel in the first image to the next extracted pixel is equal to or greater than a predetermined value, and the number of pixels constituting the foreground in the second image in the direction opposite to the said direction is equal to or greater than a predetermined value, all the pixels forming a straight line traced back from the pixel extracted as the contour from the certain pixel are identified as non-character pixels. The image processing apparatus according to claim 2, characterized by the above.

4. When the number of pixels constituting the foreground in the direction from a certain pixel to the next pixel extracted from the second image is equal to or greater than a predetermined value, and the number of pixels constituting the foreground of the first image in the direction opposite to the said direction is equal to or greater than a predetermined value, all pixels that form a straight line by tracing back the pixels extracted as the contour from the said certain pixel are specified as pixels that do not constitute characters. The image processing apparatus according to claim 2, characterized in that.

5. The second image is an image from which a set of pixels estimated to constitute handwritten characters has been extracted from the foreground of the first image. The image processing apparatus according to claim 1, characterized in that.

6. A contour extraction unit that extracts the contour of the foreground of the first image and extracts the contour of the foreground of the second image from which a set of pixels estimated to constitute characters has been extracted from the foreground of the first image; A non-character pixel specifying unit that specifies, based on a comparison between the contour extracted for the first image and the contour extracted for the second image, pixels in the set of pixels that do not constitute characters; A non-character pixel removal unit that generates a third image having, as the foreground, a set of pixels excluding the pixels specified as not constituting characters in the set of pixels; An image processing system, characterized by comprising.

7. A contour extraction procedure for extracting the contour of the foreground of the first image and extracting the contour of the foreground of the second image from which a set of pixels estimated to constitute characters has been extracted from the foreground of the first image; A non-character pixel specifying procedure for specifying, based on a comparison between the contour extracted for the first image and the contour extracted for the second image, pixels in the set of pixels that do not constitute characters; A non-character pixel removal procedure for generating a third image having, as the foreground, a set of pixels excluding the pixels specified as not constituting characters in the set of pixels; A program characterized by causing a computer to execute.

Citation Information

Patent Citations

  • Pattern extraction device, pattern re-recognition table creation device, and pattern recognition device

    JP3345224B2